Proceedings chapter
OA Policy
English

Summarizing Sets of Categorical Sequences: selecting and visualizing representative sequences

Presented atMadeira, 6-8 October 2009
Publication date2009
Abstract

This paper is concerned with the summarization of a set of categorical sequence data. More specifically, the problem studied is the determination of the smallest possible number of representative sequences that ensure a given coverage of the whole set, i.e. that have together a given percentage of sequences in their neighborhood. The goal is to yield a representative set that exhibits the key features of the whole sequence data set and permits easy sounded interpretation. We propose an heuristic for determining the representative set that first builds a list of candidates using a representativeness score and then eliminates redundancy. We propose also a visualization tool for rendering the results and quality measures for evaluating them. The proposed tools have been implemented in TraMineR our R package for mining and visualizing sequence data and we demonstrate their efficiency on a real world example from social sciences. The methods are nonetheless by no way limited to social science data and should prove useful in many other domains.

Keywords
  • Categorical sequence data
  • Representativeness
  • Dissimilarity
  • Discrepancy of sequences
  • Summarizing
Citation (ISO format)
GABADINHO, Alexis et al. Summarizing Sets of Categorical Sequences: selecting and visualizing representative sequences. In: International Conference on Knowledge Discovery and Information Retrieval. Madeira. [s.l.] : [s.n.], 2009.
Main files (1)
Proceedings chapter
accessLevelPublic
Identifiers
  • PID : unige:4528
682views
715downloads

Technical informations

Creation01/12/2009 15:28:05
First validation01/12/2009 15:28:05
Update14/03/2023 15:19:06
Status update14/03/2023 15:19:06
Last indexation29/10/2024 12:44:53
All rights reserved by Archive ouverte UNIGE and the University of GenevaunigeBlack