C-HTS: A Concept-based Hierarchical Text Segmentation Approach
File Type:
PDFItem Type:
Conference PaperDate:
2018Access:
openAccessCitation:
Mostafa Bayomi and Seamus Lawless, C-HTS: A Concept-based Hierarchical Text Segmentation Approach, Language Resources and Evaluation Conference, LREC 2018, Miyazaki, Japan, 7th-12th May 2018, 2018Download Item:
Abstract:
Hierarchical Text Segmentation is the task of building a hierarchical structure out of text to reflect its sub-topic hierarchy. Current text segmentation approaches are based upon using lexical and/or syntactic similarity to identify the coherent segments of text. However, the relationship between segments may be semantic, rather than lexical or syntactic. In this paper we propose C-HTS, a Concept-based Hierarchical Text Segmentation approach that uses the semantic relatedness between text constituents. In this approach, we use the explicit semantic representation of text, a method that replaces keyword-based text representation with concept-based features, automatically extracted from massive human knowledge repositories such as Wikipedia. C-HTS represents the meaning of a piece of text as a weighted vector of knowledge concepts, in order to reason about text. We evaluate the performance of C-HTS on two publicly available datasets. The results show that C-HTS compares favourably with previous state-of-the-art approaches. As Wikipedia is continuously growing, we measured the impact of its growth on segmentation performance. We used three different snapshots of Wikipedia from different years in order to achieve this. The experimental results show that an increase in the size of the knowledge base leads, on average, to greater improvements in hierarchical text segmentation.
Sponsor
Grant Number
Science Foundation Ireland (SFI)
13/RC/2106
Author's Homepage:
http://people.tcd.ie/selawleshttp://people.tcd.ie/bayomim
Author: Lawless, Seamus; BAYOMI, MOSTAFA MOHAMED
Sponsor:
Science Foundation Ireland (SFI)Other Titles:
Language Resources and Evaluation Conference, LREC 2018Type of material:
Conference PaperCollections
Availability:
Full text availableKeywords:
Hierarchical Text Segmentation, Explicit Semantic Analysis, Semantic Relatedness, WikipediaSubject (TCD):
Digital Engagement , Natural Language Processing , SEMANTIC ANALYSIS , SEMANTIC WEB , Text SegmentationMetadata
Show full item recordLicences: