Logical structure based semantic relationship extraction from semi-structured documents

Structured documents are composed of objects with a content and a logical structure. The effective retrieval of structured documents requires models that provide for a content-based retrieval of objects that takes into account their logical structure, so that the relevance of an object is not solely based on its content, but also on the logical structure among objects. This paper proposes a formal model for representing structured documents where the content of an object is viewed as the knowledge contained in that object, and the logical structure among objects is capture by a process of knowledge augmentation: the knowledge contained in an object is augmented with that of its structurally related objects. The knowledge augmentation process takes into account the fact that knowledge can be incomplete and become inconsistent.

Download Full-text

Semi-Structured Document Classification

Encyclopedia of Data Warehousing and Mining, Second Edition ◽

10.4018/978-1-60566-010-3.ch271 ◽

2011 ◽

pp. 1779-1786

Author(s):

Ludovic Denoyer

Keyword(s):

Machine Learning ◽

Information Sources ◽

Major Change ◽

Logical Structure ◽

Document Classification ◽

Document Collections ◽

Heterogeneous Information ◽

Structured Document ◽

Structured Documents ◽

Different Content

Document classification developed over the last ten years, using techniques originating from the pattern recognition and machine learning communities. All these methods do operate on flat text representations where word occurrences are considered independents. The recent paper (Sebastiani, 2002) gives a very good survey on textual document classification. With the development of structured textual and multimedia documents, and with the increasing importance of structured document formats like XML, the document nature is changing. Structured documents usually have a much richer representation than flat ones. They have a logical structure. They are often composed of heterogeneous information sources (e.g. text, image, video, metadata, etc). Another major change with structured documents is the possibility to access document elements or fragments. The development of classifiers for structured content is a new challenge for the machine learning and IR communities. A classifier for structured documents should be able to make use of the different content information sources present in an XML document and to classify both full documents and document parts. It should easily adapt to a variety of different sources (e.g. to different Document Type Definitions). It should be able to scale with large document collections.

Download Full-text

Semi-Structured Document Classification

Encyclopedia of Data Warehousing and Mining ◽

10.4018/978-1-59140-557-3.ch191 ◽

2011 ◽

pp. 1015-1021

Author(s):

Ludovic Denoyer ◽

Patrick Gallinari

Keyword(s):

Machine Learning ◽

Information Sources ◽

Major Change ◽

Logical Structure ◽

Document Classification ◽

Document Collections ◽

Heterogeneous Information ◽

Structured Document ◽

Structured Documents ◽

Different Content

Document classification developed over the last 10 years, using techniques originating from the pattern recognition and machine-learning communities. All these methods operate on flat text representations, where word occurrences are considered independents. The recent paper by Sebastiani (2002) gives a very good survey on textual document classification. With the development of structured textual and multimedia documents and with the increasing importance of structured document formats like XML, the document nature is changing. Structured documents usually have a much richer representation than flat ones. They have a logical structure. They are often composed of heterogeneous information sources (e.g., text, image, video, metadata, etc.). Another major change with structured documents is the possibility to access document elements or fragments. The development of classifiers for structured content is a new challenge for the machine-learning and IR communities. A classifier for structured documents should be able to make use of the different content information sources present in an XML document and to classify both full documents and document parts. It should adapt easily to a variety of different sources (e.g., different document type definitions). It should be able to scale with large document collections.

Download Full-text

Logical structure analysis and generation for structured documents: A syntactic approach

IEEE Transactions on Knowledge and Data Engineering ◽

10.1109/tkde.2003.1232278 ◽

2003 ◽

Vol 15 (5) ◽

pp. 1277-1294 ◽

Cited By ~ 7

Author(s):

Kyong-Ho Lee ◽

Yoon-Chul Choy ◽

Sung-Bae Cho

Keyword(s):

Structure Analysis ◽

Logical Structure ◽

Syntactic Approach ◽

Structured Documents

Download Full-text

Rewiev of current text representation technics for semantic relationship extraction

Computer Science and Mathematical Modelling ◽

10.5604/01.3001.0015.2733 ◽

2021 ◽

Vol 0 (11-12/2020) ◽

pp. 13-22

Author(s):

Michał Gałusza

Keyword(s):

Text Processing ◽

Text Representation ◽

Semantic Relationship ◽

Relationship Extraction

Article provides review on current most popular text processing technics; sketches their evolution and compares sequence and dependency models in detecting semantic relationship between words.

Download Full-text

Capturing Logical Structure of Visually Structured Documents with Multimodal Transition Parser

10.18653/v1/2021.nllp-1.15 ◽

2021 ◽

Author(s):

Yuta Koreeda ◽

Christopher Manning

Keyword(s):

Logical Structure ◽

Structured Documents

Download Full-text

Methodolo- gical Aspects of Semantic Relationship Extraction for Automatic Thesaurus Generation

Modeling and Analysis of Information Systems ◽

10.18255/1818-1015-2016-6-826-840 ◽

2016 ◽

Vol 23 (6) ◽

pp. 826-840 ◽

Cited By ~ 4

Author(s):

N. S. Lagutina ◽

K. V. Lagutina ◽

E. I. Mamedov ◽

I. V. Paramonov

Keyword(s):

Semantic Relationship ◽

Relationship Extraction

Download Full-text

Semantic Relationship Extraction and Ontology Building using Wikipedia: A Comprehensive Survey

International Journal of Computer Applications ◽

10.5120/1661-2236 ◽

2010 ◽

Vol 12 (3) ◽

pp. 6-12 ◽

Cited By ~ 3

Author(s):

Nora I. Al- Rajebah ◽

Hend S. Al- Khalifa

Keyword(s):

Semantic Relationship ◽

Relationship Extraction ◽

Comprehensive Survey ◽

Ontology Building

Download Full-text

Marking Up is Not Enough

Methods of Information in Medicine ◽

10.1055/s-0038-1634939 ◽

1993 ◽

Vol 32 (04) ◽

pp. 272-273 ◽

Cited By ~ 3

Author(s):

A. L. Rector

Keyword(s):

Health Care ◽

Intelligent Processing ◽

Structured Documents ◽

Electronic Health ◽

Health Care Records

Response to: Essin DJ. Intelligent processing of loosely structured documents as a strategy for organizing electronic health care records. Meth Inform Med 1993; 32: 265.

Download Full-text

Logical structure based semantic relationship extraction from semi-structured documents

Identifying Logical Structure and Content Structure in Loosely-Structured Documents

FOUR-VALUED KNOWLEDGE AUGMENTATION FOR STRUCTURED DOCUMENT RETRIEVAL

Semi-Structured Document Classification

Semi-Structured Document Classification

Logical structure analysis and generation for structured documents: A syntactic approach

Rewiev of current text representation technics for semantic relationship extraction

Capturing Logical Structure of Visually Structured Documents with Multimodal Transition Parser

Methodolo- gical Aspects of Semantic Relationship Extraction for Automatic Thesaurus Generation

Semantic Relationship Extraction and Ontology Building using Wikipedia: A Comprehensive Survey

Marking Up is Not Enough

Export Citation Format