An Application of Latent Semantic Analysis for Text Categorization

In this paper the authors propose a semantic approach to document categorization. The idea is to create for each category a semantic index (representative term vector) by performing a local Latent Semantic Analysis (LSA) followed by a clustering process. A second use of LSA (Global LSA) is adopted on a term-Class matrix in order to retrieve the class which is the most similar to the query (document to classify) in the same way where the LSA is used to retrieve documents which are the most similar to a query in Information Retrieval. The proposed system is evaluated on a popular dataset which is 20 Newsgroup corpus. Obtained results show the effectiveness of the method compared with those obtained with the classic KNN and SVM classifiers as well as with methods presented in the literature. Experimental results show that the new method has high precision and recall rates and classification accuracy is significantly improved.

Download Full-text

Local Latent Semantic Analysis Based on Support Vector Machine for Imbalanced Text Categorization

Communications in Computer and Information Science - Applied Informatics and Communication ◽

10.1007/978-3-642-23235-0_42 ◽

2011 ◽

pp. 321-329 ◽

Cited By ~ 1

Author(s):

Yuan Wan ◽

Hengqing Tong ◽

Yanfang Deng

Keyword(s):

Support Vector Machine ◽

Latent Semantic Analysis ◽

Text Categorization ◽

Semantic Analysis ◽

Support Vector

Download Full-text

Latent semantic analysis for text categorization using neural network

Knowledge-Based Systems ◽

10.1016/j.knosys.2008.03.045 ◽

2008 ◽

Vol 21 (8) ◽

pp. 900-904 ◽

Cited By ~ 70

Author(s):

Bo Yu ◽

Zong-ben Xu ◽

Cheng-hua Li

Keyword(s):

Neural Network ◽

Latent Semantic Analysis ◽

Text Categorization ◽

Semantic Analysis

Download Full-text

Text categorization based on combination of modified back propagation neural network and latent semantic analysis

Neural Computing and Applications ◽

10.1007/s00521-008-0193-3 ◽

2008 ◽

Vol 18 (8) ◽

pp. 875-881 ◽

Cited By ~ 24

Author(s):

Wei Wang ◽

Bo Yu

Keyword(s):

Neural Network ◽

Latent Semantic Analysis ◽

Text Categorization ◽

Semantic Analysis ◽

Back Propagation ◽

Back Propagation Neural Network

Download Full-text

Local and Global Latent Semantic Analysis for Text Categorization

Information Retrieval and Management ◽

10.4018/978-1-5225-5191-1.ch060 ◽

2018 ◽

pp. 1360-1374 ◽

Cited By ~ 1

Author(s):

Khadoudja Ghanem

Keyword(s):

Information Retrieval ◽

High Precision ◽

Latent Semantic Analysis ◽

Classification Accuracy ◽

Text Categorization ◽

Semantic Analysis ◽

Experimental Results ◽

New Method ◽

Semantic Approach ◽

Document Categorization

In this paper the authors propose a semantic approach to document categorization. The idea is to create for each category a semantic index (representative term vector) by performing a local Latent Semantic Analysis (LSA) followed by a clustering process. A second use of LSA (Global LSA) is adopted on a term-Class matrix in order to retrieve the class which is the most similar to the query (document to classify) in the same way where the LSA is used to retrieve documents which are the most similar to a query in Information Retrieval. The proposed system is evaluated on a popular dataset which is 20 Newsgroup corpus. Obtained results show the effectiveness of the method compared with those obtained with the classic KNN and SVM classifiers as well as with methods presented in the literature. Experimental results show that the new method has high precision and recall rates and classification accuracy is significantly improved.

Download Full-text

Convolutional neural networks for text categorization with latent semantic analysis

2017 International Conference on Energy, Communication, Data Analytics and Soft Computing (ICECDS) ◽

10.1109/icecds.2017.8390217 ◽

2017 ◽

Cited By ~ 4

Author(s):

Sojwal Patil ◽

Aishwarya Gune ◽

Mayura Nene

Keyword(s):

Neural Networks ◽

Convolutional Neural Networks ◽

Latent Semantic Analysis ◽

Text Categorization ◽

Semantic Analysis

Download Full-text

Improving Website Usability with Latent Semantic Analysis

PsycEXTRA Dataset ◽

10.1037/e577712012-027 ◽

2006 ◽

Author(s):

Sarah A. Nuehring ◽

Peter W. Foltz

Keyword(s):

Latent Semantic Analysis ◽

Semantic Analysis ◽

Website Usability

Download Full-text

Task Estimation Using Latent Semantic Analysis of Visual Scenes and Spoken Words

IEEJ Transactions on Electronics Information and Systems ◽

10.1541/ieejeiss.132.1473 ◽

2012 ◽

Vol 132 (9) ◽

pp. 1473-1480

Author(s):

Masashi Kimura ◽

Shinta Sawada ◽

Yurie Iribe ◽

Kouichi Katsurada ◽

Tsuneo Nitta

Keyword(s):

Latent Semantic Analysis ◽

Semantic Analysis ◽

Spoken Words ◽

Visual Scenes

Download Full-text

Similarity Detection Using Latent Semantic Analysis Algorithm

International Journal of Emerging Research in Management and Technology ◽

10.23956/ijermt.v6i8.124 ◽

2018 ◽

Vol 6 (8) ◽

pp. 102

Author(s):

Priyanka R. Patil ◽

Shital A. Patil

Keyword(s):

Latent Semantic Analysis ◽

Latent Dirichlet Allocation ◽

Semantic Analysis ◽

Mining Method ◽

Research Papers ◽

Information Measures ◽

Automated Software ◽

Day By Day ◽

Ways Of Life ◽

Dirichlet Allocation

Similarity View is an application for visually comparing and exploring multiple models of text and collection of document. Friendbook finds ways of life of clients from client driven sensor information, measures the closeness of ways of life amongst clients, and prescribes companions to clients if their ways of life have high likeness. Roused by demonstrate a clients day by day life as life records, from their ways of life are separated by utilizing the Latent Dirichlet Allocation Algorithm. Manual techniques can't be utilized for checking research papers, as the doled out commentator may have lacking learning in the exploration disciplines. For different subjective views, causing possible misinterpretations. An urgent need for an effective and feasible approach to check the submitted research papers with support of automated software. A method like text mining method come to solve the problem of automatically checking the research papers semantically. The proposed method to finding the proper similarity of text from the collection of documents by using Latent Dirichlet Allocation (LDA) algorithm and Latent Semantic Analysis (LSA) with synonym algorithm which is used to find synonyms of text index wise by using the English wordnet dictionary, another algorithm is LSA without synonym used to find the similarity of text based on index. LSA with synonym rate of accuracy is greater when the synonym are consider for matching.

Download Full-text

LATENT-SEMANTIC ANALYSIS, SOCIAL NETWORKS AND NON-STRUCTURED DATA: INTERACTION METHOD

Visnyk Universytetu “Ukraina” ◽

10.36994/2707-4110-2019-2-23-29 ◽

2019 ◽

Keyword(s):

Social Networks ◽

Latent Semantic Analysis ◽

Semantic Analysis ◽

Statistical Processing ◽

Unstructured Data ◽

Average Score ◽

Text Type ◽

Internet Users ◽

The Matrix ◽

Human Thinking

This article examines the method of latent-semantic analysis, its advantages, disadvantages, and the possibility of further transformation for use in arrays of unstructured data, which make up most of the information that Internet users deal with. To extract context-dependent word meanings through the statistical processing of large sets of textual data, an LSA method is used, based on operations with numeric matrices of the word-text type, the rows of which correspond to words, and the columns of text units to texts. The integration of words into themes and the representation of text units in the theme space is accomplished by applying one of the matrix expansions to the matrix data: singular decomposition or factorization of nonnegative matrices. The results of LSA studies have shown that the content of the similarity of words and text is obtained in such a way that the results obtained closely coincide with human thinking. Based on the methods described above, the author has developed and proposed a new way of finding semantic links between unstructured data, namely, information on social networks. The method is based on latent-semantic and frequency analyzes and involves processing the search result received, splitting each remaining text (post) into separate words, each of which takes the round in n words right and left, counting the number of occurrences of each term, working with a pre-created semantic resource (dictionary, ontology, RDF schema, ...). The developed method and algorithm have been tested on six well-known social networks, the interaction of which occurs through the ARI of the respective social networks. The average score for author's results exceeded that of their own social network search. The results obtained in the course of this dissertation can be used in the development of recommendation, search and other systems related to the search, rubrication and filtering of information.

Download Full-text