A Comparative Study on Cosine Similarity Algorithm and Vector Space Model Algorithm on Document Searching

Kesulitan untuk mengorganisir data kuesioner yang bersifat konvensional melatarbelakangi penelitian ini. Oleh karena itu dibuat sistem yang memudahkan pengelompokan data kuesioner secara otomatis yang lengkap dengan sentimen yang terkandung didalamnya. Dataset yang digunakan dalam penelitian ini adalah data kuesioner rumah sakit Muhammadiyah lamongan. Penelitian ini hanya menangani kuesioner yang berbentuk teks. Data dengan fisik kertas direkap kemudian diinput ke database lengkap dengan kategori unit kerja dan sentiment. Selanjutnya dataset tersebut di dilakukan pre-prosesing yang meliputi penanganan negasi case folding, tokenizing, filtering dan stemming. Sebagai data uji komentar dari kuesioner akan dilakukan pre-prosesing selanjutnya dihitung tingkat kemiripan document dengan menggunakan metode K- Nearest Neighbor dan Vector Space Model. Jumlah data yang ditangani mempengaruhi performa system terutama dari akurasi dan kecepatan pada saat proses klasifikasi. Hasil dari sistem yang dibuat berupa ranking dokumen yang paling mirip dengan dataset berdasarkan urutan nilai cosine similarity. Ujicoba klasifikasi berdasarkan kelas kategori menghasilkan nilai akurasi 91 %. Ujicoba berdasarkan Kelas Sentimen sebesar 94 %.dari kombinasi keduanya system berhasil mendapat akurasi sebesar 86 %

Download Full-text

A Comprehensive Comparative Study Using Vector Space Model with K-Nearest Neighbor on Text Categorization Data

Asian Journal of Information Management ◽

10.3923/ajim.2008.14.22 ◽

2007 ◽

Vol 2 (1) ◽

pp. 14-22 ◽

Cited By ~ 1

Author(s):

Wa`el Musa Hadi ◽

Fadi Thabtah ◽

Salahideen Mousa ◽

Samer Al Hawari ◽

Ghassan Kanaan ◽

...

Keyword(s):

Comparative Study ◽

Vector Space ◽

Text Categorization ◽

Nearest Neighbor ◽

Vector Space Model ◽

K Nearest Neighbor ◽

Space Model

Download Full-text

Myanmar News Retrieval in Vector Space Model using Cosine Similarity Measure

2020 IEEE Conference on Computer Applications(ICCA) ◽

10.1109/icca49400.2020.9022845 ◽

2020 ◽

Author(s):

Hay Man Oo ◽

Win Pa Pa

Keyword(s):

Vector Space ◽

Similarity Measure ◽

Vector Space Model ◽

Cosine Similarity ◽

Space Model ◽

Cosine Similarity Measure ◽

News Retrieval

Download Full-text

Automating case definitions using literature-based reasoning

Applied Clinical Informatics ◽

10.4338/aci-2013-04-ra-0028 ◽

2013 ◽

Vol 04 (04) ◽

pp. 515-527 ◽

Cited By ~ 4

Author(s):

R. Ball ◽

T. Botsis

Keyword(s):

Vector Space ◽

Semantic Network ◽

Vector Space Model ◽

Classification Performance ◽

Case Definition ◽

Cosine Similarity ◽

Case Definitions ◽

Space Model ◽

Research Activities ◽

Clinical Surveillance

SummaryBackground: Establishing a Case Definition (CDef) is a first step in many epidemiological, clinical, surveillance, and research activities. The application of CDefs still relies on manual steps and this is a major source of inefficiency in surveillance and research.Objective: Describe the need and propose an approach for automating the useful representation of CDefs for medical conditions.Methods: We translated the existing Brighton Collaboration CDef for anaphylaxis by mostly relying on the identification of synonyms for the criteria of the CDef using the NLM MetaMap tool. We also generated a CDef for the same condition using all the related PubMed abstracts, processing them with a text mining tool, and further treating the synonyms with the above strategy. The co-occur-rence of the anaphylaxis and any other medical term within the same sentence of the abstracts supported the construction of a large semantic network. The ‘islands’ algorithm reduced the network and revealed its densest region including the nodes that were used to represent the key criteria of the CDef. We evaluated the ability of the “translated” and the “generated” CDef to classify a set of 6034 H1N1 reports for anaphylaxis using two similarity approaches and comparing them with our previous semi-automated classification approach.Results: Overall classification performance across approaches to producing CDefs was similar, with the generated CDef and vector space model with cosine similarity having the highest accuracy (0.825±0.003) and the semi-automated approach and vector space model with cosine similarity having the highest recall (0.809±0.042). Precision was low for all approaches.Conclusion: The useful representation of CDefs is a complicated task but potentially offers substantial gains in efficiency to support safety and clinical surveillance.Citation: Botsis T, Ball R. Automating case definitions using literature-based reasoning. Appl Clin Inf 2013; 4: 515–527http://dx.doi.org/10.4338/ACI-2013-04-RA-0028

Download Full-text

Information retrieval from heterogeneous data sets using moderated IDF-cosine similarity in vector space model

2017 International Conference on Energy, Communication, Data Analytics and Soft Computing (ICECDS) ◽

10.1109/icecds.2017.8390174 ◽

2017 ◽

Cited By ~ 1

Author(s):

Bhagyashree Pathak ◽

Niranjan Lal

Keyword(s):

Information Retrieval ◽

Vector Space ◽

Vector Space Model ◽

Heterogeneous Data ◽

Cosine Similarity ◽

Data Sets ◽

Space Model

Download Full-text

Text similarity algorithm based on semantic vector space model

2016 IEEE/ACIS 15th International Conference on Computer and Information Science (ICIS) ◽

10.1109/icis.2016.7550928 ◽

2016 ◽

Cited By ~ 5

Author(s):

LiHong Xu ◽

ShuTao Sun ◽

Qi Wang

Keyword(s):

Vector Space ◽

Vector Space Model ◽

Text Similarity ◽

Space Model ◽

Similarity Algorithm

Download Full-text

An Efficient Journal Articles Searching using Vector Space Model Algorithm

IJID (International Journal on Informatics for Development) ◽

10.14421/ijid.2020.09104 ◽

2020 ◽

Vol 9 (1) ◽

pp. 21

Author(s):

Azis Alvriyanto ◽

Muhammad Taufiq Nuruzzaman ◽

Maria Ulfah Siregar ◽

Rahmat Hidayat

Keyword(s):

Vector Space ◽

Evaluation Criteria ◽

Vector Space Model ◽

Journal Articles ◽

Space Model ◽

Search Results ◽

Document Frequency ◽

Traditional Algorithm ◽

Baseline Algorithm ◽

Model Algorithm

One of the main feature of digital library is a search engine which depends on keywords submitted by a user. However, in the traditional algorithm, the computation performance, searching speed, significantly relies on the number of journal articles stored in the databases. Some irrelevant search results also increase the speed of article searching process. To solve the problem, in this paper we propose vector space model (VSM) algorithm to search for relevant journal articles. The VSM algorithm considers a term frequency - inversed document frequency (TF-IDF). The VSM algorithm will be compared to the baseline algorithm namely traditional algorithm. Both algorithms will be evaluated using combination of keywords which can be a synonym, phrase, error typography, or suffix and prefix. By using the data consist of 635 journal articles, both algorithms are compared in terms of 11 evaluation criteria. The results show that VSM algorithm is able to obtain the intended journal at 5th rank on average as compared to the traditional algorithm which can obtain the intended journal at rank of 171st on average. Therefore, our proposed algorithm can improve the performance to accurately sort the journal articles based on the submitted keywords as compared to traditional algorithm.

Download Full-text

Measuring the Level of Plagiarism of Thesis using Vector Space Model and Cosine Similarity Methods

IOP Conference Series Materials Science and Engineering ◽

10.1088/1757-899x/662/2/022111 ◽

2019 ◽

Vol 662 ◽

pp. 022111

Author(s):

I Indriyanto ◽

I D Sumitra

Keyword(s):

Vector Space ◽

Vector Space Model ◽

Cosine Similarity ◽

Space Model

Download Full-text

PENENTUAN MULTIPLE MEMBERSHIP DOKUMEN

Majalah Ilmiah UNIKOM ◽

10.34010/miu.v15i2.560 ◽

2017 ◽

Vol 15 (2) ◽

Author(s):

Stephanie Betha R.H

Keyword(s):

Vector Space ◽

Vector Space Model ◽

Cosine Similarity ◽

Space Model ◽

Multiple Membership

Multiple membership merupakan keanggotaan yang dimiliki oleh seseorang pada beberapa komunitas. Multiple membership pada dokumen artinya suatu dokumen dapat mengandung konten dari beberapa jenis kategori. Jenis kategori pada dokumen dapat ditentukan dengan mengukur kemiripan dokumen tersebut dengan kategori yang ada. Vector Space Model adalah suatu model yang digunakan untuk mengukur kemiripan antara suatu dokumen dan suatu query dengan mewakili setiap dokumen dalam sebuah koleksi sebagai sebuah titik dalam ruang vektor. Hasil dari pengukuran kemiripan tersebut merupakan nilai cosine similarity antara vektor query dari dokumen terhadap vektor kategori. Permasalahan yang terjadi adalah suatu pengukuran kemiripan vektor query dokumen, dapat menghasilkan nilai cosine similarity dengan selisih yang kecil antara vektor kategori satu dengan vektor kategori lain. Hal ini menyebabkan kedua vektor kategori tersebut menjadi saling dominan satu sama lain pada dokumen. Oleh karena itu, dibutuhkan suatu nilai batas untuk menentukan kondisi kapan suatu vektor kategori dapat dinyatakan sebagai vektor kategori yang saling dominan. Penetapan nilai batas ini menggunakan K-Means Clustering. Nilai batas ini ditetapkan berdasarkan pengelompokkan nilai jarak antar presentase cosine similarity pada suatu dokumen. Penentuan multiple membership dokumen ini akan dilakukan pada atribut judul dan kata kunci pada dokumen publikasi ilmiah.

Download Full-text

Information Retrieval for Gujarati Language Using Cosine Similarity Based Vector Space Model

Advances in Intelligent Systems and Computing - Proceedings of the 5th International Conference on Frontiers in Intelligent Computing: Theory and Applications ◽

10.1007/978-981-10-3156-4_1 ◽

2017 ◽

pp. 1-9 ◽

Cited By ~ 3

Author(s):

Rajnish M. Rakholia ◽

Jatinderkumar R. Saini

Keyword(s):

Information Retrieval ◽

Vector Space ◽

Vector Space Model ◽

Cosine Similarity ◽

Space Model ◽

Gujarati Language

Download Full-text