Analisis Perbandingan Metode Vector Space Model dan Weighted Tree Similarity dengan Cosine Similarity pada kasus Pencarian Informasi Pedoman Pengobatan Dasar di Puskesmas

<p class="keywords">Sistem pencarian merupakan salah satu solusi yang dapat membantu dalam mendapatkan informasi yang diinginkan. Dengan sistem pencarian, proses pencarian informasi akan menjadi lebih efisien. Sistem pencarian informasi pada ebook pedoman pengobatan di puskesmas sangat dibutuhkan karena terdapat banyak data penyakit di dalamnya. Dalam mengembangkan sistem pencarian pada pedoman pengobatan di puskesmas, dapat memanfaatkan metode Vector Space Model (VSM) atau Weighted Tree Similarity (WTS). Penelitian ini membandingkan metode VSM dengan WTS untuk mendapatkan metode terbaik. Selain itu, ditambahkan algoritma Hamming Distance untuk mengetahui pengaruh eksekusi waktu sistem.</p><p class="keywords">Penelitian ini menunjukkan bahwa WTS lebih baik dibandingkan VSM, hal ini dapat dilihat pada hasil pengujian, nilai precision pada WTS lebih baik dibandingkan VSM. Karena pada metode pencarian yang efektif adalah yang memberikan nilai ketepatan(precision) terbaik, meskipun nilai recall lebih rendah. Pada pengujian sistem, VSM menunjukkan hasil nilai rata – rata precision sebesar 44,82983 % dan recall sebesar 99,08165 %. Sedangkan pada WTS nilai rata – rata precision sebesar 52,17332% dan recall sebesar 98,61761%. Kemudian pada pengujian pakar menunjukkan precision WTS dengan rata – rata sebesar 46,675% dan recall sebesar 73,6111%. Sedangkan nilai precision VSM sebesar 33,6737% dan nilai recall sebesar 86,8056%.</p>Algoritma Hamming Distance sangat membantu dalam mempercepat eksekusi sistem. Pengaruh penggunaan algoritma Hamming Distance pada VSM memberikan hasil denganrata – rata waktu pengujian adalah 4,512 detik, sedangkan tanpa Hamming Distance adalah 9,185 detik. Kemudian Pada metode WTS dengan Hamming Distance memberikan hasil rata – rata dengan waktu pengujian adalah 6,042 detik, sedangkan tanpa Hamming Distance adalah 14,421 detik

Download Full-text

ASPECT BASED SENTIMENT ANALYSIS DATA KUESIONER DI RUMAH SAKIT MUHAMMADIYAH LAMONGAN MENGGUNAKAN ALGORITMA K-NN.

JOUTICA ◽

10.30736/jti.v6i2.677 ◽

2021 ◽

Vol 6 (2) ◽

pp. 506

Author(s):

Mustain Mustain Mustain

Keyword(s):

Vector Space ◽

Sentiment Analysis ◽

Nearest Neighbor ◽

Vector Space Model ◽

Analysis Data ◽

Cosine Similarity ◽

K Nearest Neighbor ◽

Space Model

Kesulitan untuk mengorganisir data kuesioner yang bersifat konvensional melatarbelakangi penelitian ini. Oleh karena itu dibuat sistem yang memudahkan pengelompokan data kuesioner secara otomatis yang lengkap dengan sentimen yang terkandung didalamnya. Dataset yang digunakan dalam penelitian ini adalah data kuesioner rumah sakit Muhammadiyah lamongan. Penelitian ini hanya menangani kuesioner yang berbentuk teks. Data dengan fisik kertas direkap kemudian diinput ke database lengkap dengan kategori unit kerja dan sentiment. Selanjutnya dataset tersebut di dilakukan pre-prosesing yang meliputi penanganan negasi case folding, tokenizing, filtering dan stemming. Sebagai data uji komentar dari kuesioner akan dilakukan pre-prosesing selanjutnya dihitung tingkat kemiripan document dengan menggunakan metode K- Nearest Neighbor dan Vector Space Model. Jumlah data yang ditangani mempengaruhi performa system terutama dari akurasi dan kecepatan pada saat proses klasifikasi. Hasil dari sistem yang dibuat berupa ranking dokumen yang paling mirip dengan dataset berdasarkan urutan nilai cosine similarity. Ujicoba klasifikasi berdasarkan kelas kategori menghasilkan nilai akurasi 91 %. Ujicoba berdasarkan Kelas Sentimen sebesar 94 %.dari kombinasi keduanya system berhasil mendapat akurasi sebesar 86 %

Download Full-text

Myanmar News Retrieval in Vector Space Model using Cosine Similarity Measure

2020 IEEE Conference on Computer Applications(ICCA) ◽

10.1109/icca49400.2020.9022845 ◽

2020 ◽

Author(s):

Hay Man Oo ◽

Win Pa Pa

Keyword(s):

Vector Space ◽

Similarity Measure ◽

Vector Space Model ◽

Cosine Similarity ◽

Space Model ◽

Cosine Similarity Measure ◽

News Retrieval

Download Full-text

Automating case definitions using literature-based reasoning

Applied Clinical Informatics ◽

10.4338/aci-2013-04-ra-0028 ◽

2013 ◽

Vol 04 (04) ◽

pp. 515-527 ◽

Cited By ~ 4

Author(s):

R. Ball ◽

T. Botsis

Keyword(s):

Vector Space ◽

Semantic Network ◽

Vector Space Model ◽

Classification Performance ◽

Case Definition ◽

Cosine Similarity ◽

Case Definitions ◽

Space Model ◽

Research Activities ◽

Clinical Surveillance

SummaryBackground: Establishing a Case Definition (CDef) is a first step in many epidemiological, clinical, surveillance, and research activities. The application of CDefs still relies on manual steps and this is a major source of inefficiency in surveillance and research.Objective: Describe the need and propose an approach for automating the useful representation of CDefs for medical conditions.Methods: We translated the existing Brighton Collaboration CDef for anaphylaxis by mostly relying on the identification of synonyms for the criteria of the CDef using the NLM MetaMap tool. We also generated a CDef for the same condition using all the related PubMed abstracts, processing them with a text mining tool, and further treating the synonyms with the above strategy. The co-occur-rence of the anaphylaxis and any other medical term within the same sentence of the abstracts supported the construction of a large semantic network. The ‘islands’ algorithm reduced the network and revealed its densest region including the nodes that were used to represent the key criteria of the CDef. We evaluated the ability of the “translated” and the “generated” CDef to classify a set of 6034 H1N1 reports for anaphylaxis using two similarity approaches and comparing them with our previous semi-automated classification approach.Results: Overall classification performance across approaches to producing CDefs was similar, with the generated CDef and vector space model with cosine similarity having the highest accuracy (0.825±0.003) and the semi-automated approach and vector space model with cosine similarity having the highest recall (0.809±0.042). Precision was low for all approaches.Conclusion: The useful representation of CDefs is a complicated task but potentially offers substantial gains in efficiency to support safety and clinical surveillance.Citation: Botsis T, Ball R. Automating case definitions using literature-based reasoning. Appl Clin Inf 2013; 4: 515–527http://dx.doi.org/10.4338/ACI-2013-04-RA-0028

Download Full-text

Information retrieval from heterogeneous data sets using moderated IDF-cosine similarity in vector space model

2017 International Conference on Energy, Communication, Data Analytics and Soft Computing (ICECDS) ◽

10.1109/icecds.2017.8390174 ◽

2017 ◽

Cited By ~ 1

Author(s):

Bhagyashree Pathak ◽

Niranjan Lal

Keyword(s):

Information Retrieval ◽

Vector Space ◽

Vector Space Model ◽

Heterogeneous Data ◽

Cosine Similarity ◽

Data Sets ◽

Space Model

Download Full-text

Measuring the Level of Plagiarism of Thesis using Vector Space Model and Cosine Similarity Methods

IOP Conference Series Materials Science and Engineering ◽

10.1088/1757-899x/662/2/022111 ◽

2019 ◽

Vol 662 ◽

pp. 022111

Author(s):

I Indriyanto ◽

I D Sumitra

Keyword(s):

Vector Space ◽

Vector Space Model ◽

Cosine Similarity ◽

Space Model

Download Full-text

PENENTUAN MULTIPLE MEMBERSHIP DOKUMEN

Majalah Ilmiah UNIKOM ◽

10.34010/miu.v15i2.560 ◽

2017 ◽

Vol 15 (2) ◽

Author(s):

Stephanie Betha R.H

Keyword(s):

Vector Space ◽

Vector Space Model ◽

Cosine Similarity ◽

Space Model ◽

Multiple Membership

Multiple membership merupakan keanggotaan yang dimiliki oleh seseorang pada beberapa komunitas. Multiple membership pada dokumen artinya suatu dokumen dapat mengandung konten dari beberapa jenis kategori. Jenis kategori pada dokumen dapat ditentukan dengan mengukur kemiripan dokumen tersebut dengan kategori yang ada. Vector Space Model adalah suatu model yang digunakan untuk mengukur kemiripan antara suatu dokumen dan suatu query dengan mewakili setiap dokumen dalam sebuah koleksi sebagai sebuah titik dalam ruang vektor. Hasil dari pengukuran kemiripan tersebut merupakan nilai cosine similarity antara vektor query dari dokumen terhadap vektor kategori. Permasalahan yang terjadi adalah suatu pengukuran kemiripan vektor query dokumen, dapat menghasilkan nilai cosine similarity dengan selisih yang kecil antara vektor kategori satu dengan vektor kategori lain. Hal ini menyebabkan kedua vektor kategori tersebut menjadi saling dominan satu sama lain pada dokumen. Oleh karena itu, dibutuhkan suatu nilai batas untuk menentukan kondisi kapan suatu vektor kategori dapat dinyatakan sebagai vektor kategori yang saling dominan. Penetapan nilai batas ini menggunakan K-Means Clustering. Nilai batas ini ditetapkan berdasarkan pengelompokkan nilai jarak antar presentase cosine similarity pada suatu dokumen. Penentuan multiple membership dokumen ini akan dilakukan pada atribut judul dan kata kunci pada dokumen publikasi ilmiah.

Download Full-text

Information Retrieval for Gujarati Language Using Cosine Similarity Based Vector Space Model

Advances in Intelligent Systems and Computing - Proceedings of the 5th International Conference on Frontiers in Intelligent Computing: Theory and Applications ◽

10.1007/978-981-10-3156-4_1 ◽

2017 ◽

pp. 1-9 ◽

Cited By ~ 3

Author(s):

Rajnish M. Rakholia ◽

Jatinderkumar R. Saini

Keyword(s):

Information Retrieval ◽

Vector Space ◽

Vector Space Model ◽

Cosine Similarity ◽

Space Model ◽

Gujarati Language

Download Full-text

A Comparative Study on Cosine Similarity Algorithm and Vector Space Model Algorithm on Document Searching

Advanced Science Letters ◽

10.1166/asl.2015.6481 ◽

2015 ◽

Vol 21 (10) ◽

pp. 3321-3323

Author(s):

Warnia Nengsih

Keyword(s):

Comparative Study ◽

Vector Space ◽

Vector Space Model ◽

Cosine Similarity ◽

Space Model ◽

Similarity Algorithm ◽

Model Algorithm

Download Full-text

Word Sense Disambiguation Using Cosine Similarity Collaborates with Word2vec and WordNet

Future Internet ◽

10.3390/fi11050114 ◽

2019 ◽

Vol 11 (5) ◽

pp. 114 ◽

Cited By ~ 5

Author(s):

Korawit Orkphol ◽

Wu Yang

Keyword(s):

Vector Space ◽

Language Processing ◽

Semantic Analysis ◽

Word Sense Disambiguation ◽

Vector Space Model ◽

Word Embedding ◽

Cosine Similarity ◽

Word Sense ◽

Lexical Database ◽

Space Model

Words have different meanings (i.e., senses) depending on the context. Disambiguating the correct sense is important and a challenging task for natural language processing. An intuitive way is to select the highest similarity between the context and sense definitions provided by a large lexical database of English, WordNet. In this database, nouns, verbs, adjectives, and adverbs are grouped into sets of cognitive synonyms interlinked through conceptual semantics and lexicon relations. Traditional unsupervised approaches compute similarity by counting overlapping words between the context and sense definitions which must match exactly. Similarity should compute based on how words are related rather than overlapping by representing the context and sense definitions on a vector space model and analyzing distributional semantic relationships among them using latent semantic analysis (LSA). When a corpus of text becomes more massive, LSA consumes much more memory and is not flexible to train a huge corpus of text. A word-embedding approach has an advantage in this issue. Word2vec is a popular word-embedding approach that represents words on a fix-sized vector space model through either the skip-gram or continuous bag-of-words (CBOW) model. Word2vec is also effectively capturing semantic and syntactic word similarities from a huge corpus of text better than LSA. Our method used Word2vec to construct a context sentence vector, and sense definition vectors then give each word sense a score using cosine similarity to compute the similarity between those sentence vectors. The sense definition also expanded with sense relations retrieved from WordNet. If the score is not higher than a specific threshold, the score will be combined with the probability of that sense distribution learned from a large sense-tagged corpus, SEMCOR. The possible answer senses can be obtained from high scores. Our method shows that the result (50.9% or 48.7% without the probability of sense distribution) is higher than the baselines (i.e., original, simplified, adapted and LSA Lesk) and outperforms many unsupervised systems participating in the SENSEVAL-3 English lexical sample task.

Download Full-text

Pemanfaatan Metode Vector Space Model dan Metode Cosine Similarity pada Fitur Deteksi Hama dan Penyakit Tanaman Padi

Jurnal Teknologi & Informasi ITSmart ◽

10.20961/its.v3i2.704 ◽

2016 ◽

Vol 3 (2) ◽

pp. 90 ◽

Cited By ~ 1

Author(s):

Ana Triana ◽

Ristu Saptono ◽

Meiyanto Eko Sulistyo

Keyword(s):

Vector Space ◽

Vector Space Model ◽

Cosine Similarity ◽

Space Model

Download Full-text