regularized em
Recently Published Documents


TOTAL DOCUMENTS

9
(FIVE YEARS 2)

H-INDEX

2
(FIVE YEARS 0)

Author(s):  
Anastasia Ianina ◽  
Konstantin Vorontsov

Real-time monitoring of scientific papers and technological news requires fast processing of complicated search demands motivated by thematically relevant information acquisition. For this case, the authors develop an exploratory search engine based on probabilistic hierarchical topic modeling. Topic model gives a low dimensional sparse interpretable vector representation (topical embedding) of a text, which is used for ranking documents by their similarity to the query. They explore several ways of comparing topical vectors including searching with thematically homogeneous text segments. Topical hierarchies are built using the regularized EM-algorithm from BigARTM project. The topic-based search achieves better precision and recall than other approaches (TF-IDF, fastText, LSTM, BERT) and even human assessors who spend up to an hour to complete the same search task. They also discover that blending hierarchical topic vectors with neural pretrained embeddings is a promising way of enriching both models that helps to get precision and recall higher than 90%.


IEEE Access ◽  
2020 ◽  
Vol 8 ◽  
pp. 211576-211584
Author(s):  
Junwei Shi ◽  
Daiki Hara ◽  
Wensi Tao ◽  
Nesrin Dogan ◽  
Alan Pollack ◽  
...  

2016 ◽  
Vol 23 (4) ◽  
pp. 327-351
Author(s):  
Hidetaka Kamigaito ◽  
Taro Watanabe ◽  
Hiroya Takamura ◽  
Manabu Okumura ◽  
Eiichiro Sumita

2014 ◽  
Author(s):  
Hidetaka Kamigaito ◽  
Taro Watanabe ◽  
Hiroya Takamura ◽  
Manabu Okumura

2010 ◽  
Author(s):  
Soo-Jin Lee ◽  
Van-Giang Nguyen ◽  
Mi No Lee
Keyword(s):  

Sign in / Sign up

Export Citation Format

Share Document