probabilistic retrieval
Recently Published Documents


TOTAL DOCUMENTS

45
(FIVE YEARS 2)

H-INDEX

13
(FIVE YEARS 0)

2022 ◽  
Vol 40 (3) ◽  
pp. 1-24
Author(s):  
Jiaul H. Paik ◽  
Yash Agrawal ◽  
Sahil Rishi ◽  
Vaishal Shah

Existing probabilistic retrieval models do not restrict the domain of the random variables that they deal with. In this article, we show that the upper bound of the normalized term frequency ( tf ) from the relevant documents is much smaller than the upper bound of the normalized tf from the whole collection. As a result, the existing models suffer from two major problems: (i) the domain mismatch causes data modeling error, (ii) since the outliers have very large magnitude and the retrieval models follow tf hypothesis, the combination of these two factors tends to overestimate the relevance score. In an attempt to address these problems, we propose novel weighted probabilistic models based on truncated distributions. We evaluate our models on a set of large document collections. Significant performance improvement over six existing probabilistic models is demonstrated.


2014 ◽  
Vol 52 (3) ◽  
pp. 1892-1900 ◽  
Author(s):  
Andrea Antonini ◽  
Riccardo Benedetti ◽  
Alberto Ortolani ◽  
Luca Rovai ◽  
Giovanni Schiavon

Sign in / Sign up

Export Citation Format

Share Document