z-logo
open-access-imgOpen Access
How to interpret PubMed queries and why it matters
Author(s) -
Yeganova Lana,
Comeau Donald C.,
Kim Won,
Wilbur W. John
Publication year - 2009
Publication title -
journal of the american society for information science and technology
Language(s) - English
Resource type - Journals
eISSN - 1532-2890
pISSN - 1532-2882
DOI - 10.1002/asi.20979
Subject(s) - computer science , phrase , information retrieval , conjunction (astronomy) , class (philosophy) , parsing , relevance (law) , quality (philosophy) , fraction (chemistry) , simple (philosophy) , natural language processing , artificial intelligence , philosophy , chemistry , physics , organic chemistry , epistemology , astronomy , political science , law
A significant fraction of queries in PubMed™ are multiterm queries without parsing instructions. Generally, search engines interpret such queries as collections of terms, and handle them as a Boolean conjunction of these terms. However, analysis of queries in PubMed™ indicates that many such queries are meaningful phrases, rather than simple collections of terms. In this study, we examine whether or not it makes a difference, in terms of retrieval quality, if such queries are interpreted as a phrase or as a conjunction of query terms. And, if it does, what is the optimal way of searching with such queries. To address the question, we developed an automated retrieval evaluation method, based on machine learning techniques, that enables us to evaluate and compare various retrieval outcomes. We show that the class of records that contain all the search terms, but not the phrase, qualitatively differs from the class of records containing the phrase. We also show that the difference is systematic, depending on the proximity of query terms to each other within the record. Based on these results, one can establish the best retrieval order for the records. Our findings are consistent with studies in proximity searching.

The content you want is available to Zendy users.

Already have an account? Click here to sign in.
Having issues? You can contact us here