NDLI: Exploiting External Collections for Query Expansion

Please wait, while we are loading the content...

Exploiting External Collections for Query Expansion

Content Provider	ACM Digital Library
Author	De rijke, Maarten Weerkamp, Wouter Balog, Krisztian
Copyright Year	2012
Abstract	A persisting challenge in the field of information retrieval is the vocabulary mismatch between a user’s information need and the relevant documents. One way of addressing this issue is to apply query modeling: to add terms to the original query and reweigh the terms. In social media, where documents usually contain creative and noisy language (e.g., spelling and grammatical errors), query modeling proves difficult. To address this, attempts to use external sources for query modeling have been made and seem to be successful. In this article we propose a general generative query expansion model that uses external document collections for term generation: the External Expansion Model (EEM). The main rationale behind our model is our hypothesis that each query requires its own mixture of external collections for expansion and that an expansion model should account for this. For some queries we expect, for example, a news collection to be most beneficial, while for other queries we could benefit more by selecting terms from a general encyclopedia. EEM allows for query-dependent weighing of the external collections. We put our model to the test on the task of blog post retrieval and we use four external collections in our experiments: (i) a news collection, (ii) a Web collection, (iii) Wikipedia, and (iv) a blog post collection. Experiments show that EEM outperforms query expansion on the individual collections, as well as the Mixture of Relevance Models that was previously proposed by Diaz and Metzler [2006]. Extensive analysis of the results shows that our naive approach to estimating query-dependent collection importance works reasonably well and that, when we use “oracle” settings, we see the full potential of our model. We also find that the query-dependent collection importance has more impact on retrieval performance than the independent collection importance (i.e., a collection prior).
Starting Page	1
Ending Page	29
Page Count	29
File Format	PDF
ISSN	15591131
e-ISSN	1559114X
DOI	10.1145/2382616.2382621
Volume Number	6
Issue Number	4
Journal	ACM Transactions on the Web (TWEB)
Language	English
Publisher	Association for Computing Machinery (ACM)
Publisher Date	2012-11-01
Publisher Place	New York
Access Restriction	One Nation One Subscription (ONOS)
Subject Keyword	Query modeling Blog post retrieval External expansion
Content Type	Text
Resource Type	Article
Subject	Computer Networks and Communications

Sl.	Authority	Responsibilities	Communication Details
1	Ministry of Education (GoI), Department of Higher Education	Sanctioning Authority	https://www.education.gov.in/ict-initiatives
2	Indian Institute of Technology Kharagpur	Host Institute of the Project: The host institute of the project is responsible for providing infrastructure support and hosting the project	https://www.iitkgp.ac.in
3	National Digital Library of India Office, Indian Institute of Technology Kharagpur	The administrative and infrastructural headquarters of the project	Dr. B. Sutradhar bsutra@ndl.gov.in
4	Project PI / Joint PI	Principal Investigator and Joint Principal Investigators of the project	Dr. B. Sutradhar bsutra@ndl.gov.in Prof. Saswat Chakrabarti will be added soon
5	Website/Portal (Helpdesk)	Queries regarding NDLI and its services	support@ndl.gov.in
6	Contents and Copyright Issues	Queries related to content curation and copyright issues	content@ndl.gov.in
7	National Digital Library of India Club (NDLI Club)	Queries related to NDLI Club formation, support, user awareness program, seminar/symposium, collaboration, social media, promotion, and outreach	clubsupport@ndl.gov.in
8	Digital Preservation Centre (DPC)	Assistance with digitizing and archiving copyright-free printed books	dpc@ndl.gov.in
9	IDR Setup or Support	Queries related to establishment and support of Institutional Digital Repository (IDR) and IDR workshops	idr@ndl.gov.in

Exploiting external collections for query expansion (2012)

A generative blog post retrieval model that uses query expansion based on external collections.

A Generative Blog Post Retrieval Model that Uses Query Expansion based on External Collections

External Query Expansion in the Blogosphere

A generative blog post retrieval model that uses query expansion based on external collections (2009)

Exploiting query logs for cross-lingual query suggestions

Document representation and query expansion models for blog recommendation.

A Combined Query Expansion Technique for Retrieving Opinions from Blogs

Improving opinionated blog retrieval effectiveness with quality measures and temporal features

Exploiting External Collections for Query Expansion

Similar Documents

Exploiting external collections for query expansion (2012)

A generative blog post retrieval model that uses query expansion based on external collections.

A Generative Blog Post Retrieval Model that Uses Query Expansion based on External Collections

External Query Expansion in the Blogosphere

A generative blog post retrieval model that uses query expansion based on external collections (2009)

Exploiting query logs for cross-lingual query suggestions

Document representation and query expansion models for blog recommendation.

A Combined Query Expansion Technique for Retrieving Opinions from Blogs

Improving opinionated blog retrieval effectiveness with quality measures and temporal features

Exploiting External Collections for Query Expansion