NDLI: Batch Mode Active Sampling Based on Marginal Probability Distribution Matching

Please wait, while we are loading the content...

Batch Mode Active Sampling Based on Marginal Probability Distribution Matching

Content Provider	ACM Digital Library
Author	Chattopadhyay, Rita Wang, Zheng Panchanathan, Sethuraman Davidson, Ian Ye, Jieping Fan, Wei
Copyright Year	2013
Abstract	Active Learning is a machine learning and data mining technique that selects the most informative samples for labeling and uses them as training data; it is especially useful when there are large amount of unlabeled data and labeling them is expensive. Recently, batch-mode active learning, where a set of samples are selected concurrently for labeling, based on their collective merit, has attracted a lot of attention. The objective of batch-mode active learning is to select a set of informative samples so that a classifier learned on these samples has good generalization performance on the unlabeled data. Most of the existing batch-mode active learning methodologies try to achieve this by selecting samples based on certain criteria. In this article we propose a novel criterion which achieves good generalization performance of a classifier by specifically selecting a set of query samples that minimize the difference in distribution between the labeled and the unlabeled data, after annotation. We explicitly measure this difference based on all candidate subsets of the unlabeled data and select the best subset. The proposed objective is an NP-hard integer programming optimization problem. We provide two optimization techniques to solve this problem. In the first one, the problem is transformed into a convex quadratic programming problem and in the second method the problem is transformed into a linear programming problem. Our empirical studies using publicly available UCI datasets and two biomedical image databases demonstrate the effectiveness of the proposed approach in comparison with the state-of-the-art batch-mode active learning methods. We also present two extensions of the proposed approach, which incorporate uncertainty of the predicted labels of the unlabeled data and transfer learning in the proposed formulation. In addition, we present a joint optimization framework for performing both transfer and active learning simultaneously unlike the existing approaches of learning in two separate stages, that is, typically, transfer learning followed by active learning. We specifically minimize a common objective of reducing distribution difference between the domain adapted source, the queried and labeled samples and the rest of the unlabeled target domain data. Our empirical studies on two biomedical image databases and on a publicly available 20 Newsgroups dataset show that incorporation of uncertainty information and transfer learning further improves the performance of the proposed active learning based classifier. Our empirical studies also show that the proposed transfer-active method based on the joint optimization framework performs significantly better than a framework which implements transfer and active learning in two separate stages.
Starting Page	1
Ending Page	25
Page Count	25
File Format	PDF
ISSN	15564681
e-ISSN	1556472X
DOI	10.1145/2513092.2513094
Volume Number	7
Issue Number	3
Journal	ACM Transactions on Knowledge Discovery from Data (TKDD)
Language	English
Publisher	Association for Computing Machinery (ACM)
Publisher Date	2013-09-01
Publisher Place	New York
Access Restriction	One Nation One Subscription (ONOS)
Subject Keyword	Active learning Marginal probability distribution Maximum mean discrepancy Transfer learning
Content Type	Text
Resource Type	Article
Subject	Computer Science

Sl.	Authority	Responsibilities	Communication Details
1	Ministry of Education (GoI), Department of Higher Education	Sanctioning Authority	https://www.education.gov.in/ict-initiatives
2	Indian Institute of Technology Kharagpur	Host Institute of the Project: The host institute of the project is responsible for providing infrastructure support and hosting the project	https://www.iitkgp.ac.in
3	National Digital Library of India Office, Indian Institute of Technology Kharagpur	The administrative and infrastructural headquarters of the project	Dr. B. Sutradhar bsutra@ndl.gov.in
4	Project PI / Joint PI	Principal Investigator and Joint Principal Investigators of the project	Dr. B. Sutradhar bsutra@ndl.gov.in Prof. Saswat Chakrabarti will be added soon
5	Website/Portal (Helpdesk)	Queries regarding NDLI and its services	support@ndl.gov.in
6	Contents and Copyright Issues	Queries related to content curation and copyright issues	content@ndl.gov.in
7	National Digital Library of India Club (NDLI Club)	Queries related to NDLI Club formation, support, user awareness program, seminar/symposium, collaboration, social media, promotion, and outreach	clubsupport@ndl.gov.in
8	Digital Preservation Centre (DPC)	Assistance with digitizing and archiving copyright-free printed books	dpc@ndl.gov.in
9	IDR Setup or Support	Queries related to establishment and support of Institutional Digital Repository (IDR) and IDR workshops	idr@ndl.gov.in

Batch mode active sampling based on marginal probability distribution matching

Batch mode active sampling based on marginal probability distribution matching.

Batch Mode Active Sampling based on Marginal Probability Distribution Matching

Querying Discriminative and Representative Samples for Batch Mode Active Learning

Multi-label active learning by model guided distribution matching

Batch Mode Active Learning for Object Detection Based on Maximum Mean Discrepancy

Querying discriminative and representative samples for batch mode active learning

Batch Mode Active Learning for Networked Data

Semisupervised SVM batch mode active learning with applications to image retrieval

Batch Mode Active Sampling Based on Marginal Probability Distribution Matching

Similar Documents

Batch mode active sampling based on marginal probability distribution matching

Batch mode active sampling based on marginal probability distribution matching.

Batch Mode Active Sampling based on Marginal Probability Distribution Matching

Querying Discriminative and Representative Samples for Batch Mode Active Learning

Multi-label active learning by model guided distribution matching

Batch Mode Active Learning for Object Detection Based on Maximum Mean Discrepancy

Querying discriminative and representative samples for batch mode active learning

Batch Mode Active Learning for Networked Data

Semisupervised SVM batch mode active learning with applications to image retrieval

Batch Mode Active Sampling Based on Marginal Probability Distribution Matching