DocumentCode
3458215
Title
Online crowdsourcing: Rating annotators and obtaining cost-effective labels
Author
Welinder, Peter ; Perona, Pietro
Author_Institution
California Inst. of Technol., Pasadena, CA, USA
fYear
2010
fDate
13-18 June 2010
Firstpage
25
Lastpage
32
Abstract
Labeling large datasets has become faster, cheaper, and easier with the advent of crowdsourcing services like Amazon Mechanical Turk. How can one trust the labels obtained from such services? We propose a model of the labeling process which includes label uncertainty, as well a multi-dimensional measure of the annotators´ ability. From the model we derive an online algorithm that estimates the most likely value of the labels and the annotator abilities. It finds and prioritizes experts when requesting labels, and actively excludes unreliable annotators. Based on labels already obtained, it dynamically chooses which images will be labeled next, and how many labels to request in order to achieve a desired level of confidence. Our algorithm is general and can handle binary, multi-valued, and continuous annotations (e.g. bounding boxes). Experiments on a dataset containing more than 50,000 labels show that our algorithm reduces the number of labels required, and thus the total cost of labeling, by a large factor while keeping error rates low on a variety of datasets.
Keywords
Web services; costing; image classification; outsourcing; amazon mechanical turk; cost effective label; dataset labeling service; label uncertainty; multidimensional annotator ability measurement; multivalued continuous annotation; online crowdsourcing; rating annotator; Adaptation model; Computer vision; Costs; Error analysis; Labeling; Noise figure; Outsourcing;
fLanguage
English
Publisher
ieee
Conference_Titel
Computer Vision and Pattern Recognition Workshops (CVPRW), 2010 IEEE Computer Society Conference on
Conference_Location
San Francisco, CA
ISSN
2160-7508
Print_ISBN
978-1-4244-7029-7
Type
conf
DOI
10.1109/CVPRW.2010.5543189
Filename
5543189
Link To Document