Multimodal Learning in Loosely-Organized Web Images

Kun Duan, David J. Crandall, Dhruv Batra
2014 2014 IEEE Conference on Computer Vision and Pattern Recognition  
Photo-sharing websites have become very popular in the last few years, leading to huge collections of online images. In addition to image data, these websites collect a variety of multimodal metadata about photos including text tags, captions, GPS coordinates, camera metadata, user profiles, etc. However, this metadata is not well constrained and is often noisy, sparse, or missing altogether. In this paper, we propose a framework to model these "loosely organized" multimodal datasets, and show
more » ... ow to perform looselysupervised learning using a novel latent Conditional Random Field framework. We learn parameters of the LCRF automatically from a small set of validation data, using Information Theoretic Metric Learning (ITML) to learn distance functions and a structural SVM formulation to learn the potential functions. We apply our framework on four datasets of images from Flickr, evaluating both qualitatively and quantitatively against several baselines.
doi:10.1109/cvpr.2014.316 dblp:conf/cvpr/DuanCB14 fatcat:ddg2rdsvqjdzjcjf7qsr4prhui