Using Visual Features based on MPEG-7 and Deep Learning for Movie Recommendation

Published in International Journal of Multimedia Information Retrieval, 2018

Recommended citation: Yashar Deldjoo, Mehdi Elahi, Massimo Quadrana, Paolo Cremonesi International Journal of Multimedia Information Retrieval 2018.



Item features play an important role in movie recommender systems, where recommendations can be generated by using explicit or implicit preferences of users on traditional features (attributes) such as tag, genre, and cast. Typically, movie features are human- generated, either editorially (e.g., genre and cast) or by leveraging the wisdom of the crowd (e.g., tag), and as such, they are prone to noise and are expensive to collect. Moreover, these features are often rare or absent for new items, making it difficult or even impossible to provide good quality recommendations.

In this paper, we show that users’ preferences on movies can be well or even better described in terms of the mise-en-scene features, i.e., the visual aspects of a movie that characterize design, aesthetics and style (e.g., colors, textures). We use both MPEG-7 visual descriptors and Deep Learning hidden layers as examples of mise-en-scene features that can visually describe movies. These features can be computed automatically from any video file, offering the flexibility in handling new items, avoiding the need for costly and error-prone human-based tagging, and providing good scalability. We have conducted a set of experiments on a large catalog of 4K movies. Results show that recommendations based on mise-en-scene features consistently out- perform traditional metadata attributes (e.g., genre and tag).