about
Using Posters to Recommend Anime and Mangas in a Cold-Start Scenario (arxiv.org)
65 points by adulau on Sep 8, 2017 | hide | past | pdf | 15 comments on HN

In plain words: When a manga has few ratings, the system reads its poster for clues like swords or ponytails and mixes those tags with community ratings to guess who will like it. On real data, it recommends obscure titles better than ratings alone and explains user tastes.

Abstract

Item cold-start is a classical issue in recommender systems that affects anime and manga recommendations as well. This problem can be framed as follows: how to predict whether a user will like a manga that received few ratings from the community? Content-based techniques can alleviate this issue but require extra information, that is usually expensive to gather. In this paper, we use a deep learning technique, Illustration2Vec, to easily extract tag information from the manga and anime posters (e.g., sword, or ponytail). We propose BALSE (Blended Alternate Least Squares with Explanation), a new model for collaborative filtering, that benefits from this extra information to recommend mangas. We show, using real data from an online manga recommender system called Mangaki, that our model improves substantially the quality of recommendations, especially for less-known manga, and is able to provide an interpretation of the taste of the users.

Jill-Jênn Vie, Florian Yger, Ryan Lahfa, Basile Clement, Kévin Cocchi, Thomas Chalumeau, Hisashi Kashima
arXiv:1709.01584 · cs.IR, cs.LG, stat.ML · submitted Sep 3, 2017 · updated Sep 7, 2017
abstract · pdf · html · 6 pages, 3 figures, 1 table, accepted at the MANPU 2017 workshop, co-located with ICDAR 2017 in Kyoto on November 10, 2017

add comment on HN

> We propose BALSE (Blended Alternate Least Squares with Explanation), a new model for collaborative filtering

http://knowyourmeme.com/memes/events/balse

(for anyone who didn't spot the joke :) )

Essentially, this is about creating alternative training data to "seed" a recommender engine - if the engine has no previous data about a video series, what other data (cover art) can be used to recommend a given viewer to watch or not watch?

Is a recommender a type of Machine Learning?

yes, recommendation systems are a very vibrant branch of machine learning.
Co-author here. We will provide our dataset for the sake of reproducibility, in the meantime we strongly encourage you to compete to our ongoing Mangaki Data Challenge, organized with Kyoto University: http://research.mangaki.fr/2017/07/18/mangaki-data-challenge... Deadline October 1.
Co-author here, feel free to ask any question!
Can someone eli5?
(skimmed paper)

The classic recommendation problem is the following: given a user and the items (mangas) that they like, out of some universe of items (mangas), how can we recommend new items (mangas) that they are also likely to enjoy? Typically this is done via collaborative filtering or some other method, i.e. people who like the same mangas as the original user also enjoy other mangas, so we recommend these to the original user.

A very common problem occurs when you have a new or obscure manga, aka the cold start problem. There are no reviews to use when finding similar mangas (which is our input to the recommendation system). The authors propose extracting visual information from the posters of these less commonly reviewed mangas that finds characteristics of the manga. The theory is that users that like mangas with 'girl with sword' will also like other mangas that have 'girl with sword' or perhaps 'girl with bow' but probably not 'girl with book' (I made these tags up).

I'm not sure this would work in practice. The story is what matters, not the visuals. Valvrave is good, Code Geass is amazing, but Girls Und Panzer sort of sucks. Yet it looks visually similar.
Huh? Those series look completely different?

Code Geass has sharp lines and a neon palette similar to many modern mecha anime, whereas Girls und Panzer has softer artwork and a more muted pastoral palette, similar to many highschool sports anime. And while neither series fits perfectly into those categories, I'd still expect it to be a good predictor of preferences.

Indeed, while there are some outliers, you can get a pretty good idea what an anime is like from its poster image, which is not really surprising because they're made to appeal to their target audience.

Gotta say I've never met anyone with these anime opinions before.
It's pretty non controversial to say that Code Geass is amazing. It's #16 in the top 100 of MyAnimeList (anime equivalent of IMDb Top 250). As for Valvrave and Girls Und Panzer they're rated 7.3 and 7.6 so it's perfectly possible that someone might like one and not the other.

[1] - https://myanimelist.net/topanime.php

Thanks!
I don't understand the paper fully but the general idea seems to be, if an item like an anime don't have enough ratings from users for standard techniques to apply in analyzing it for recommendations, using metadata derived from its poster (e.g. a girl holding a sword) can help.
Most content-recommendation systems you interact with are collaborative; with data about what users like, they use "likes" from one user as "recommendations" for another user, when those users generally like the same things. So if I like content A, B, C, and you like content A, B, D, they might recommend C to you and D to me.

This approach is (comparatively) easy to build, since your recommendation system never needs to actually look at the content directly, just at whether users like it or not.

There are problems with this approach, though, and a big one is the "cold start" problem–what if you have content in your system that few (or no) users have interacted with? Even though you might have information about the content available to you (like a text description, or, in this case, a poster/cover image), a collaborative system has no way to read this information, and so it can't recommend the content to users until people have found it organically first.

So, a cooler option is to build a system that actually understands the content it recommends. In this case, there was an existing project called Illustration2Vec that used a deep neural net to predict tags from anime images. These authors have built a hybrid recommendation system that uses collaborative filtering when there is a lot of data available, and then uses tag similarity (getting the tags by running Illustration2Vec on the posters/cover images) when data is sparser. It can also "explain" the results by telling you which tags it used to give you the recommendation.

The authors are using data from (and presumably, incorporating the resulting system into) a French anime/manga site called Mangaki (https://mangaki.fr/about/en).

Thanks for this great summing up! I will add that the whole Mangaki platform is on GitHub: https://github.com/mangaki/mangaki/