In plain words: Epicure cleans 4.14 million recipes to 1,790 ingredient names, then learns an ingredient map from graphs of which ingredients appear together and which share flavor chemicals. Three versions differ in which graph paths they follow, landing at different points between chemistry and recipe context.
Abstract · Epicure: Navigating the Emergent Geometry of Food Ingredient Embeddings
We present Epicure, a family of three sibling skip-gram ingredient embeddings retrained from scratch on a multilingual recipe corpus. We aggregate 4.14M recipes from 11 sources spanning seven languages, English, Chinese, Russian, Vietnamese, Spanish, Turkish, Indonesian, German, and Indian-English, and normalise the raw ingredient strings to 1,790 canonical entries via an LLM-augmented pipeline. A 203,508-edge ingredient-ingredient NPMI graph and an 80,019-edge typed FlavorDB ingredient-compound graph, 2,247 typed compound nodes across 15 categories, seed three Metapath2Vec variants that share architecture and hyperparameters and differ only in the random-walk schema: Cooc walks the co-occurrence graph only, Chem walks the typed compound metapaths only, and Core blends both via injected ingredient-ingredient walks at controlled mixing, placing each model at a distinct point on the chemistry-vs-recipe-context spectrum.
Jakub Radzikowski, Josef Chen
arXiv:2605.22391 · cs.AI, cs.CL, cs.CY · submitted May 21, 2026
abstract · pdf · html
A better title would be: "all of human ingredients compressed into 1,800 primitives"
There is little to substantively nothing about the actual cooking: preparation methods, proportions, etc.
But the idea that tomato goes well with beef the whole world over is very interesting and useful for creating flavors that will go together, perhaps surprisingly. It will be a nice resource in the future.