about
GraphLab: A New Framework for Parallel Machine Learning (arxiv.org)
31 points by bsaunder on Jun 29, 2010 | hide | past | pdf | 4 comments on HN

In plain words: GraphLab lays out machine-learning work as a graph of data and update steps, so many pieces run at once while shared data stays consistent. Unlike MapReduce-style tools, it compactly expressed five algorithms and ran fast on large real-world problems.

Abstract

Designing and implementing efficient, provably correct parallel machine learning (ML) algorithms is challenging. Existing high-level parallel abstractions like MapReduce are insufficiently expressive while low-level tools like MPI and Pthreads leave ML experts repeatedly solving the same design challenges. By targeting common patterns in ML, we developed GraphLab, which improves upon abstractions like MapReduce by compactly expressing asynchronous iterative algorithms with sparse computational dependencies while ensuring data consistency and achieving a high degree of parallel performance. We demonstrate the expressiveness of the GraphLab framework by designing and implementing parallel versions of belief propagation, Gibbs sampling, Co-EM, Lasso and Compressed Sensing. We show that using GraphLab we can achieve excellent parallel performance on large scale real-world problems.

Yucheng Low, Joseph Gonzalez, Aapo Kyrola, Danny Bickson, Carlos Guestrin, Joseph M. Hellerstein
arXiv:1006.4990 · cs.LG, cs.DC · submitted Jun 25, 2010
abstract · pdf · html

add comment on HN

Thanks for posting this, its very relevant to what I am currently working on.
You're welcome. I'm working on something similar as well. I haven't fully digested the paper yet, but it resonates very well with my current implementation.
Maybe we should get a beer together and discuss implementations :)
Sure, or at least chat on IRC some evening.