about
Knowledge-Based Trust: Estimating the Trustworthiness of Web Sources [pdf] (arxiv.org)
10 points by ScottBurson on Mar 6, 2015 | hide | past | pdf | discuss on HN

In plain words: Instead of ranking sites by how many sites link to them, this scores a site by how many of its facts are true, separating extraction mistakes from real errors. It recovered true trust levels on synthetic data and passed manual checks across 119 million pages.

Abstract · Knowledge-Based Trust: Estimating the Trustworthiness of Web Sources

The quality of web sources has been traditionally evaluated using exogenous signals such as the hyperlink structure of the graph. We propose a new approach that relies on endogenous signals, namely, the correctness of factual information provided by the source. A source that has few false facts is considered to be trustworthy. The facts are automatically extracted from each source by information extraction methods commonly used to construct knowledge bases. We propose a way to distinguish errors made in the extraction process from factual errors in the web source per se, by using joint inference in a novel multi-layer probabilistic model. We call the trustworthiness score we computed Knowledge-Based Trust (KBT). On synthetic data, we show that our method can reliably compute the true trustworthiness levels of the sources. We then apply it to a database of 2.8B facts extracted from the web, and thereby estimate the trustworthiness of 119M webpages. Manual evaluation of a subset of the results confirms the effectiveness of the method.

Xin Luna Dong, Evgeniy Gabrilovich, Kevin Murphy, Van Dang, Wilko Horn, Camillo Lugaresi, Shaohua Sun, Wei Zhang
arXiv:1502.03519 · cs.DB, cs.IR · submitted Feb 12, 2015
abstract · pdf · html

add comment on HN
Also discussed: Nov 2015 (3 points, 0 comments) · Mar 2015 (5 points, 1 comment) · Mar 2015 (1 point, 0 comments) · Mar 2015 (4 points, 0 comments)