In plain words: A startup-success predictor feeds a company's past numbers into a time-series forecaster, then pulls in facts about its rivals and partners from a knowledge graph so the model sees how firms connect. It beat the usual time-series-only approach at guessing which startups succeed.
Abstract · Enhancing Startup Success Predictions in Venture Capital: A GraphRAG Augmented Multivariate Time Series Method
In the Venture Capital (VC) industry, predicting the success of startups is challenging due to limited financial data and the need for subjective revenue forecasts. Previous methods based on time series analysis often fall short as they fail to incorporate crucial inter-company relationships such as competition and collaboration. To fill the gap, this paper aims to introduce a novel approach using GraphRAG augmented time series model. With GraphRAG, time series predictive methods are enhanced by integrating these vital relationships into the analysis framework, allowing for a more dynamic understanding of the startup ecosystem in venture capital. Our experimental results demonstrate that our model significantly outperforms previous models in startup success predictions.
Zitian Gao, Yihao Xiao
arXiv:2408.09420 · q-fin.CP, cs.CL, cs.LG · submitted Aug 18, 2024 · updated Mar 22, 2025
abstract · pdf · html · ICLR 2025 Financial AI
Our approach consists of two stages: First, we input a large amount of unstructured news text into GraphRAG, where during the clustering and retrieval process, global statistics are dynamically updated to accurately cluster relationships between entities. This results in a directed knowledge graph of companies, with edges representing both direct and indirect relationships between companies. In the second stage, inspired by Ibrahim et al. (2022) and Barigozzi & Brownlees (2019), we transform the knowledge graph into a mask matrix using the Leiden algorithm Traag et al. (2019). This matrix is then used as a regularizer in the multivariate time series model, combined with relevant covariates. This method effectively enhances the generalization ability of models on scarce company data and significantly improves prediction performance, as illustrated in Figure 1. The specific algorithm details are provided in Algorithms 1 and 2.