In plain words: A decision tree where a language model picks each split, so no hand-built features are needed, and people can step in to fix wrong paths. It found future unicorn startups with 7.8% precision, beating few-shot GPT-4o and the best humans at 3.1% to 5.6%.
Abstract · GPTree: Towards Explainable Decision-Making via LLM-powered Decision Trees
Traditional decision tree algorithms are explainable but struggle with non-linear, high-dimensional data, limiting its applicability in complex decision-making. Neural networks excel at capturing complex patterns but sacrifice explainability in the process. In this work, we present GPTree, a novel framework combining explainability of decision trees with the advanced reasoning capabilities of LLMs. GPTree eliminates the need for feature engineering and prompt chaining, requiring only a task-specific prompt and leveraging a tree-based structure to dynamically split samples. We also introduce an expert-in-the-loop feedback mechanism to further enhance performance by enabling human intervention to refine and rebuild decision paths, emphasizing the harmony between human expertise and machine intelligence. Our decision tree achieved a 7.8% precision rate for identifying "unicorn" startups at the inception stage of a startup, surpassing gpt-4o with few-shot learning as well as the best human decision-makers (3.1% to 5.6%).
Sichao Xiong, Yigit Ihlamur, Fuat Alican, Aaron Ontoyin Yin
arXiv:2411.08257 · cs.LG, cs.AI, cs.CE · submitted Nov 13, 2024
abstract · pdf · html
GPTree outperforms the randomness by ~10x and the world's best experts by ~3x.
The cool part is that:
1. It's explainable, not a blackbox. A decision tree that humans can understand.
2. Human-machine harmony wins over machine-only and human-only decision making processes.
3. It's potentially applicable and high performant to any decision making use-case.
We fine-tuned the model for our own use-case at Vela Partners, picking outlier startups at their inception stage.
The reason why we love this research problem is that humans are so bad at picking startups at their inception stage.
For context, only 2% of the US-based investor-backed startups become an outlier return at the inception stage. Y Combinator and tier-1 VCs hovers around ~3% and ~6%, respectively. Ten-fold cross-validated GPTree is at 8% and most fined-tuned version is at ~18%.
Please take a moment to take a deep breath and let that sync in...
GPTree can find 1 outlier startup out of 5 of its investments at the inception stage. This may translate into a 10x+ return fund for whoever uses it if the future behaves as it forecasts.
Excited to hear the HN's feedback.