In plain words: Instead of smoothing synthetic text into perfect-sounding prose, this system maps the writer's thinking and rebuilds the text with the small quirks real people make. In a 2015 stock-crash test, strategies using this data cut their worst losses by 47.4%.
Abstract · The Necessity of Imperfection:Reversing Model Collapse via Simulating Cognitive Boundedness
Although synthetic data is widely promoted as a remedy, its prevailing production paradigm -- one optimizing for statistical smoothness -- systematically removes the long-tail, cognitively grounded irregularities that characterize human text. Prolonged training on such statistically optimal but cognitively impoverished data accelerates model collapse. This paper proposes a paradigm shift: instead of imitating the surface properties of data, we simulate the cognitive processes that generate human text. We introduce the Prompt-driven Cognitive Computing Framework (PMCSF), whose core consists of a Cognitive State Decoder (CSD) that reverse-engineers unstructured text into structured cognitive vectors, and a Cognitive Text Encoder (CTE) that re-materializes these states into text enriched with human-typical imperfections via mathematically defined Cognitive Perturbation Operators. The framework is validated through a two-stage objective evaluation pipeline. First, in cognitive codec verification, CTE text yields a Jensen-Shannon divergence of 0.0614 from human text (vs. 0.4431 for standard LLM output), passes double-blind professional media review, and achieves an intraclass correlation coefficient ICC > 0.9 for cognitive profile alignment across heterogeneous models. Second, in functional gain evaluation, isomorphic stress tests in the A-share market show that strategies incorporating CTE-generated data reduce maximum drawdown by 47.4% during the 2015 crash and deliver 8.6% Defensive Alpha, exceeding transaction costs by a factor of 33. Our findings demonstrate that modelling human cognitive limitations -- not copying surface data -- enables synthetic data with genuine functional gain, offering a viable technical pathway toward resolving the AI data-collapse crisis.
Zhongjie Jiang
arXiv:2512.01354 · cs.AI, cs.CL, cs.CY, cs.LG, q-fin.TR · submitted Dec 1, 2025 · updated Dec 8, 2025
abstract · pdf · html · 60 pages,9 figures. v3: Major update. Added 3D topological visualization (Figure 1) and independent computational verification of the Adaptive Markets Hypothesis (AMH). Includes comprehensive Supplementary Materials (algorithmic pseudocode, system architecture, and real-time GARCH logs) for technical reproducibility
I'm the author of this paper/project. I am a humanities researcher turned quant architect, working solo.
The Problem: I noticed that LLMs are suffering from "Model Collapse" because they optimize for statistical smoothness. They are too perfect, which makes them dumb and easily detectable.
The Solution: Instead of cleaning data, I went back to Herbert Simon's "Bounded Rationality". I built a pipeline (PMCSF) that injects mathematical "cognitive noise" (e.g., sentence length oscillation, hesitation) back into the generation process.
The Results:
Anti-Detection: It achieves a Jensen-Shannon divergence of 0.0614 against human text (vs 0.44 for standard AI).
Financial Alpha: In a blind backtest of the 2015 Crash, the cognitive signal (MDI) predicted the liquidity freeze, reducing drawdown by 47%.
Safety Note: Because this architecture can effectively generate undetectable disinformation and manipulate sentiment, I have redacted the core prompts and safety constraints in the open release. I believe we need to build the "Radar" (detection) before distributing the "Missile".
Happy to answer questions about the Neuro-Symbolic architecture or the backtest data!