about
Mind the Gaps: Mixture-of-Minds for Human Simulation (arxiv.org)
2 points by jtewright 54 days ago | hide | past | pdf | 1 comment on HN

In plain words: A new simulator groups respondents into clusters and gives each its own trained brain, so it can guess how one person—not just the crowd—would answer. On an outside survey it scored 0.775 for matching individual answers, the best result, with slight bias.

Abstract

Predicting how a population will answer a new question is a long-standing goal. Statistical methods succeed at the level of the mass but falter at the level of the individual. Large language model simulators inherit this gap. They recover a population's central tendencies while flattening its heterogeneity, and they carry social biases and prompt brittleness that distort individual predictions. This paper introduces Anacreon, an audience simulation model that targets the individual level within a narrow, well-specified domain. Anacreon learns an authorship embedding that separates individuals, clusters a real qualitative corpus around seed people, and trains a dedicated adapter for each cluster, a mixture of minds, on a Gemma~4 12B base. It harvests demographics, psychological traits, and survey responses from public text, and augments each record with a chain-of-emotion. It reduces prompt brittleness by shuffling response options and reduces positive bias by balancing the training distribution. On a large, externally sourced survey, Anacreon reaches a state-of-the-art ordinal alignment of 0.775, the individual-level accuracy measure on which the field has converged, with a small residual bias. The work is a step toward drawing aggregate insight from faithfully simulated individuals.

Pranav Dahiya
arXiv:2608.06115 · cs.AI · submitted Aug 6, 2026
abstract · pdf · html

add comment on HN

Semilattice founder here. pranavdahiya did the work and wrote the paper, also here for questions.

Short version: instead of prompting one large model to simulate different people, we train one small model per cluster of similar people. In the paper that's 420 LoRA adapters on Gemma 4 12B.

The motivation is that LLM simulators are good at predicting the average but bad at individuals. They flatten the heterogeneity, minority views, and disagreement that make a sample real, and answers change with reworded questions. Making LLMs better assistants makes them worse human simulators.

Privacy: there are no models of specific real people. The pipeline clusters a real corpus into groups, each of which represents a probable person rather than a named individual, based on how they respond to stimuli within a specific domain. What gets trained is simulacrum 283, not a model of someone real.

Evaluation: we split train and test 80:20 on a date cutoff rather than randomly, so each model learns to predict the answers to future questions based on what it saw in the past. We also randomise answer option ordering and flip the sentiment of questions to make sure the models learn underlying predictors rather than meaningless signals like answer position or sentiment patterns.

Results: we score top-1 accuracy, which is simple exact match accuracy, and ordinal alignment, the individual-level metric the field has converged around, which measures the accuracy of ordered Likert scale questions. On top-1, we score 67.9%, and on ordinal alignment, 0.775. For scale, ask real people the same questions twice and they only match their own earlier answers about 80% of the time, so that is the target rather than 100%.

The paper compares these against prior published methods on the same metrics, and we come out ahead of all of them by a few points. However it's on our own population rather than a shared benchmark, so it's not apples to apples.

Happy to get into any of it.