
AI Summary
→ WHAT IT COVERS Joon Sung Park, creator of the 2023 Generative Agents (Smallville) paper and cofounder of Simile AI, explains how behavioral simulation differs from prediction, why frontier LLMs fail at modeling real human behavior, and how Simile's approach of collecting RCT data and post-training models achieves 85% accuracy in replicating individual human responses. → KEY INSIGHTS - **Behavioral vs. Attitudinal Data:** Frontier LLMs train on web data—what people say they do—not what they actually do. Simile identifies three data categories: qualitative interview data (life stories), observational behavioral data (transactions, web scraping), and causal RCT data. The RCT category is the hardest to acquire but most valuable, because it captures the *why* behind decisions, enabling counterfactual simulation rather than mere prediction. - **85% Individual Replication Accuracy:** Simile's "Generative Agent Simulations of 1,000 People" paper recruited a representative US sample, collected two hours of data per person, built digital twins, then had twins complete Big Five personality tests, General Social Survey items, and behavioral economics games. Twins matched source individuals' responses 85% of the time—comparable to how accurately people replicate their own prior answers across sessions. - **Simulation vs. Prediction Framing:** Decision-makers rarely need to know an outcome will be bad—they need the intervention path to change it. Simulation surfaces counterintuitive step-by-step action sequences toward a goal, analogous to Asimov's psychohistory in Foundation, where the optimal first move (exiling scientists to Terminus) appears irrational in isolation but is correct within the full causal chain. - **When to Prompt vs. When to Retrain:** Use prompting when the model already contains the relevant "social physics" and only needs to react to a new environment. Retrain or post-train when the model lacks the underlying causal mechanics of a population's behavior. Simile trains two distinct model types: a population-level model and an individual-level model, both taking a population description plus a stimulus as input. - **Synthetic Panels Replace Human Panels at Scale:** Simile collects data on tens of thousands of people weekly and maintains panel partnerships reaching tens of millions globally. Once a person's model is built, it is domain-agnostic and reusable across studies. Stable traits like risk tolerance don't change over time, making the initial data collection a durable asset deployable across concept testing, focus groups, AB testing, and earnings-call modeling. - **Scaling Laws Emerging for Simulation:** Simile observes early evidence that increasing human behavioral data and compute produces predictable, measurable gains in simulation accuracy—mirroring the scaling laws seen in LLM pretraining. The long-term architecture goal is a multi-agent simulation of 8 billion people, with costs potentially reaching foundation-model training scale, justified by societal decisions around climate coordination, democratic stability, and policy design like UBI. → NOTABLE MOMENT Park argues that the billion-persona paper from Tencent—which generates synthetic populations by cross-multiplying demographic and personality variables—only retrieves statistics already embedded in model weights. When populations are niche or behaviors are causally complex, frontier model accuracy drops to 30%, making those outputs unreliable for enterprise decisions. 💼 SPONSORS None detected 🏷️ AI Simulation, Behavioral Modeling, Synthetic Panels, Agent-Based Modeling, Human Digital Twins, Foundation Models
