Skip to main content
Latent Space

Simulation: the new Scaling Law — Joon Sung Park, Simile AI

69 min episode · 3 min read
·
Joon Sung Park

Episode

69 min

Read time

3 min

Topics

Productivity, Relationships, Startups

AI-Generated Summary

Key Takeaways

  • Behavioral vs. Attitudinal Data: Frontier LLMs train on web data—what people say they do—not what they actually do. Simile identifies three data categories: qualitative interview data (life stories), observational behavioral data (transactions, web scraping), and causal RCT data. The RCT category is the hardest to acquire but most valuable, because it captures the *why* behind decisions, enabling counterfactual simulation rather than mere prediction.
  • 85% Individual Replication Accuracy: Simile's "Generative Agent Simulations of 1,000 People" paper recruited a representative US sample, collected two hours of data per person, built digital twins, then had twins complete Big Five personality tests, General Social Survey items, and behavioral economics games. Twins matched source individuals' responses 85% of the time—comparable to how accurately people replicate their own prior answers across sessions.
  • Simulation vs. Prediction Framing: Decision-makers rarely need to know an outcome will be bad—they need the intervention path to change it. Simulation surfaces counterintuitive step-by-step action sequences toward a goal, analogous to Asimov's psychohistory in Foundation, where the optimal first move (exiling scientists to Terminus) appears irrational in isolation but is correct within the full causal chain.
  • When to Prompt vs. When to Retrain: Use prompting when the model already contains the relevant "social physics" and only needs to react to a new environment. Retrain or post-train when the model lacks the underlying causal mechanics of a population's behavior. Simile trains two distinct model types: a population-level model and an individual-level model, both taking a population description plus a stimulus as input.
  • Synthetic Panels Replace Human Panels at Scale: Simile collects data on tens of thousands of people weekly and maintains panel partnerships reaching tens of millions globally. Once a person's model is built, it is domain-agnostic and reusable across studies. Stable traits like risk tolerance don't change over time, making the initial data collection a durable asset deployable across concept testing, focus groups, AB testing, and earnings-call modeling.

What It Covers

Joon Sung Park, creator of the 2023 Generative Agents (Smallville) paper and cofounder of Simile AI, explains how behavioral simulation differs from prediction, why frontier LLMs fail at modeling real human behavior, and how Simile's approach of collecting RCT data and post-training models achieves 85% accuracy in replicating individual human responses.

Key Questions Answered

  • Behavioral vs. Attitudinal Data: Frontier LLMs train on web data—what people say they do—not what they actually do. Simile identifies three data categories: qualitative interview data (life stories), observational behavioral data (transactions, web scraping), and causal RCT data. The RCT category is the hardest to acquire but most valuable, because it captures the *why* behind decisions, enabling counterfactual simulation rather than mere prediction.
  • 85% Individual Replication Accuracy: Simile's "Generative Agent Simulations of 1,000 People" paper recruited a representative US sample, collected two hours of data per person, built digital twins, then had twins complete Big Five personality tests, General Social Survey items, and behavioral economics games. Twins matched source individuals' responses 85% of the time—comparable to how accurately people replicate their own prior answers across sessions.
  • Simulation vs. Prediction Framing: Decision-makers rarely need to know an outcome will be bad—they need the intervention path to change it. Simulation surfaces counterintuitive step-by-step action sequences toward a goal, analogous to Asimov's psychohistory in Foundation, where the optimal first move (exiling scientists to Terminus) appears irrational in isolation but is correct within the full causal chain.
  • When to Prompt vs. When to Retrain: Use prompting when the model already contains the relevant "social physics" and only needs to react to a new environment. Retrain or post-train when the model lacks the underlying causal mechanics of a population's behavior. Simile trains two distinct model types: a population-level model and an individual-level model, both taking a population description plus a stimulus as input.
  • Synthetic Panels Replace Human Panels at Scale: Simile collects data on tens of thousands of people weekly and maintains panel partnerships reaching tens of millions globally. Once a person's model is built, it is domain-agnostic and reusable across studies. Stable traits like risk tolerance don't change over time, making the initial data collection a durable asset deployable across concept testing, focus groups, AB testing, and earnings-call modeling.
  • Scaling Laws Emerging for Simulation: Simile observes early evidence that increasing human behavioral data and compute produces predictable, measurable gains in simulation accuracy—mirroring the scaling laws seen in LLM pretraining. The long-term architecture goal is a multi-agent simulation of 8 billion people, with costs potentially reaching foundation-model training scale, justified by societal decisions around climate coordination, democratic stability, and policy design like UBI.

Notable Moment

Park argues that the billion-persona paper from Tencent—which generates synthetic populations by cross-multiplying demographic and personality variables—only retrieves statistics already embedded in model weights. When populations are niche or behaviors are causally complex, frontier model accuracy drops to 30%, making those outputs unreliable for enterprise decisions.

Know someone who'd find this useful?

Episode Transcript

Today, we have June in the podcast. Excited to kick this one off. Very exciting company. I wanna kick off and ask you the question. You know, talk us through the story of your life. How have you gotten here? Yeah. For sure. So really excited to be here. A story of of my life. So I was born in Korea, and I lived there for good eleven years or so of my life. And then my family moved to Boston. So we moved, when I was 11. And my parents were doctors, so they were basically going through their postdoctoral studies. My dad was a surgeon, so he was doing, his sabbatical years actually at the Boston Children's Hospital. So I grew up there, not too close to tech, actually. I was very much, like, you know, music, articy, painting, like, that kind of guy. I I actually got into painting a little bit later, in high school, but that's what I used to do. And then I grew up mostly in the East Coast after Korea. So I lived good number of years in New Hampshire, and then I went to college in Pennsylvania. And I got into more of this tech scene, in college. So I was originally trained to be an artist. I actually thought that would be my actually professional career. So I it wasn't a hobby. It was actually like, hey. Let's make a living out of this. And then gradually, I got really interested in this idea of, hey. The greatest artist often creates their own medium. And the best medium that we had available today was actually in computation. So I decided to go deeper into that, and one thing's, led to another. And, obviously, we can go deeper into this, but I decided that research was something that gradually, that I still got in got interested in, and here I am. So there's obviously a lot that you packed into the research components. You had one of the best papers of 2023, which was the generative agents paper, commonly known as the Smallville paper. Yeah. Feel free to call back to anything else that you mentioned. But most people would have heard of you from this, obviously. Do you have any statistics of how many people, have, like, read it? Archive gives you something. Right? Some some stats. Yeah. It's a good question. How many people have read it? I'm actually not sure. I know that's that I mean, we do keep track of the number of citations, which I know is going up, quite fast. But the readers go to Google Google Scholar has seven seventy two thousand It made a bigger hit, and it was actually a pretty instrumental paper. It was, like, one that got cited so many times. Frequently, like, when people ask what is the best paper of the year, like, best paper you've read recently, it's it's this one. I thought the the memory component …

Get the full transcript (12,760 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all Latent Space transcripts →

You just read a 3-minute summary of a 66-minute episode.

Get Latent Space summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

Books, tools, and gear mentioned in this episode

SignalCast may earn commission on purchases via these links. As an Amazon Associate, SignalCast earns from qualifying purchases.

Books

  • by Isaac Asimov

    Simulation surfaces counterintuitive step-by-step action sequences toward a goal, analogous to Asimov's psychohistory in Foundation, where the optimal first move (exiling scientists to Terminus) appears irrational in isolation but is correct within the full causal chain.

other

  • Joon Sung Park, creator of the 2023 Generative Agents (Smallville) paper and cofounder of Simile AI, explains how behavioral simulation differs from prediction

company

  • Simile AIBy guest
    Joon Sung Park, creator of the 2023 Generative Agents (Smallville) paper and cofounder of Simile AI, explains how behavioral simulation differs from prediction

More from Latent Space

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into Latent Space.

Every Monday, we deliver AI summaries of the latest episodes from Latent Space and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime