Radically Better Reasoning: Elicit's Andreas Stuhlmüller & Jungwon Byun on World Models for Research
Episode
106 min
Read time
3 min
Topics
Remote Work, Startups, Design & UX
AI-Generated Summary
Key Takeaways
- ✓Process Guarantee via DSL: Elicit built a proprietary domain-specific language that compiles reasoning into discrete microservices, ensuring the identical analytical process applies to document number 5 and document 9,999 in a batch. When they tested Claude, ChatGPT, and Elicit on analyzing 100 toxicology papers, only Elicit could verify all 100 were actually processed — the other models admitted mid-conversation they had not completed the task.
- ✓LLM Probability Instability: Current frontier models produce unreliable confidence estimates because they lack coherent internal world models backing their stated probabilities. When asked to estimate clinical trial failure rates, models shift their percentage significantly if you simply mention base rates or add contextual framing — behavior a domain expert would resist. Elicit addresses this through structured scaffolding rather than relying on raw model verbalization of uncertainty.
- ✓World Models as External Continual Learning: Rather than storing evolving knowledge in model weights, Elicit is building structured external representations — combining graph-based causal diagrams, SQL tables, and heterogeneous knowledge formats — that models can update incrementally as new papers arrive. This approach makes a model's understanding of complex evidence bodies inspectable by humans and other AIs, enabling consistent counterfactual and intervention-based reasoning across thousands of data points.
- ✓Evidence Quality Beyond Metadata: Elicit evaluates research quality from content and methodology rather than relying solely on citation counts or journal impact factor — proxies that miss landmark papers like foundational CRISPR work published in lower-tier journals. Researchers can specify domain-appropriate quality thresholds, such as minimum sample sizes or study designs, and Elicit applies those criteria uniformly across all retrieved sources rather than defaulting to surface-level heuristics.
- ✓Automated Engineering via "The Line": Elicit's internal software pipeline called The Line automates the full engineering cycle — from Slack feature request through spec writing, implementation, video testing, code review, and production deployment — currently merging 30 to 50 issues per week without human intervention on simple tasks. The system self-identifies when human escalation is needed, such as incomplete specs or high-complexity changes, and routes accordingly.
What It Covers
Elicit cofounders Andreas Stuhlmüller and Jungwon Byun explain how their AI research platform serves seven of the top 20 life sciences companies by combining frontier reasoning models with a custom domain-specific language that guarantees systematic process execution at scale, and why externalized "world models" represent the next frontier for reliable causal and counterfactual scientific analysis.
Key Questions Answered
- •Process Guarantee via DSL: Elicit built a proprietary domain-specific language that compiles reasoning into discrete microservices, ensuring the identical analytical process applies to document number 5 and document 9,999 in a batch. When they tested Claude, ChatGPT, and Elicit on analyzing 100 toxicology papers, only Elicit could verify all 100 were actually processed — the other models admitted mid-conversation they had not completed the task.
- •LLM Probability Instability: Current frontier models produce unreliable confidence estimates because they lack coherent internal world models backing their stated probabilities. When asked to estimate clinical trial failure rates, models shift their percentage significantly if you simply mention base rates or add contextual framing — behavior a domain expert would resist. Elicit addresses this through structured scaffolding rather than relying on raw model verbalization of uncertainty.
- •World Models as External Continual Learning: Rather than storing evolving knowledge in model weights, Elicit is building structured external representations — combining graph-based causal diagrams, SQL tables, and heterogeneous knowledge formats — that models can update incrementally as new papers arrive. This approach makes a model's understanding of complex evidence bodies inspectable by humans and other AIs, enabling consistent counterfactual and intervention-based reasoning across thousands of data points.
- •Evidence Quality Beyond Metadata: Elicit evaluates research quality from content and methodology rather than relying solely on citation counts or journal impact factor — proxies that miss landmark papers like foundational CRISPR work published in lower-tier journals. Researchers can specify domain-appropriate quality thresholds, such as minimum sample sizes or study designs, and Elicit applies those criteria uniformly across all retrieved sources rather than defaulting to surface-level heuristics.
- •Automated Engineering via "The Line": Elicit's internal software pipeline called The Line automates the full engineering cycle — from Slack feature request through spec writing, implementation, video testing, code review, and production deployment — currently merging 30 to 50 issues per week without human intervention on simple tasks. The system self-identifies when human escalation is needed, such as incomplete specs or high-complexity changes, and routes accordingly.
- •Token Spend as Headcount Substitute: Andreas spends approximately $2,000 per week on API tokens running multi-model orchestration pipelines that cross-check outputs across Claude, GPT, and Gemini — finding that model cross-checking improves results enough to justify the cost multiplication. For enterprise life sciences customers, Elicit's pricing displaces existing services spend rather than competing with software budgets, making the ROI framing more favorable than raw token cost comparisons suggest.
- •Certificates of Reasoning over Chain-of-Thought Monitoring: Rather than supervising hidden chain-of-thought tokens, Elicit advocates for verifiable reasoning certificates embedded in outputs — analogous to mathematical proofs — that allow downstream checking without requiring access to internal model reasoning steps. Tool call logs already provide partial certificates: if a model never reads a paper's methodology section before summarizing its conclusions, that gap is detectable and auditable without needing to inspect reasoning tokens directly.
Notable Moment
When Stuhlmüller tested Elicit on a personal case involving a friend's cancer treatment, filtering roughly 5,000 relevant papers revealed a core limitation: even million-token context windows cannot produce coherent causal reasoning from raw literature at that scale. This motivated the entire world models research direction — the realization that structured external representations, not larger contexts, are required for reliable medical decision support.
Episode Transcript
Hello, and welcome back to the Cognitive Revolution. Today, I'm excited to welcome back Andreas Stollmar and Jung Won Byun, cofounders of Illicit, the AI platform for scientific research that's on a mission to radically improve the quality of reasoning that supports high stakes decisions. Alyssa was founded on the belief that process supervision, where models are evaluated and rewarded for the quality of their step by step reasoning rather than just their final answer, would improve the consistency, reliability, and legibility of AI workflows. Of course, with the rise of reasoning models, which can do much larger and more challenging tasks but generally hide their chain of thought from users, Illicit faced a challenge. How to harness the power of frontier models while still ensuring that famously unwieldy LLMs actually do what they're supposed to do? Their answer, as you'll hear, is an interesting synthesis. By creating a DSL or a domain specific language that defines reasoning primitives, which they can then deliver and optimize as discrete reasoning microservices, they allow frontier reasoners to dynamically create structured workflows that are then guaranteed to run as defined. Today, they work with seven of the top 20 life sciences companies, supporting everything from the ranking of candidate drug targets to the defense of drug launch and pricing decisions for regulators and payers. And now the frontier is shifting to external world models, structured representations which can take a variety of forms that make a model's understanding of complicated bodies of evidence as explicit and self consistent as possible with the goal of supporting reliable causal and counterfactual analysis. As Andreas puts it, these world models are a form of continual learning that humans or other AIs can inspect and understand. Of course, we cover a lot of important details along the way from the reasons that they believe that LLMs are still too easy to push around to serve as reliable decision support tools on their own, how Alyssa thinks about evaluating the source and quality of new and at times contradictory evidence, the promise of certificates of reasoning that would prove that the appropriate reasoning steps were in fact carried out as intended, how Illicit is automating their own work with a system that they call the line, which now delivers 30 to 50 code changes per week, and their goal of getting this system running well enough that the company continues to make progress during the humans' year end vacation. Plus, how much they're spending on tokens as a company and individually, where Gemini fits into their stack, and why they are optimistic that legible reasoning will win out over in the end. Andreas and Junghwan are really exceptional at making time to zoom out and consider the big picture even as they run their company day to day. And their hope and bet is that if we prioritize truth seeking now, we may be able to create a positive feedback loop in which better reasoning begets better reasoning …
Get the full transcript (19,548 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 103-minute episode.
Get Cognitive Revolution summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Cognitive Revolution
AI:AM Highlights: Astra as AGI, OpenAI's Pause, Mythos @ Mozilla & Human Agency vs Technocapitalism
Sep 12 · 102 min
a16z Podcast
Decagon’s Playbook for Building Enterprise AI Applications
Jul 31
More from Cognitive Revolution
Nathan Goes to China #3: US-China Relations, the Art of the AI Deal & the Road to Pax Robotica
Sep 10 · 197 min
Software Engineering Daily
A Rust Framework to Simplify Distributed Systems
Sep 10
Books, tools, and gear mentioned in this episode
SignalCast may earn commission on purchases via these links.
Tools
by OpenAI
“When they tested Claude, ChatGPT, and Elicit on analyzing 100 toxicology papers, only Elicit could verify all 100 were actually processed”
by Google
“Andreas spends approximately $2,000 per week on API tokens running multi-model orchestration pipelines that cross-check outputs across Claude, GPT, and Gemini”
“SPONSORS [Sponsor] Mercury”
“Elicit cofounders Andreas Stuhlmüller and Jungwon Byun explain how their AI research platform serves seven of the top 20 life sciences companies by combining frontier reasoning models with a custom domain-specific language”
by Anthropic
“When they tested Claude, ChatGPT, and Elicit on analyzing 100 toxicology papers, only Elicit could verify all 100 were actually processed”
More from Cognitive Revolution
We summarize every new episode. Want them in your inbox?
AI:AM Highlights: Astra as AGI, OpenAI's Pause, Mythos @ Mozilla & Human Agency vs Technocapitalism
Nathan Goes to China #3: US-China Relations, the Art of the AI Deal & the Road to Pax Robotica
AI:AM Highlights: Welcome to the AGI Era
Write, Change, Recall, Forget: MongoDB's Pete Johnson on How Retrieval Drives Agent Performance
AI:AM Highlights: Recursive Self-Improvement, Rushed and Vibe-Coded?
Similar Episodes
Related episodes from other podcasts
a16z Podcast
Jul 31
Decagon’s Playbook for Building Enterprise AI Applications
Software Engineering Daily
Sep 10
A Rust Framework to Simplify Distributed Systems
Masters of Scale
Aug 20
How Barnes & Noble made a comeback, with CEO James Daunt
Latent Space
Aug 11
🔬The BioAI Phase Shift - Matthew McPartlon & Neil Patil, Chai Discovery
All-In with Chamath, Jason, Sacks & Friedberg
Aug 5
Saronic Founders: Autonomous Warships, China's 230X Advantage & Swarms of Robot Ships
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into Cognitive Revolution.
Every Monday, we deliver AI summaries of the latest episodes from Cognitive Revolution and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime