🔬Scaling Past Informal AI - Carina Hong, Axiom Math
Episode
93 min
Read time
3 min
Topics
Productivity, Remote Work, Investing
AI-Generated Summary
Key Takeaways
- ✓Verified Generation as Performance Gain: Formal verification is not a quality-control tax but a direct performance multiplier. Axiom's system scored 120/120 on the 2025 Putnam exam, outperforming the best human score of 110 and DeepSeek's 103, using orders of magnitude less compute and data than frontier labs. This demonstrates that verified generation produces higher sample efficiency, allowing smaller teams to exceed frontier lab benchmarks on structured reasoning tasks.
- ✓Lean as Dual-Purpose Infrastructure: Lean functions simultaneously as a functional programming language and a formal proof checker via the Curry-Howard correspondence, which maps proofs to programs. Developers can write autograd in Lean, verify distributed systems components, or prove mathematical theorems within the same environment. Practitioners building AI reasoning pipelines should evaluate Lean not as a niche academic tool but as a Turing-complete substrate for co-generating code and correctness proofs together.
- ✓Formal Math Transfer Learning Parallels Coding: Anthropic's coding focus in 2023-2024 was underestimated because structured, formal data transfers horizontally across reasoning domains rather than staying vertical. Axiom applies the same thesis to formal math: Lean proof data provides structured, verifiable training signal that transfers to software verification, hardware verification, and general reasoning. Teams building reasoning systems should prioritize formally verifiable data sources over volume of informal chain-of-thought data.
- ✓Hardware Verification as Near-Term Revenue Anchor: Chip verification currently requires a 1:3 to 1:4 ratio of engineering time and team size dedicated purely to verification relative to design. There is no partial credit for a mostly-verified GPU. Axiom positions formal proof generation as a direct replacement for manual verification labor in ASIC projects, where a single unverified edge case invalidates the entire design. This represents a concrete, high-value enterprise deployment path beyond pure mathematics research.
- ✓Axle Tooling Enables Claude-Based Lean Workflows: Axiom released Axle (Axiom Lean Engine), a free suite of 14 Lean meta-programming tools including VerifyProof, which runs 100x faster than the prior standard tool Comparator. On the Verina code-verification benchmark of 189 problems, Axiom's system solved 187 with no benchmark-specific modifications. Developers can integrate Axle directly with Claude Code today to generate and verify Lean proofs without configuring a local Lean toolchain, lowering the barrier to formal verification in production workflows.
What It Covers
Carina Hong, CEO of Axiom Math, explains how formal verification using the Lean proof language enables verified AI reasoning rather than merely correcting hallucinations. Axiom scored 120/120 on the 2025 Putnam exam, raised $200M at a $1.6B valuation, and argues formal math provides transfer learning advantages that informal LLM scaling cannot replicate at superintelligence scale.
Key Questions Answered
- •Verified Generation as Performance Gain: Formal verification is not a quality-control tax but a direct performance multiplier. Axiom's system scored 120/120 on the 2025 Putnam exam, outperforming the best human score of 110 and DeepSeek's 103, using orders of magnitude less compute and data than frontier labs. This demonstrates that verified generation produces higher sample efficiency, allowing smaller teams to exceed frontier lab benchmarks on structured reasoning tasks.
- •Lean as Dual-Purpose Infrastructure: Lean functions simultaneously as a functional programming language and a formal proof checker via the Curry-Howard correspondence, which maps proofs to programs. Developers can write autograd in Lean, verify distributed systems components, or prove mathematical theorems within the same environment. Practitioners building AI reasoning pipelines should evaluate Lean not as a niche academic tool but as a Turing-complete substrate for co-generating code and correctness proofs together.
- •Formal Math Transfer Learning Parallels Coding: Anthropic's coding focus in 2023-2024 was underestimated because structured, formal data transfers horizontally across reasoning domains rather than staying vertical. Axiom applies the same thesis to formal math: Lean proof data provides structured, verifiable training signal that transfers to software verification, hardware verification, and general reasoning. Teams building reasoning systems should prioritize formally verifiable data sources over volume of informal chain-of-thought data.
- •Hardware Verification as Near-Term Revenue Anchor: Chip verification currently requires a 1:3 to 1:4 ratio of engineering time and team size dedicated purely to verification relative to design. There is no partial credit for a mostly-verified GPU. Axiom positions formal proof generation as a direct replacement for manual verification labor in ASIC projects, where a single unverified edge case invalidates the entire design. This represents a concrete, high-value enterprise deployment path beyond pure mathematics research.
- •Axle Tooling Enables Claude-Based Lean Workflows: Axiom released Axle (Axiom Lean Engine), a free suite of 14 Lean meta-programming tools including VerifyProof, which runs 100x faster than the prior standard tool Comparator. On the Verina code-verification benchmark of 189 problems, Axiom's system solved 187 with no benchmark-specific modifications. Developers can integrate Axle directly with Claude Code today to generate and verify Lean proofs without configuring a local Lean toolchain, lowering the barrier to formal verification in production workflows.
- •Blueprint Authorship Remains the Human Bottleneck: Large-scale formalization projects like sphere packing in 8 dimensions still rely on human-authored blueprints that decompose theorems into subtasks assignable across contributors. Auto-generated blueprints — high-level proof sketches that structure collaborative formalization — remain an unsolved technical problem that multiple groups are racing to crack. Mathematicians and AI researchers targeting collaborative theorem proving should focus engineering effort on blueprint generation as the current rate-limiting step, not proof search itself.
- •Specification Gap Limits Enterprise Deployment: Formal verification guarantees correctness only relative to a written specification, and humans consistently underspecify what they want from complex systems. A financial audit system or flight controller cannot be fully verified if the specification omits edge cases. Axiom's near-term mitigation combines mutation-based LLM unit-test generation to surface unspecified cases as conjecture proposals, then feeds confirmed specifications to the prover. Teams adopting formal verification should invest in specification tooling and conjecture generation before expecting end-to-end automated correctness guarantees.
Notable Moment
When discussing why informal LLM scaling cannot reach mathematical superintelligence, Hong points out that frontier math benchmarks required collaboration with EPFL because expert evaluators are genuinely scarce — there are not enough humans who understand results in the Langlands program to grade outputs at scale. Infinite compute budgets cannot solve a human-attention bottleneck.
Episode Transcript
But it's for the first time now, I think verified AI is to open up collaboration. Either it's human AI collaboration. Well, before blueprinting, that's human human collaboration, and lean was the grounding, was the verification, formal language. And then human AI collaboration, like we're seeing now, future AI agent agent agent like collaboration. Like, I think verified AI is for openness. It's not for meeting the requirements of closed industries. And I think just like I think verification should not be about, oh, I remember, like, you know, was there this article, like, chatbots mixed up. Oh, there's math solution hallucination. Verification to me is not about lousiness. Verification to me is about scaling brilliance, compounding brilliance. It's like just kind of going back to the collaboration point. It's about Ramanujan being a much stronger mathematician. He was already a really strong one, but verification helps him extend the brilliance, like, both kind of, like, scale up and scale out. Welcome to the Latent Space AR for Science podcast. I'm Brandon Anderson. I build, RNA therapeutics at Atomic AI, and I'm joined by RJ Honeke, the CTO of Mirroromics, working on spatial transcriptomics. It's a pleasure to have Karina Hong, CEO and founder of Acxiom Math. Acxiom has made a splash in several different areas. First, they were they got a perfect score in the Putnam, last, December, I think. They also have the claim of the first AI to prove research conjectures using formal verification. And, very exciting, they just yesterday announced, quite a large, series a. Yeah. Welcome to the show. Thank you for having me. You just raised $200,000,000, which as one of your colleagues said, this is, like, basically the entire, like, US math budget for math research each year. Is that true, actually? It's according to his LinkedIn post. Yeah. Okay. Wow. 250,000,000 is our fairly annual, math budget. I think we should spend more on math research. Yes. Yeah. It's kinda sad. Yeah. I know. But anyway, like, you know, as a, you know, as a nerd who loves math, that's, like, really cool. But, I mean, I'm just like that kinda blew my mind. Like, what? Whenever I heard that, like okay. So, like, yeah. How is it $202,100,000,000? I guess, 1,600,000,000 valuation? Yeah. I don't know. Yeah. Well, super super excited to be here. Also, I think, like, you know, this is a Thursday, so it's a very, very interesting timely timely podcast. We're, like, a seven, eight months old company, so it definitely means a lot to us. It's a really cool milestone. We're currently about, like, 30 people now. Right? So kind of going into I think this amount of funding will, like, give us a feel that we it needs to to to accelerate Yeah. The strong execution momentum that we have so far. I think, like, people think of us, like, there are many kind of ways to think about Acxiom. People think of us as us …
Get the full transcript (17,153 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 90-minute episode.
Get Latent Space summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Latent Space
🔬“We have foundation models for language, not for physics” — Anima Anandkumar, Bren Professor of Computing
Aug 26 · 83 min
Gradient Dissent
The $64M Bet on an AI That Has to Be Right | Carina Hong, CEO of Axiom
Feb 5
More from Latent Space
Simulation: the new Scaling Law — Joon Sung Park, Simile AI
Aug 21 · 69 min
The TWIML AI Podcast
Building an AI Mathematician with Carina Hong - #754
Nov 4
Books, tools, and gear mentioned in this episode
SignalCast may earn commission on purchases via these links.
Tools
- LeanRecommended
“Lean as Dual-Purpose Infrastructure: Lean functions simultaneously as a functional programming language and a formal proof checker via the Curry-Howard correspondence, which maps proofs to programs. Developers can write autograd in Lean, verify distributed systems components, or prove mathematical theorems within the same environment.”
- ClaudeRecommended
by Anthropic
“Developers can integrate Axle directly with Claude Code today to generate and verify Lean proofs without configuring a local Lean toolchain, lowering the barrier to formal verification in production workflows.”
by Axiom Math
“Axiom released Axle (Axiom Lean Engine), a free suite of 14 Lean meta-programming tools including VerifyProof, which runs 100x faster than the prior standard tool Comparator.”
“Axle (Axiom Lean Engine), a free suite of 14 Lean meta-programming tools including VerifyProof, which runs 100x faster than the prior standard tool Comparator.”
by Axiom Math
“Axiom released Axle (Axiom Lean Engine), a free suite of 14 Lean meta-programming tools including VerifyProof, which runs 100x faster than the prior standard tool Comparator.”
More from Latent Space
We summarize every new episode. Want them in your inbox?
🔬“We have foundation models for language, not for physics” — Anima Anandkumar, Bren Professor of Computing
Simulation: the new Scaling Law — Joon Sung Park, Simile AI
🔬The BioAI Phase Shift - Matthew McPartlon & Neil Patil, Chai Discovery
The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten
Codex from 0 to 10M Users: Building ChatGPT Work — Akshay Nathan, OpenAI
Similar Episodes
Related episodes from other podcasts
Gradient Dissent
Feb 5
The $64M Bet on an AI That Has to Be Right | Carina Hong, CEO of Axiom
The TWIML AI Podcast
Nov 4
Building an AI Mathematician with Carina Hong - #754
Cognitive Revolution
Feb 18
Mathematical Superintelligence: Harmonic's Vlad Tenev & Tudor Achim on IMO Gold & Theories of Everything
Lex Fridman Podcast
Jun 15
#472 – Terence Tao: Hardest Problems in Mathematics, Physics & the Future of AI
Cognitive Revolution
Jul 12
Alignment with Awakening: Davidad on Moral Realism, AI Wisdom, & why His p(Doom) is Down to 5%
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Investing & Markets Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into Latent Space.
Every Monday, we deliver AI summaries of the latest episodes from Latent Space and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime