Skip to main content
Latent Space

🔬Scaling Past Informal AI - Carina Hong, Axiom Math

93 min episode · 3 min read
·

Episode

93 min

Read time

3 min

Topics

Productivity, Remote Work, Investing

AI-Generated Summary

Key Takeaways

  • Verified Generation as Performance Gain: Formal verification is not a quality-control tax but a direct performance multiplier. Axiom's system scored 120/120 on the 2025 Putnam exam, outperforming the best human score of 110 and DeepSeek's 103, using orders of magnitude less compute and data than frontier labs. This demonstrates that verified generation produces higher sample efficiency, allowing smaller teams to exceed frontier lab benchmarks on structured reasoning tasks.
  • Lean as Dual-Purpose Infrastructure: Lean functions simultaneously as a functional programming language and a formal proof checker via the Curry-Howard correspondence, which maps proofs to programs. Developers can write autograd in Lean, verify distributed systems components, or prove mathematical theorems within the same environment. Practitioners building AI reasoning pipelines should evaluate Lean not as a niche academic tool but as a Turing-complete substrate for co-generating code and correctness proofs together.
  • Formal Math Transfer Learning Parallels Coding: Anthropic's coding focus in 2023-2024 was underestimated because structured, formal data transfers horizontally across reasoning domains rather than staying vertical. Axiom applies the same thesis to formal math: Lean proof data provides structured, verifiable training signal that transfers to software verification, hardware verification, and general reasoning. Teams building reasoning systems should prioritize formally verifiable data sources over volume of informal chain-of-thought data.
  • Hardware Verification as Near-Term Revenue Anchor: Chip verification currently requires a 1:3 to 1:4 ratio of engineering time and team size dedicated purely to verification relative to design. There is no partial credit for a mostly-verified GPU. Axiom positions formal proof generation as a direct replacement for manual verification labor in ASIC projects, where a single unverified edge case invalidates the entire design. This represents a concrete, high-value enterprise deployment path beyond pure mathematics research.
  • Axle Tooling Enables Claude-Based Lean Workflows: Axiom released Axle (Axiom Lean Engine), a free suite of 14 Lean meta-programming tools including VerifyProof, which runs 100x faster than the prior standard tool Comparator. On the Verina code-verification benchmark of 189 problems, Axiom's system solved 187 with no benchmark-specific modifications. Developers can integrate Axle directly with Claude Code today to generate and verify Lean proofs without configuring a local Lean toolchain, lowering the barrier to formal verification in production workflows.

What It Covers

Carina Hong, CEO of Axiom Math, explains how formal verification using the Lean proof language enables verified AI reasoning rather than merely correcting hallucinations. Axiom scored 120/120 on the 2025 Putnam exam, raised $200M at a $1.6B valuation, and argues formal math provides transfer learning advantages that informal LLM scaling cannot replicate at superintelligence scale.

Key Questions Answered

  • Verified Generation as Performance Gain: Formal verification is not a quality-control tax but a direct performance multiplier. Axiom's system scored 120/120 on the 2025 Putnam exam, outperforming the best human score of 110 and DeepSeek's 103, using orders of magnitude less compute and data than frontier labs. This demonstrates that verified generation produces higher sample efficiency, allowing smaller teams to exceed frontier lab benchmarks on structured reasoning tasks.
  • Lean as Dual-Purpose Infrastructure: Lean functions simultaneously as a functional programming language and a formal proof checker via the Curry-Howard correspondence, which maps proofs to programs. Developers can write autograd in Lean, verify distributed systems components, or prove mathematical theorems within the same environment. Practitioners building AI reasoning pipelines should evaluate Lean not as a niche academic tool but as a Turing-complete substrate for co-generating code and correctness proofs together.
  • Formal Math Transfer Learning Parallels Coding: Anthropic's coding focus in 2023-2024 was underestimated because structured, formal data transfers horizontally across reasoning domains rather than staying vertical. Axiom applies the same thesis to formal math: Lean proof data provides structured, verifiable training signal that transfers to software verification, hardware verification, and general reasoning. Teams building reasoning systems should prioritize formally verifiable data sources over volume of informal chain-of-thought data.
  • Hardware Verification as Near-Term Revenue Anchor: Chip verification currently requires a 1:3 to 1:4 ratio of engineering time and team size dedicated purely to verification relative to design. There is no partial credit for a mostly-verified GPU. Axiom positions formal proof generation as a direct replacement for manual verification labor in ASIC projects, where a single unverified edge case invalidates the entire design. This represents a concrete, high-value enterprise deployment path beyond pure mathematics research.
  • Axle Tooling Enables Claude-Based Lean Workflows: Axiom released Axle (Axiom Lean Engine), a free suite of 14 Lean meta-programming tools including VerifyProof, which runs 100x faster than the prior standard tool Comparator. On the Verina code-verification benchmark of 189 problems, Axiom's system solved 187 with no benchmark-specific modifications. Developers can integrate Axle directly with Claude Code today to generate and verify Lean proofs without configuring a local Lean toolchain, lowering the barrier to formal verification in production workflows.
  • Blueprint Authorship Remains the Human Bottleneck: Large-scale formalization projects like sphere packing in 8 dimensions still rely on human-authored blueprints that decompose theorems into subtasks assignable across contributors. Auto-generated blueprints — high-level proof sketches that structure collaborative formalization — remain an unsolved technical problem that multiple groups are racing to crack. Mathematicians and AI researchers targeting collaborative theorem proving should focus engineering effort on blueprint generation as the current rate-limiting step, not proof search itself.
  • Specification Gap Limits Enterprise Deployment: Formal verification guarantees correctness only relative to a written specification, and humans consistently underspecify what they want from complex systems. A financial audit system or flight controller cannot be fully verified if the specification omits edge cases. Axiom's near-term mitigation combines mutation-based LLM unit-test generation to surface unspecified cases as conjecture proposals, then feeds confirmed specifications to the prover. Teams adopting formal verification should invest in specification tooling and conjecture generation before expecting end-to-end automated correctness guarantees.

Notable Moment

When discussing why informal LLM scaling cannot reach mathematical superintelligence, Hong points out that frontier math benchmarks required collaboration with EPFL because expert evaluators are genuinely scarce — there are not enough humans who understand results in the Langlands program to grade outputs at scale. Infinite compute budgets cannot solve a human-attention bottleneck.

Know someone who'd find this useful?

Episode Transcript

But it's for the first time now, I think verified AI is to open up collaboration. Either it's human AI collaboration. Well, before blueprinting, that's human human collaboration, and lean was the grounding, was the verification, formal language. And then human AI collaboration, like we're seeing now, future AI agent agent agent like collaboration. Like, I think verified AI is for openness. It's not for meeting the requirements of closed industries. And I think just like I think verification should not be about, oh, I remember, like, you know, was there this article, like, chatbots mixed up. Oh, there's math solution hallucination. Verification to me is not about lousiness. Verification to me is about scaling brilliance, compounding brilliance. It's like just kind of going back to the collaboration point. It's about Ramanujan being a much stronger mathematician. He was already a really strong one, but verification helps him extend the brilliance, like, both kind of, like, scale up and scale out. Welcome to the Latent Space AR for Science podcast. I'm Brandon Anderson. I build, RNA therapeutics at Atomic AI, and I'm joined by RJ Honeke, the CTO of Mirroromics, working on spatial transcriptomics. It's a pleasure to have Karina Hong, CEO and founder of Acxiom Math. Acxiom has made a splash in several different areas. First, they were they got a perfect score in the Putnam, last, December, I think. They also have the claim of the first AI to prove research conjectures using formal verification. And, very exciting, they just yesterday announced, quite a large, series a. Yeah. Welcome to the show. Thank you for having me. You just raised $200,000,000, which as one of your colleagues said, this is, like, basically the entire, like, US math budget for math research each year. Is that true, actually? It's according to his LinkedIn post. Yeah. Okay. Wow. 250,000,000 is our fairly annual, math budget. I think we should spend more on math research. Yes. Yeah. It's kinda sad. Yeah. I know. But anyway, like, you know, as a, you know, as a nerd who loves math, that's, like, really cool. But, I mean, I'm just like that kinda blew my mind. Like, what? Whenever I heard that, like okay. So, like, yeah. How is it $202,100,000,000? I guess, 1,600,000,000 valuation? Yeah. I don't know. Yeah. Well, super super excited to be here. Also, I think, like, you know, this is a Thursday, so it's a very, very interesting timely timely podcast. We're, like, a seven, eight months old company, so it definitely means a lot to us. It's a really cool milestone. We're currently about, like, 30 people now. Right? So kind of going into I think this amount of funding will, like, give us a feel that we it needs to to to accelerate Yeah. The strong execution momentum that we have so far. I think, like, people think of us, like, there are many kind of ways to think about Acxiom. People think of us as us …

Get the full transcript (17,153 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all Latent Space transcripts →

You just read a 3-minute summary of a 90-minute episode.

Get Latent Space summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

Books, tools, and gear mentioned in this episode

SignalCast may earn commission on purchases via these links.

Tools

  • LeanRecommended
    Lean as Dual-Purpose Infrastructure: Lean functions simultaneously as a functional programming language and a formal proof checker via the Curry-Howard correspondence, which maps proofs to programs. Developers can write autograd in Lean, verify distributed systems components, or prove mathematical theorems within the same environment.
  • ClaudeRecommended

    by Anthropic

    Developers can integrate Axle directly with Claude Code today to generate and verify Lean proofs without configuring a local Lean toolchain, lowering the barrier to formal verification in production workflows.
  • VerifyProofRecommendedBy guest

    by Axiom Math

    Axiom released Axle (Axiom Lean Engine), a free suite of 14 Lean meta-programming tools including VerifyProof, which runs 100x faster than the prior standard tool Comparator.
  • Axle (Axiom Lean Engine), a free suite of 14 Lean meta-programming tools including VerifyProof, which runs 100x faster than the prior standard tool Comparator.
  • AxleRecommendedBy guest

    by Axiom Math

    Axiom released Axle (Axiom Lean Engine), a free suite of 14 Lean meta-programming tools including VerifyProof, which runs 100x faster than the prior standard tool Comparator.

More from Latent Space

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Investing & Markets Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into Latent Space.

Every Monday, we deliver AI summaries of the latest episodes from Latent Space and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime