Skip to main content
Machine Learning Street Talk

New top score on ARC-AGI-2-pub (29.4%) - Jeremy Berman

68 min episode · 3 min read
·
Jeremy Berman

Episode

68 min

Read time

3 min

Topics

Startups, Fundraising & VC, Artificial Intelligence

AI-Generated Summary

Key Takeaways

  • Natural language over Python for ARC-AGI-2: Switching from Python programs to plain English descriptions of transformation rules dramatically improves performance on ARC-AGI-2 because every task can be described in five to ten bullet points. Python becomes brittle and verbose for compositional grid tasks, while natural language lets the model's inductive bias express itself fully, yielding higher accuracy at roughly $30 per task on v2 versus $8 on v1.
  • Breadth over depth for thinking models: On ARC-AGI-1, iterative revision loops were critical because models lacked internal reasoning. On ARC-AGI-2, RL-trained thinking models like Grok 4 perform deep revision internally, so the optimal strategy shifts toward maximizing entropy and breadth of initial generation rather than deep iterative refinement. Artificially increasing entropy in prompts consistently outperformed narrow, constrained prompting strategies.
  • Model selection is domain-specific and spiky: ARC leaderboard performance varies dramatically by model in ways other benchmarks do not. Grok 4 outperforms GPT-class models on ARC-AGI-2 grid reasoning, while Sonnet 3.5 remains superior for code generation tasks. Testing each model directly on the target domain rather than relying on general leaderboard rankings is necessary to identify the right tool for a specific problem type.
  • Reasoning as the meta-skill for AGI: Current LLMs acquire domain-specific reasoning circuits — math reasoning stays in math weights, science reasoning in science weights — with limited cross-domain transfer. The core AGI gap is not skill acquisition but the meta-skill of creating new skills. Aligning models purely toward general reasoning through RL, before layering domain knowledge, is the proposed path toward a foundation for general intelligence.
  • Knowledge trees versus knowledge webs: Pretraining treats all knowledge as an associative web of embeddings without guaranteed causal structure. Reinforcement learning with verifiable rewards functions as a pruning mechanism, replacing web-like associations with deductive trees where each node is causally consistent with its ancestors. The hypothesis is that models with weight configurations reflecting actual deductive structure will generalize to novel problems, while web-based models cannot.

What It Covers

Jeremy Berman, research scientist at Reflection AI, explains how he reached 29.4% on the ARC-AGI-2 public leaderboard using an evolutionary algorithm that generates and refines natural language descriptions of transformation rules rather than Python code, then discusses why reasoning is the meta-skill required for AGI and the fundamental gap between knowledge webs and deductive knowledge trees.

Key Questions Answered

  • Natural language over Python for ARC-AGI-2: Switching from Python programs to plain English descriptions of transformation rules dramatically improves performance on ARC-AGI-2 because every task can be described in five to ten bullet points. Python becomes brittle and verbose for compositional grid tasks, while natural language lets the model's inductive bias express itself fully, yielding higher accuracy at roughly $30 per task on v2 versus $8 on v1.
  • Breadth over depth for thinking models: On ARC-AGI-1, iterative revision loops were critical because models lacked internal reasoning. On ARC-AGI-2, RL-trained thinking models like Grok 4 perform deep revision internally, so the optimal strategy shifts toward maximizing entropy and breadth of initial generation rather than deep iterative refinement. Artificially increasing entropy in prompts consistently outperformed narrow, constrained prompting strategies.
  • Model selection is domain-specific and spiky: ARC leaderboard performance varies dramatically by model in ways other benchmarks do not. Grok 4 outperforms GPT-class models on ARC-AGI-2 grid reasoning, while Sonnet 3.5 remains superior for code generation tasks. Testing each model directly on the target domain rather than relying on general leaderboard rankings is necessary to identify the right tool for a specific problem type.
  • Reasoning as the meta-skill for AGI: Current LLMs acquire domain-specific reasoning circuits — math reasoning stays in math weights, science reasoning in science weights — with limited cross-domain transfer. The core AGI gap is not skill acquisition but the meta-skill of creating new skills. Aligning models purely toward general reasoning through RL, before layering domain knowledge, is the proposed path toward a foundation for general intelligence.
  • Knowledge trees versus knowledge webs: Pretraining treats all knowledge as an associative web of embeddings without guaranteed causal structure. Reinforcement learning with verifiable rewards functions as a pruning mechanism, replacing web-like associations with deductive trees where each node is causally consistent with its ancestors. The hypothesis is that models with weight configurations reflecting actual deductive structure will generalize to novel problems, while web-based models cannot.
  • Catastrophic forgetting blocks continual learning more than compute does: The fundamental barrier to adaptive, continuously learning AI is not computational cost but catastrophic forgetting — fine-tuning on new data drifts weights away from previously correct solutions. Proposed directions include freezing expert layers, composable model architectures analogous to Docker's immutable layers, and selective data mixtures during fine-tuning. Solving this problem is framed as the next S-curve after the current RL scaling wave.

Notable Moment

Berman argues that heavy pretraining may actively slow reasoning development rather than accelerate it. His analogy contrasts consultants who know terminology but cannot derive conclusions with Feynman-style thinkers who deduce everything from first principles — and frames RL post-training as the process of converting one into the other.

Know someone who'd find this useful?

Episode Transcript

Get in the game with the college branded Venmo debit card. Rack your team with every tap and earn up to 5% cash back with Venmo Stash, a new rewards program from Venmo. No monthly fee, no minimum balance. Just school pride and spending power. Get in the game and sign up for the Venmo debit card at venmo.com/collegecard. The Venmo Mastercard is issued by the Bancorp Bank NA. Select schools available. Venmo's stash terms and exclusions apply at venmo.me/stashterms. Max, $100 cash back per month. This episode is brought to you by White Claw Surge. Nice choice hitting up this podcast. No surprises. You're all about diving into tastes everyone in the room can enjoy, just like White Claw Surge. It's for celebrating those moments when connections have been made and the night's just begun. With bold flavors and 8% alcohol by volume, unleash the night. Unleash White Claw Surge. Please drink responsibly. Hard seltzer with flavors, 8% alcohol by volume, White Claw Seltzer Works, Chicago, Illinois. You can describe every single Arc v two task in 10 bullet points of plain English, most of them in five bullet points. And I think that this actually gets to the heart of Arc. Right? Everything is quite simple. It's not very hard. And I think this is also how we do it too. Right? Like, we when we look at these arc graphs, we're coming up with these bullet points in our head and we're, you know, checking them. Okay. This was right. This was right. And Python doesn't have these features. It's just not as expressive as natural language. MLST is sponsored by Cyberfund. Link in the description. I get actually, even more fundamentally, like, the ideal system would be we have a set of data. Our language model is bad at a certain thing. We can just give it this data and then all of a sudden, it keeps all of its knowledge and then also gets really good at this new thing. We we are not there yet and that to me is, like, a fundamental, missing part. Really, what you want is more expressive program. And so that's why I switched from Python to English, which is a much more expressive program. You can language you can always teach a language model skill. Right? But it's the meta skill. It's the skill to create the skills that is AGI. And to me, that's reasoning. Like, reasoning is that meta skill. And so, to put it another way, I think if you fundamentally learn the skill of reasoning, you should be able to then, apply that skill to learn all the other skills. That is the meta skill. You know, kick whatever weights out you need to, align the model to reason. And then from there, you have a foundation from which you can actually build general intelligence. Okay, folks. Hot off the press. Many of you would have seen last week that Jeremy Berman, who is …

Get the full transcript (14,243 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all Machine Learning Street Talk transcripts →

You just read a 3-minute summary of a 65-minute episode.

Get Machine Learning Street Talk summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

Books, tools, and gear mentioned in this episode

SignalCast may earn commission on purchases via these links.

Tools

  • Grok-4Recommended

    by xAI

    RL-trained thinking models like Grok 4 perform deep revision internally, so the optimal strategy shifts toward maximizing entropy and breadth of initial generation rather than deep iterative refinement.
  • by OpenAI

    Grok 4 outperforms GPT-class models on ARC-AGI-2 grid reasoning, while Sonnet 3.5 remains superior for code generation tasks.
  • by Anthropic

    Grok 4 outperforms GPT-class models on ARC-AGI-2 grid reasoning, while Sonnet 3.5 remains superior for code generation tasks.

company

  • Jeremy Berman, research scientist at Reflection AI, explains how he reached 29.4% on the ARC-AGI-2 public leaderboard.

other

  • Jeremy Berman, research scientist at Reflection AI, explains how he reached 29.4% on the ARC-AGI-2 public leaderboard using an evolutionary algorithm that generates and refines natural language descriptions of transformation rules.

More from Machine Learning Street Talk

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into Machine Learning Street Talk.

Every Monday, we deliver AI summaries of the latest episodes from Machine Learning Street Talk and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime