Skip to main content
The TWIML AI Podcast

Rethinking Pre-Training for Agentic AI with Aakanksha Chowdhery - #759

52 min episode · 2 min read
·
Aakanksha Chowdhery

Episode

52 min

Read time

2 min

Topics

Health & Wellness, Fundraising & VC, Artificial Intelligence

AI-Generated Summary

Key Takeaways

  • Pretraining for agents: Current models train on static benchmarks like GLUE or GSM8K, but agentic tasks require interactive environment capabilities. Pretraining must fundamentally change attention mechanisms, loss objectives, and data composition, not just rely on post-training fixes to achieve multistep reasoning and tool use.
  • Long context reasoning: Models need attention mechanisms that enable reasoning over millions of tokens while maintaining retrieval and synthesis capabilities. Current transformers struggle with multi-hop reasoning benchmarks like MRCR v2 and LOFT, even with long context windows, requiring architectural modifications for agent workflows.
  • Training data augmentation: Dominant pretraining sources like internet articles must be augmented with reasoning traces at comparable token volumes. Masking specific portions during training, similar to fill-in-the-middle for code models, teaches models which tools to use and how to plan across multiple steps.
  • Failure recovery capability: Models must learn to recognize failed trajectory steps in their context and choose different action spaces rather than repeating probabilistic mistakes. This requires both reinforcement learning objectives and pretraining formats that help models notice and correct from previous errors during multistep problem solving.

What It Covers

Aakanksha Chowdhery from Reflection explains why pretraining language models specifically for agentic capabilities requires rethinking attention mechanisms, loss objectives, and training data composition beyond current post-training approaches that optimize static benchmarks.

Key Questions Answered

  • Pretraining for agents: Current models train on static benchmarks like GLUE or GSM8K, but agentic tasks require interactive environment capabilities. Pretraining must fundamentally change attention mechanisms, loss objectives, and data composition, not just rely on post-training fixes to achieve multistep reasoning and tool use.
  • Long context reasoning: Models need attention mechanisms that enable reasoning over millions of tokens while maintaining retrieval and synthesis capabilities. Current transformers struggle with multi-hop reasoning benchmarks like MRCR v2 and LOFT, even with long context windows, requiring architectural modifications for agent workflows.
  • Training data augmentation: Dominant pretraining sources like internet articles must be augmented with reasoning traces at comparable token volumes. Masking specific portions during training, similar to fill-in-the-middle for code models, teaches models which tools to use and how to plan across multiple steps.
  • Failure recovery capability: Models must learn to recognize failed trajectory steps in their context and choose different action spaces rather than repeating probabilistic mistakes. This requires both reinforcement learning objectives and pretraining formats that help models notice and correct from previous errors during multistep problem solving.

Notable Moment

Chowdhery reveals that halfway through training PaLM, the team tested it on a crowdsourced reasoning benchmark and discovered a sudden step change in performance, indicating emergent reasoning capabilities that would not have been detected without diverse evaluation tasks beyond standard metrics.

Know someone who'd find this useful?

Episode Transcript

For the longest time we were measuring, pre training on static benchmarks. If we want these models to be useful as agents, they need to be able to interact with environments. And when we start caring about those agentic tasks, pretraining needs to rethink from fundamentals. This is not just a post training problem to achieve these set of capabilities that we want in the next generation of models. And the kind of benchmarks we need for measuring this kind of intelligence is sometimes not available today. Alright, everyone. Welcome to another episode of the TwiML AI podcast. I'm your host, Sam Charrington. Today, I'm joined by Akansha Choudhury. Akansha is a member of technical staff at Reflection. Before we get going, be sure to hit that subscribe button wherever you're listening to today's show. Kansha, welcome to the podcast. Thank you, Sam. You have a really interesting background. You've trained some of the earliest large language models, including POM and, Gemini one dot o, Gemini one dot five. Tell us a little bit about those experiences. I got into large language models at Google while building one of the distributed systems, that led to the training of, PALM, which is which was our large our largest language model at the time. It had five forty billion parameters. And, people stopped publishing the number of parameters after that. One and, that led me to be, at the forefront of pretraining solving one set of problems after the other that come with scale in first two generations of palm models and then the first two generations of Gemini models. And I think the one thing that you learn when you do pretraining is that at scale, every problem magnifies and things go wrong at every possible part of the stack. So it's always fun and it's always exciting. And you said you were on the infrastructure side? I I was on the ML side, but I've done both. One thing that is super interesting about, pretraining is that you have to be able to think across the stack. So, otherwise, you if you're gonna train something for two or two and a half months, you need to be able to think about every single part of the system. Nice. And so tell us about reflection. What is reflection focused on? So the mission of Reflection is to build frontier open, intelligence, for agentic capabilities. And, the company has been focused on building, post training stack for agentic tasks. And with the most recent fund raise, we are, doing training end to end. So we're building, different tier open agentic models, which are both pretrained and post trained in house. And that's really gonna be a a focus of what we're talking about today, some of the reasons why you think pretraining a different approach to pretraining is key? Is that kinda the way you think about it? Well, the way I really put it is that, for the longest time, we …

Get the full transcript (8,434 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all The TWIML AI Podcast transcripts →

You just read a 3-minute summary of a 49-minute episode.

Get The TWIML AI Podcast summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

Books, tools, and gear mentioned in this episode

SignalCast may earn commission on purchases via these links.

company

  • Sponsors: Google - https://ai.studio/build
  • Aakanksha Chowdhery from Reflection explains why pretraining language models specifically for agentic capabilities requires rethinking attention mechanisms.

More from The TWIML AI Podcast

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Health & Longevity Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into The TWIML AI Podcast.

Every Monday, we deliver AI summaries of the latest episodes from The TWIML AI Podcast and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime