⚡️GPT5-Codex-Max: Training Agents with Personality, Tools & Trust — Brian Fioca + Bill Chen, OpenAI
Read time
2 min
Topics
Productivity, Investing, Fundraising & VC
AI-Generated Summary
Key Takeaways
- ✓Model Personality Training: GPT-5 coding models are trained on behavioral characteristics like communication, planning, and self-checking rather than just code completion. These software engineering best practices become measurable personality traits that build developer trust and enable longer autonomous operation without human intervention.
- ✓Tool Usage Habits: Codex develops specific tool preferences during training, performing better when tools are named exactly as trained. For example, naming a search tool "rg" instead of "grep" significantly improves performance because the model learned ripgrep conventions, demonstrating how training creates exploitable usage patterns.
- ✓Agent Abstraction Layer: The development paradigm shifts from optimizing individual model releases to packaging complete agents like Codex that platforms can integrate directly. This allows developers to build one layer above the model, avoiding constant updates to harnesses, sandboxing, and API changes while maintaining cutting-edge capabilities.
- ✓Multi-Turn Evaluation Challenge: Real-world agent evaluation requires assessing entire task trajectories, not single responses. Teams use LLM-as-judge to grade complete workflows, identify suboptimal steps, and have models self-improve by writing better instructions for future runs, creating a meta-prompting feedback loop that enhances agent performance over time.
What It Covers
OpenAI's Brian Fioca and Bill Chen explain how GPT-5 and Codex Max are trained with personality traits, tool usage patterns, and trust-building behaviors to create coding agents that run autonomously for twenty-four hours or more.
Key Questions Answered
- •Model Personality Training: GPT-5 coding models are trained on behavioral characteristics like communication, planning, and self-checking rather than just code completion. These software engineering best practices become measurable personality traits that build developer trust and enable longer autonomous operation without human intervention.
- •Tool Usage Habits: Codex develops specific tool preferences during training, performing better when tools are named exactly as trained. For example, naming a search tool "rg" instead of "grep" significantly improves performance because the model learned ripgrep conventions, demonstrating how training creates exploitable usage patterns.
- •Agent Abstraction Layer: The development paradigm shifts from optimizing individual model releases to packaging complete agents like Codex that platforms can integrate directly. This allows developers to build one layer above the model, avoiding constant updates to harnesses, sandboxing, and API changes while maintaining cutting-edge capabilities.
- •Multi-Turn Evaluation Challenge: Real-world agent evaluation requires assessing entire task trajectories, not single responses. Teams use LLM-as-judge to grade complete workflows, identify suboptimal steps, and have models self-improve by writing better instructions for future runs, creating a meta-prompting feedback loop that enhances agent performance over time.
Notable Moment
Brian Fioca reveals he has not written a single line of code by hand in months, relying entirely on Codex for all development work including launching open source projects, demonstrating the trust threshold senior engineers now place in AI coding agents.
Episode Transcript
Okay. We're here at AIECODE, and we we have, two of our speakers, Bill and Brian. Welcome. Hi. In in space. Thank you for having us. Bill, Brian, I I know you've been the listener for a little bit. Oh, yeah. What what's your take on lint space? Like, how how how does it what role does it perform in your functions at OpenAI? Yeah. I mean, first of all, love the name. Okay. I'm a massive latent space context management person. Tell the story behind the name. Give me the the chance. Yeah. So it start, we never had Latent Space as a name at the start. It was called lspace. Interesting. And, one of my readers, donated the domain name, latent.space. He's like, you want it? I'm like, yeah. Awesome. So latin just, like, came accidentally. So you're in the ether, but, like, I didn't have this domain yet. So I just I just, like, called it lspace. Lspace is like a physical Nice. Domain. Yeah. No. It's it's it's amazing. I love it because it's you're, like, always on the cutting edge and it goes into a lot of detail about all the things that, like, I should be keeping up with as part of my job, and there's so much to keep up with. Right? So there's only so many sources of of really good high quality information for what's, like, happening on a deep level. Well, you guys have your own podcast now. So I'm like, you know, I get a competition. Yeah. Well, I still listen to yours, and I I still think yours is really good. So you guys, I guess, are representing, like, a start ups team, codex Yeah. All the things. You just launched codex Yeah. The the Clarus Max Yep. For the journey Yep. Yesterday. Yep. We're good on namings. Yeah. I I do. The people do make friends. I think Tivo was like, yeah. You know, we're good at a lot of things, but not on eBay. I was like, well, why call it Max? Like, was there any, like, internal discussion? Yeah. I mean, it's complicated because it needs to be differentiated from the previous one. And the idea is, like, Max can run for a really long time. We can go twenty four hours or more. I've actually, like, sort of had it gone for more than that. And the name is is, you know It's inside codecs on the web. Is that how do you when you say a really long time, twenty four hours Oh, I on my on my oh, that's I think that was on the web inside of Kronos. I'm not sure. But I've actually done it on my local computer for for quite a bit longer than twenty four hours over the course of a couple days with closing my laptop and reopening it. But but the the name, you know, it you could come up with something like pro, but …
Get the full transcript (5,616 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
Get Latent Space summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Latent Space
🔬“We have foundation models for language, not for physics” — Anima Anandkumar, Bren Professor of Computing
Aug 26 · 83 min
Lenny's Podcast
AI’s third era: the rise of persistent AI coworkers | Tara Seshan (OpenAI’s product lead)
Aug 30
More from Latent Space
Simulation: the new Scaling Law — Joon Sung Park, Simile AI
Aug 21 · 69 min
Lenny's Podcast
OpenAI’s Head of Design: This is the best time in history to be a designer | Ian Silber
Aug 16
More from Latent Space
We summarize every new episode. Want them in your inbox?
🔬“We have foundation models for language, not for physics” — Anima Anandkumar, Bren Professor of Computing
Simulation: the new Scaling Law — Joon Sung Park, Simile AI
🔬The BioAI Phase Shift - Matthew McPartlon & Neil Patil, Chai Discovery
The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten
Codex from 0 to 10M Users: Building ChatGPT Work — Akshay Nathan, OpenAI
Similar Episodes
Related episodes from other podcasts
Lenny's Podcast
Aug 30
AI’s third era: the rise of persistent AI coworkers | Tara Seshan (OpenAI’s product lead)
Lenny's Podcast
Aug 16
OpenAI’s Head of Design: This is the best time in history to be a designer | Ian Silber
Cognitive Revolution
May 9
Milliseconds to Match: Criteo's AdTech AI & the Future of Commerce w/ Diarmuid Gill & Liva Ralaivola
This Week in Startups
Mar 27
The $60 billion resource hiding in space, and the start trying to mine it (feat. Matt Gialich, Astroforge) | E2268
The Prof G Pod
Mar 1
First Time Founders: Is Cohere the Next AI Powerhouse?
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Investing & Markets Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into Latent Space.
Every Monday, we deliver AI summaries of the latest episodes from Latent Space and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime