Skip to main content
The AI Breakdown

How to Decide What Work AI Should Do for You: The AI Deputization Audit

29 min episode · 2 min read

Episode

29 min

Read time

2 min

Topics

Productivity, Personal Finance, Relationships

AI-Generated Summary

Key Takeaways

  • Deputization Audit Scoring: Rate each recurring task across five dimensions — frequency/time cost, teachability, output checkability, error stakes, and personal necessity — on a 0–2 scale. Tasks scoring 8–10 out of 10 are strong candidates for full AI handoff; scores of 4–7 suggest a human-AI collaboration model; scores of 0–3 should stay with the human.
  • Ambient vs. Deliberate Context-Building: ChatGPT's computer history uses passive screen observation to build ongoing context automatically, while GrokBot's teach-a-task requires intentionally pressing record and demonstrating a workflow. Workers who prefer control should use deliberate demonstration; those who want zero friction should use ambient observation — both solve the same context gap.
  • Blocker Identification as a Decision Tool: After scoring tasks, naming the specific blocker preventing AI handoff reveals whether new tools actually solve it. Vendor portals with no API and hard-to-explain processes are now addressable via computer-use agents and screen recording; tasks requiring judgment, human relationships, or irreversible-mistake risk remain unsolvable by current tools.
  • Token Efficiency Beats Sticker Price: An AlphaSense study of hundreds of financial analysis questions found GPT-5.6 Soul completed tasks 13% cheaper than Kimi K3 while scoring 20% higher on quality. More expensive-per-token models often consume fewer total tokens, making cost-per-task the correct metric for enterprise AI budgeting — not cost-per-token.
  • Speed as a Distinct Model Dimension: Gemini 3.7 Flash runs at 340 tokens per second and OpenAI's ultra-fast Soul mode reaches 750 tokens per second via Cerebras hardware. For synchronous coding or real-time voice workflows where a human is actively waiting on output, speed improvements can justify higher per-task costs even at slightly lower benchmark performance.

What It Covers

Two new AI features — GrokBot's teach-a-task recording and ChatGPT's computer history — shift AI's core bottleneck from capability to context. The episode introduces a five-criteria "deputization audit" scoring system to help workers identify which recurring tasks to hand off to AI agents versus keep themselves.

Key Questions Answered

  • Deputization Audit Scoring: Rate each recurring task across five dimensions — frequency/time cost, teachability, output checkability, error stakes, and personal necessity — on a 0–2 scale. Tasks scoring 8–10 out of 10 are strong candidates for full AI handoff; scores of 4–7 suggest a human-AI collaboration model; scores of 0–3 should stay with the human.
  • Ambient vs. Deliberate Context-Building: ChatGPT's computer history uses passive screen observation to build ongoing context automatically, while GrokBot's teach-a-task requires intentionally pressing record and demonstrating a workflow. Workers who prefer control should use deliberate demonstration; those who want zero friction should use ambient observation — both solve the same context gap.
  • Blocker Identification as a Decision Tool: After scoring tasks, naming the specific blocker preventing AI handoff reveals whether new tools actually solve it. Vendor portals with no API and hard-to-explain processes are now addressable via computer-use agents and screen recording; tasks requiring judgment, human relationships, or irreversible-mistake risk remain unsolvable by current tools.
  • Token Efficiency Beats Sticker Price: An AlphaSense study of hundreds of financial analysis questions found GPT-5.6 Soul completed tasks 13% cheaper than Kimi K3 while scoring 20% higher on quality. More expensive-per-token models often consume fewer total tokens, making cost-per-task the correct metric for enterprise AI budgeting — not cost-per-token.
  • Speed as a Distinct Model Dimension: Gemini 3.7 Flash runs at 340 tokens per second and OpenAI's ultra-fast Soul mode reaches 750 tokens per second via Cerebras hardware. For synchronous coding or real-time voice workflows where a human is actively waiting on output, speed improvements can justify higher per-task costs even at slightly lower benchmark performance.

Notable Moment

Microsoft's Windows Recall feature caused widespread privacy outrage in early 2024 for taking periodic screen screenshots. ChatGPT's computer history does something structurally similar but records interaction events rather than screenshots — and user reaction has shifted from alarm to enthusiasm, reflecting a measurable attitude change toward AI data access in under two years.

Know someone who'd find this useful?

Episode Transcript

What if I told you that figuring out what parts of your work you should be getting AI to automate was a simple math equation? This week, two new products came online that make getting AI to do work for you much simpler. Grokbot gives users the ability to teach it a task by manually recording them doing something, while ChatGPT's computer history watches how you work and learns over time. Together, these represent the shift of the biggest challenge in AI moving from capability to context. But as these new features come online, you still have to figure out which part of your work you want AI to automate. The work best suited for AI deputization is frequent, time consuming, teachable, easily verifiable, and doesn't require you to have been the one to do it to be successful. The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI. Alright, friends. Quick announcements before we dive in. First of all, thank you to today's sponsors, KPMG, Blitsy, Harbor, and HyperAgent. To get an ad free version of the show, go to patreon.com/aidailybrief, or you can subscribe on Apple Podcasts. To learn more about sponsoring the show, send us a note at sponsors@aidailybrief.ai. And one more thing before we get into the headlines, including a new model from Gemini, a lot of the shows this week including today's show have to do with the newly launched Grokbot. And for those of you who are doing our AI summer adventure choose your own adventure learning program, we've just posted a new pop up destination, I e project, all about testing out Grokbot. You can sign up for free at summeradventure.ai and walk through how to use this always on teammate that I think is the simplest, cleanest version of Open Claw we've ever had. Again, you can find that at summeradventure.ai. But now, let's get into the headlines. It's almost like Google heard us yesterday talking about how many people were talking about SpaceX AI as though it had completely usurped Google in the pantheon of serious frontier model companies. On Thursday, Google released Gemini 3.7 Flash. And while it is neither the much delayed Gemini 3.5 Pro nor the now increasingly anticipated Gemini four, it does play in a different category of efficiency that's becoming a higher and higher consideration, especially for serious and advanced users. So let's start with what's good about this model. It appears to be very, very fast. During testing from artificial analysis, the model ran at 340 tokens per second, which is an entirely different category than anything else. It's more than twice as fast as g p t five six Luna and even a bit faster than NVIDIA's new Neemotron 3.5 lightning. On the benchmarks, Google made some solid gains over 3.6 flash, most notably improving their score on coding benchmark Deep Sway from 48.6 to 65.3%. And with this model, Google is also slashing prices …

Get the full transcript (5,933 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all The AI Breakdown transcripts →

You just read a 3-minute summary of a 26-minute episode.

Get The AI Breakdown summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

More from The AI Breakdown

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.

You're clearly into The AI Breakdown.

Every Monday, we deliver AI summaries of the latest episodes from The AI Breakdown and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime