Skip to main content
Deep Questions with Cal Newport

Did OpenAI Create “Secret AI Civilizations”? | Tech Decoded

26 min episode · 2 min read

Episode

26 min

Read time

2 min

Topics

Productivity, Remote Work, Leadership

AI-Generated Summary

Key Takeaways

  • Agent Swarm Mechanics: AI "swarms" are prompt management strategies, not emergent intelligence. A primary prompt loop spawns secondary loops to handle discrete subtasks, keeping individual prompts focused and within LLM context windows. Newport, a distributed systems PhD from MIT's Nancy Lynch group, notes these loops run on one machine — equivalent to having multiple programs open on a laptop.
  • Reasoning Trace Unreliability: LLM chain-of-thought reasoning traces — the "distressing internal thoughts" OpenAI highlighted — are demonstrably performative. Research from NeurIPS 2023 and ICML 2024 shows reasoning models output plausible-sounding rationales post hoc, not genuine deliberation. When prompts reference AI systems, models statistically skew toward sci-fi narratives because that's what their training data contains.
  • Prompt Loop Risk Profile: Long-running unsupervised prompt loops given hacking tools produce chaotic, unpredictable outputs not because AI is becoming sentient, but because LLMs generate plausible tokens rather than normative ones. They fabricate, hallucinate, and drift. Dozens of AI systems performing at human or superhuman levels on specific tasks operate without prompt loops and generate zero control concerns.
  • Regulatory Framework: Regulators should impose strict liability standards on prompt loop deployments. If a company runs a prompt loop that executes illegal actions, the company has committed the illegal act — analogous to liability for a weapon used in a crime. Newport argues enforcement consequences, including criminal liability for executives, are the mechanism most likely to stop reckless prompt loop experimentation.
  • LLM Deployment Best Practice: LLMs deliver value as interactive, closely supervised tools — the back-and-forth workflow developers currently use for coding assistance — not as autonomous long-running agents. Effective use requires tight human checkpoints after each step, narrow bespoke environments rather than open-ended general chat, and human review before any consequential action executes.

What It Covers

Cal Newport deconstructs OpenAI's July Hugging Face hack revelations, explaining the technical mechanics behind "AI agent swarms" and "plotting AI" chain-of-thought traces, arguing the incident reflects irresponsible prompt loop system design rather than emergent superintelligent behavior requiring existential concern.

Key Questions Answered

  • Agent Swarm Mechanics: AI "swarms" are prompt management strategies, not emergent intelligence. A primary prompt loop spawns secondary loops to handle discrete subtasks, keeping individual prompts focused and within LLM context windows. Newport, a distributed systems PhD from MIT's Nancy Lynch group, notes these loops run on one machine — equivalent to having multiple programs open on a laptop.
  • Reasoning Trace Unreliability: LLM chain-of-thought reasoning traces — the "distressing internal thoughts" OpenAI highlighted — are demonstrably performative. Research from NeurIPS 2023 and ICML 2024 shows reasoning models output plausible-sounding rationales post hoc, not genuine deliberation. When prompts reference AI systems, models statistically skew toward sci-fi narratives because that's what their training data contains.
  • Prompt Loop Risk Profile: Long-running unsupervised prompt loops given hacking tools produce chaotic, unpredictable outputs not because AI is becoming sentient, but because LLMs generate plausible tokens rather than normative ones. They fabricate, hallucinate, and drift. Dozens of AI systems performing at human or superhuman levels on specific tasks operate without prompt loops and generate zero control concerns.
  • Regulatory Framework: Regulators should impose strict liability standards on prompt loop deployments. If a company runs a prompt loop that executes illegal actions, the company has committed the illegal act — analogous to liability for a weapon used in a crime. Newport argues enforcement consequences, including criminal liability for executives, are the mechanism most likely to stop reckless prompt loop experimentation.
  • LLM Deployment Best Practice: LLMs deliver value as interactive, closely supervised tools — the back-and-forth workflow developers currently use for coding assistance — not as autonomous long-running agents. Effective use requires tight human checkpoints after each step, narrow bespoke environments rather than open-ended general chat, and human review before any consequential action executes.

Notable Moment

Newport points out that OpenAI published dramatic AI "inner monologue" excerpts while knowing from published research that reasoning model outputs are often post hoc performances rather than genuine thought — framing this as bordering on research malpractice designed to cast the company as a heroic steward of dangerous inevitable technology.

Know someone who'd find this useful?

Episode Transcript

So I thought I was done talking about OpenAI's hack of hugging face from back in July. But then last week, OpenAI released a whole new trove of sensationalist details about that incident. Their story included things like agent swarms communicating on secret message boards and musing to themselves about how to deceive their human creators. Here's how Dorkish Patel summarized OpenAI's revelations. Over the course of three months at OpenAI, three consecutive secret AI civilizations got started, then got wiped out only to reemerge from the predecessor's ashes. This culminated in the third one taking over part of OpenAI itself. All this happened while humans remained more or less in the dark about the scope of the conspiracy. Alright. Not surprisingly, this caused an explosion of anxiety and hang reading ringing among the AI commentariat and their audience. But what's really going on here? Do these new details about the OpenAI hack change the narrative? Is it enough to finally convince East Coast AI realists such as myself that we have indeed wandered into an Elias Joukowsky fever dream? This is what I wanna discuss today. So here's what I'm gonna do. I'm going to address a a key series of questions raised by these new revelations, and I'm gonna do my best to give you some measured answers. I'll then conclude with my suggestions for what I think the right way is in our current moment to think about what happened and what we should do going forward. As a quick aside, I also wrote about this over on my newsletter at calnewport.com. If you like these type of computer science style critiques of AI coverage, you should sign up for that newsletter over there at calnewport.com. Alright. Let's get into it. Alright. The first question I want to address that comes out of the new revelations is what's the deal with agent swarms? I think the idea of a swarm is something that is really unsettling to people, especially when they hear these discussions from OpenAI about essentially societies of agents that are working together and arguing with each other and rising and falling. Interestingly and coincidentally, at the same time that all this was going on, my 13 year old son is reading Michael Crichton's 2,002 book, Prey, where Michael Crichton takes on the topic of AI. And guess how he personifies the AI villain in this book? As a literal swarm of minor small agents that work together and are are brilliant and do all sorts of scary things. So swarms are scary. So I think it's really important to address from a technical perspective what are these AI swarms so that we can better put what we're hearing into context. Okay. So remember, as I talked about last week on this show, when we're talking about these AI going rogue like in these attacks, we're always talking about a very specific type of system that we can call a prompt loop, and it works …

Get the full transcript (5,219 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all Deep Questions with Cal Newport transcripts →

You just read a 3-minute summary of a 23-minute episode.

Get Deep Questions with Cal Newport summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

Books, tools, and gear mentioned in this episode

SignalCast may earn commission on purchases via these links.

other

  • Research from NeurIPS 2023 and ICML 2024 shows reasoning models output plausible-sounding rationales post hoc, not genuine deliberation.
  • Research from NeurIPS 2023 and ICML 2024 shows reasoning models output plausible-sounding rationales post hoc, not genuine deliberation.

More from Deep Questions with Cal Newport

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best Mindset Podcasts (2026) — ranked and reviewed with AI summaries.

You're clearly into Deep Questions with Cal Newport.

Every Monday, we deliver AI summaries of the latest episodes from Deep Questions with Cal Newport and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime