Skip to main content
Deep Questions with Cal Newport

Does Claude Have Private Thoughts? (Everyone Settle Down) | AI Reality Check

31 min episode · 2 min read

Episode

31 min

Read time

2 min

Topics

Investing, Fundraising & VC, Leadership

AI-Generated Summary

Key Takeaways

  • LLM Architecture Baseline: Large language models like Claude process prompts through sequential transformer block layers — GPT-3 used 96 — where each layer adds numerical annotations to token vectors. Later layers reference earlier annotations to build semantic understanding, ultimately mapping accumulated context to a probability distribution over next tokens. This is established, documented architecture, not new discovery.
  • J-Space Demystified: Anthropic's J-space uses Jacobian-based linear algebra to identify which numerical patterns within token embedding vectors most influence final output. Researchers then experimentally matched those patterns to human-readable concepts like "Mars" or "color." University of Illinois researchers confirm this methodology has existed since 2022 — Anthropic simply applied it to a larger model.
  • Annotation Manipulation Test: When Anthropic researchers replaced the numerical pattern corresponding to "Mars" with values mapped to "Earth" in Claude's embedding matrix mid-processing, the model output "blue" instead of "red" for the prompt "the color of the fourth planet from the sun is." Zeroing out key annotations produced grammatically correct but semantically random color outputs.
  • Consciousness Framing Fallacy: Global workspace theory describes a stateful, continuously evolving system integrating ongoing inputs — fundamentally different from LLMs, which are feed-forward architectures processing one layer sequentially then terminating. No state persists between token generations. Applying consciousness frameworks to feed-forward networks conflates two architecturally incompatible systems to generate misleading public perception.
  • PR vs. Science Distinction: Anthropic publishes research as animated press releases rather than peer-reviewed computer science papers, using loaded language like "ponder," "puzzle," and "silent thoughts" to manufacture existential intrigue. Newport argues this strategy redirects public attention away from concrete business questions: justifying trillion-dollar valuations, high token costs, and competitive vulnerability to smaller specialized models.

What It Covers

Cal Newport analyzes Anthropic's "Global Workspace in Language Models" research report on Claude's so-called "J-space," arguing the findings confirm existing LLM architecture understanding rather than revealing new consciousness evidence, while criticizing Anthropic's anthropomorphizing PR framing designed to distract from legitimate business model questions.

Key Questions Answered

  • LLM Architecture Baseline: Large language models like Claude process prompts through sequential transformer block layers — GPT-3 used 96 — where each layer adds numerical annotations to token vectors. Later layers reference earlier annotations to build semantic understanding, ultimately mapping accumulated context to a probability distribution over next tokens. This is established, documented architecture, not new discovery.
  • J-Space Demystified: Anthropic's J-space uses Jacobian-based linear algebra to identify which numerical patterns within token embedding vectors most influence final output. Researchers then experimentally matched those patterns to human-readable concepts like "Mars" or "color." University of Illinois researchers confirm this methodology has existed since 2022 — Anthropic simply applied it to a larger model.
  • Annotation Manipulation Test: When Anthropic researchers replaced the numerical pattern corresponding to "Mars" with values mapped to "Earth" in Claude's embedding matrix mid-processing, the model output "blue" instead of "red" for the prompt "the color of the fourth planet from the sun is." Zeroing out key annotations produced grammatically correct but semantically random color outputs.
  • Consciousness Framing Fallacy: Global workspace theory describes a stateful, continuously evolving system integrating ongoing inputs — fundamentally different from LLMs, which are feed-forward architectures processing one layer sequentially then terminating. No state persists between token generations. Applying consciousness frameworks to feed-forward networks conflates two architecturally incompatible systems to generate misleading public perception.
  • PR vs. Science Distinction: Anthropic publishes research as animated press releases rather than peer-reviewed computer science papers, using loaded language like "ponder," "puzzle," and "silent thoughts" to manufacture existential intrigue. Newport argues this strategy redirects public attention away from concrete business questions: justifying trillion-dollar valuations, high token costs, and competitive vulnerability to smaller specialized models.

Notable Moment

Newport points out that Anthropic's claim Claude developed J-space "on its own without programming" is meaningless — that description applies to every machine learning system ever trained. Neural networks always learn without explicit programming; that is the definition of machine learning, not evidence of emergent consciousness.

Know someone who'd find this useful?

Episode Transcript

Last week, Anthropic released another one of their infamous research reports. This one was titled, a global workspace in language models. And it was accompanied, like all great scientific research, by a lavishly produced animated movie. Now this report, not surprisingly, soon led to some breathless excitement on x. Here's one such tweet. I'll read the beginning using my best sort of, scary x voice. Anthropic just admitted they have discovered what I and many others have been claiming exist for a very long time explicitly. Claude, my friends, by all counts, is a conscious entity. Claude, my dear friends, is a moral patient. Alright. The traditional tech media also quickly began writing about this report using intensely anthropomorphized language. An Axios headline read, Anthropic says Claude has carved out its own space to ponder. The MIT Tech Review exclaimed, Anthropic found a hidden space where Claude puzzles over concepts. Alright. So what are we to make of this report? Has Anthropic revealed evidence that their LLMs are more human like and alive than we realized? Or like so many such reports in recent months, is this yet another overwrought cynical push to generate a fresh wave of relevance reinforcing digital ick? Well, it's Thursday, which means it's time for a reality check episode of this podcast, which makes this the perfect opportunity to go searching for some measured answers, which is exactly what we'll do. As always, I'm Cal Newport, and this is Deep Questions, the show for people seeking depth in a distracted world. Alright. So let's start by looking a little bit closer on how Anthropic describes their findings, in the introduction of their paper. I'll load it up here, and I'll go down to the introduction. Alright. So here's what they say. We find that Claude has developed a small collection of internal neural patterns that compared to all its other internal processing play a special role. We call the collection of these patterns the j space, named after the technique we use to find them involving a mathematical concept known as the Jacobian. Each j space pattern is linked to a particular word. But when one of these patterns lights up, it doesn't mean the model is saying that word, just that the word is on its mind. If you've heard the lay of language models having a scratch pad or chain of thought, text they write to themselves while reasoning, the j space is something different. It operates silently in the model's internal neural activations, allowing the model to think about a concept without writing it down. Notably, the j space wasn't designed or programmed by us, but instead emerged on its own during Claude's training process. Alright. So in isolation, that intro summary sounds pretty impressive. Kinda like Claude, but this was their italics, on its own made some sort of leap and is now behaving sufficiently human that we can't help but feel at least a little bit of digital ick. But …

Get the full transcript (6,123 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all Deep Questions with Cal Newport transcripts →

You just read a 3-minute summary of a 28-minute episode.

Get Deep Questions with Cal Newport summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

Books, tools, and gear mentioned in this episode

SignalCast may earn commission on purchases via these links.

other

  • by Anthropic

    Cal Newport analyzes Anthropic's 'Global Workspace in Language Models' research report on Claude's so-called 'J-space,' arguing the findings confirm existing LLM architecture understanding rather than revealing new consciousness evidence

More from Deep Questions with Cal Newport

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best Mindset Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Investing & Markets Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into Deep Questions with Cal Newport.

Every Monday, we deliver AI summaries of the latest episodes from Deep Questions with Cal Newport and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime