Skip to main content
Cognitive Revolution

Nathan Goes to China – Part 2: AI Safety with Chinese Characteristics

137 min episode · 3 min read
·
Chinese Characteristics

Episode

137 min

Read time

3 min

Topics

Productivity, Relationships, Investing

AI-Generated Summary

Key Takeaways

  • US Safety Lead is Narrow Without Top Two: American AI models lead Chinese counterparts on safety benchmarks, but this advantage is almost entirely carried by OpenAI and Anthropic. Remove those two companies, and the gap between the remaining US models and Chinese models shrinks dramatically. Google's Gemini holds a modest edge, but the broader ecosystem on both sides looks far more comparable than headline comparisons suggest. Concordia AI's aimonitor.net tracks these evaluations publicly across all major models.
  • The 45-Degree Line Framework: Shanghai AI Lab director Zhao Boen introduced a guiding principle at WAIC 2024 stating that AI safety measures must grow in direct proportion to capabilities — a slope-of-one relationship. This framework is broadly accepted across China's AI ecosystem. In practice, Chinese models currently fall below this line, meaning capabilities are outpacing safety measures, but the principle itself signals institutional awareness and provides a concrete benchmark for evaluating progress over time.
  • Open-Weights vs. Service-Level Safety Thinking: Chinese AI safety thinking focuses on regulating deployed services rather than base models, because trillion-parameter models require serious infrastructure to run and cannot be casually misused by individuals. This contrasts with the US approach of evaluating worst-case model behavior in isolation. Both perspectives have merit: the Chinese framing captures realistic deployment risk, while the US framing captures tail risks from sophisticated actors who can access raw model weights.
  • China's AI Safety Research is Growing at 10x in Three Years: Concordia AI's aisafetychina.com tracks Chinese AI safety publications, which grew from a handful of papers monthly in 2023 to 50–60 papers per month by mid-2026. US and anglosphere output runs roughly 50 to a few hundred papers monthly, meaning China is approaching comparable volume. Paper topics directly mirror US research: eval faking, deception benchmarks, mechanistic interpretability, self-replication red lines, and gradient routing for hazardous capability isolation.
  • Tsinghua AI Safety Hub Targets Constellation and Lisa as Models: A new AI safety research hub launched at Tsinghua University's College of AI days before WAIC 2026, with zero coverage in English-language media. The hub has five founding board members including one European professor taking a resident position. Leadership explicitly named Constellation and Lisa in London as institutional models to emulate, committed to hosting international researchers in residence, and plans to fund Chinese students to spend time at global AI safety hubs.

What It Covers

Nathan Labenz reports from a two-week China trip on the state of Chinese AI safety, covering deployed model safeguards, academic research output, government regulation, and cross-cultural exchange between US and Chinese AI safety communities. The episode systematically challenges the assumption that China ignores AI safety, using data from Concordia AI, Xi Jinping's WAIC speech, and a new Tsinghua University AI safety hub launch.

Key Questions Answered

  • US Safety Lead is Narrow Without Top Two: American AI models lead Chinese counterparts on safety benchmarks, but this advantage is almost entirely carried by OpenAI and Anthropic. Remove those two companies, and the gap between the remaining US models and Chinese models shrinks dramatically. Google's Gemini holds a modest edge, but the broader ecosystem on both sides looks far more comparable than headline comparisons suggest. Concordia AI's aimonitor.net tracks these evaluations publicly across all major models.
  • The 45-Degree Line Framework: Shanghai AI Lab director Zhao Boen introduced a guiding principle at WAIC 2024 stating that AI safety measures must grow in direct proportion to capabilities — a slope-of-one relationship. This framework is broadly accepted across China's AI ecosystem. In practice, Chinese models currently fall below this line, meaning capabilities are outpacing safety measures, but the principle itself signals institutional awareness and provides a concrete benchmark for evaluating progress over time.
  • Open-Weights vs. Service-Level Safety Thinking: Chinese AI safety thinking focuses on regulating deployed services rather than base models, because trillion-parameter models require serious infrastructure to run and cannot be casually misused by individuals. This contrasts with the US approach of evaluating worst-case model behavior in isolation. Both perspectives have merit: the Chinese framing captures realistic deployment risk, while the US framing captures tail risks from sophisticated actors who can access raw model weights.
  • China's AI Safety Research is Growing at 10x in Three Years: Concordia AI's aisafetychina.com tracks Chinese AI safety publications, which grew from a handful of papers monthly in 2023 to 50–60 papers per month by mid-2026. US and anglosphere output runs roughly 50 to a few hundred papers monthly, meaning China is approaching comparable volume. Paper topics directly mirror US research: eval faking, deception benchmarks, mechanistic interpretability, self-replication red lines, and gradient routing for hazardous capability isolation.
  • Tsinghua AI Safety Hub Targets Constellation and Lisa as Models: A new AI safety research hub launched at Tsinghua University's College of AI days before WAIC 2026, with zero coverage in English-language media. The hub has five founding board members including one European professor taking a resident position. Leadership explicitly named Constellation and Lisa in London as institutional models to emulate, committed to hosting international researchers in residence, and plans to fund Chinese students to spend time at global AI safety hubs.
  • CAC Regulatory Process Includes a Six-Month Precedent for Slowing Companies: In 2023, China's Cyberspace Administration of China halted LLM product launches for approximately six months while establishing review standards, demonstrating willingness to impose significant commercial costs for safety reasons. Current process requires provincial-level API testing followed by national CAC review before any new AI service can launch publicly. Companies report weekly to daily contact with regulators, and incremental model updates face lighter review while major capability jumps trigger full re-evaluation.
  • Chinese Big Tech Monitors US AI Safety Discourse Daily via Agents: At least one major Chinese technology company runs an automated agent that surveys US AI safety publications, social media, and research daily, compiling a briefing for leadership. This asymmetry — Chinese companies systematically tracking US safety discourse while US companies have no equivalent China-monitoring practice — reflects a broader pattern where Chinese AI institutions cite US organizations including Apollo Research, Palisade, METR, and Redwood Research by name in published papers and conference presentations.

Notable Moment

At a lunch with a Chinese professor who has published extensively on AI safety, Nathan asked whether colleagues trade probability-of-doom estimates the way Bay Area researchers do. The professor seemed genuinely surprised this was normal American behavior, suggesting that while Chinese researchers engage seriously with AI risk, the cultural framing around existential probability remains largely absent from their professional conversations.

Know someone who'd find this useful?

You just read a 3-minute summary of a 134-minute episode.

Get Cognitive Revolution summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

Books, tools, and gear mentioned in this episode

SignalCast may earn commission on purchases via these links.

Tools

  • by Concordia AI

    Concordia AI's aimonitor.net tracks these evaluations publicly across all major models.
  • by Concordia AI

    Concordia AI's aisafetychina.com tracks Chinese AI safety publications, which grew from a handful of papers monthly in 2023 to 50–60 papers per month by mid-2026.

company

  • Chinese AI institutions cite US organizations including Apollo Research, Palisade, METR, and Redwood Research by name in published papers and conference presentations.
  • Chinese AI institutions cite US organizations including Apollo Research, Palisade, METR, and Redwood Research by name in published papers and conference presentations.
  • Chinese AI institutions cite US organizations including Apollo Research, Palisade, METR, and Redwood Research by name in published papers and conference presentations.
  • Chinese AI institutions cite US organizations including Apollo Research, Palisade, METR, and Redwood Research by name in published papers and conference presentations.

More from Cognitive Revolution

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Investing & Markets Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into Cognitive Revolution.

Every Monday, we deliver AI summaries of the latest episodes from Cognitive Revolution and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime