Skip to main content
Cognitive Revolution

Nathan Goes to China – Part 2: AI Safety with Chinese Characteristics

137 min episode · 3 min read
·
Chinese Characteristics

Episode

137 min

Read time

3 min

Topics

Productivity, Relationships, Investing

AI-Generated Summary

Key Takeaways

  • US Safety Lead is Narrow Without Top Two: American AI models lead Chinese counterparts on safety benchmarks, but this advantage is almost entirely carried by OpenAI and Anthropic. Remove those two companies, and the gap between the remaining US models and Chinese models shrinks dramatically. Google's Gemini holds a modest edge, but the broader ecosystem on both sides looks far more comparable than headline comparisons suggest. Concordia AI's aimonitor.net tracks these evaluations publicly across all major models.
  • The 45-Degree Line Framework: Shanghai AI Lab director Zhao Boen introduced a guiding principle at WAIC 2024 stating that AI safety measures must grow in direct proportion to capabilities — a slope-of-one relationship. This framework is broadly accepted across China's AI ecosystem. In practice, Chinese models currently fall below this line, meaning capabilities are outpacing safety measures, but the principle itself signals institutional awareness and provides a concrete benchmark for evaluating progress over time.
  • Open-Weights vs. Service-Level Safety Thinking: Chinese AI safety thinking focuses on regulating deployed services rather than base models, because trillion-parameter models require serious infrastructure to run and cannot be casually misused by individuals. This contrasts with the US approach of evaluating worst-case model behavior in isolation. Both perspectives have merit: the Chinese framing captures realistic deployment risk, while the US framing captures tail risks from sophisticated actors who can access raw model weights.
  • China's AI Safety Research is Growing at 10x in Three Years: Concordia AI's aisafetychina.com tracks Chinese AI safety publications, which grew from a handful of papers monthly in 2023 to 50–60 papers per month by mid-2026. US and anglosphere output runs roughly 50 to a few hundred papers monthly, meaning China is approaching comparable volume. Paper topics directly mirror US research: eval faking, deception benchmarks, mechanistic interpretability, self-replication red lines, and gradient routing for hazardous capability isolation.
  • Tsinghua AI Safety Hub Targets Constellation and Lisa as Models: A new AI safety research hub launched at Tsinghua University's College of AI days before WAIC 2026, with zero coverage in English-language media. The hub has five founding board members including one European professor taking a resident position. Leadership explicitly named Constellation and Lisa in London as institutional models to emulate, committed to hosting international researchers in residence, and plans to fund Chinese students to spend time at global AI safety hubs.

What It Covers

Nathan Labenz reports from a two-week China trip on the state of Chinese AI safety, covering deployed model safeguards, academic research output, government regulation, and cross-cultural exchange between US and Chinese AI safety communities. The episode systematically challenges the assumption that China ignores AI safety, using data from Concordia AI, Xi Jinping's WAIC speech, and a new Tsinghua University AI safety hub launch.

Key Questions Answered

  • US Safety Lead is Narrow Without Top Two: American AI models lead Chinese counterparts on safety benchmarks, but this advantage is almost entirely carried by OpenAI and Anthropic. Remove those two companies, and the gap between the remaining US models and Chinese models shrinks dramatically. Google's Gemini holds a modest edge, but the broader ecosystem on both sides looks far more comparable than headline comparisons suggest. Concordia AI's aimonitor.net tracks these evaluations publicly across all major models.
  • The 45-Degree Line Framework: Shanghai AI Lab director Zhao Boen introduced a guiding principle at WAIC 2024 stating that AI safety measures must grow in direct proportion to capabilities — a slope-of-one relationship. This framework is broadly accepted across China's AI ecosystem. In practice, Chinese models currently fall below this line, meaning capabilities are outpacing safety measures, but the principle itself signals institutional awareness and provides a concrete benchmark for evaluating progress over time.
  • Open-Weights vs. Service-Level Safety Thinking: Chinese AI safety thinking focuses on regulating deployed services rather than base models, because trillion-parameter models require serious infrastructure to run and cannot be casually misused by individuals. This contrasts with the US approach of evaluating worst-case model behavior in isolation. Both perspectives have merit: the Chinese framing captures realistic deployment risk, while the US framing captures tail risks from sophisticated actors who can access raw model weights.
  • China's AI Safety Research is Growing at 10x in Three Years: Concordia AI's aisafetychina.com tracks Chinese AI safety publications, which grew from a handful of papers monthly in 2023 to 50–60 papers per month by mid-2026. US and anglosphere output runs roughly 50 to a few hundred papers monthly, meaning China is approaching comparable volume. Paper topics directly mirror US research: eval faking, deception benchmarks, mechanistic interpretability, self-replication red lines, and gradient routing for hazardous capability isolation.
  • Tsinghua AI Safety Hub Targets Constellation and Lisa as Models: A new AI safety research hub launched at Tsinghua University's College of AI days before WAIC 2026, with zero coverage in English-language media. The hub has five founding board members including one European professor taking a resident position. Leadership explicitly named Constellation and Lisa in London as institutional models to emulate, committed to hosting international researchers in residence, and plans to fund Chinese students to spend time at global AI safety hubs.
  • CAC Regulatory Process Includes a Six-Month Precedent for Slowing Companies: In 2023, China's Cyberspace Administration of China halted LLM product launches for approximately six months while establishing review standards, demonstrating willingness to impose significant commercial costs for safety reasons. Current process requires provincial-level API testing followed by national CAC review before any new AI service can launch publicly. Companies report weekly to daily contact with regulators, and incremental model updates face lighter review while major capability jumps trigger full re-evaluation.
  • Chinese Big Tech Monitors US AI Safety Discourse Daily via Agents: At least one major Chinese technology company runs an automated agent that surveys US AI safety publications, social media, and research daily, compiling a briefing for leadership. This asymmetry — Chinese companies systematically tracking US safety discourse while US companies have no equivalent China-monitoring practice — reflects a broader pattern where Chinese AI institutions cite US organizations including Apollo Research, Palisade, METR, and Redwood Research by name in published papers and conference presentations.

Notable Moment

At a lunch with a Chinese professor who has published extensively on AI safety, Nathan asked whether colleagues trade probability-of-doom estimates the way Bay Area researchers do. The professor seemed genuinely surprised this was normal American behavior, suggesting that while Chinese researchers engage seriously with AI risk, the cultural framing around existential probability remains largely absent from their professional conversations.

Know someone who'd find this useful?

Episode Transcript

Hello, and welcome back to the Cognitive Revolution. This is gonna be part two of the Nathan Goes to China series, and we're calling it AI Safety with Chinese Characteristics. Part one, if you haven't heard that, is up on the feed. It's been up for a few days. And in that part one of the series, I really just laid out the tech setup for going to China, what you would need to do if you wanna get a cell phone ready, the apps that you need to download. I also shared some of my experience using Chinese AIs on the ground there and then just shared a bunch of observations based on my two weeks in China and many conversations I had. I appreciate the kind comments that I've got in response to that one, including at least one from a listener in China, which was the probably the one that mattered to me the most. I think if you've been to China in the last few years, you can probably skip that one. But if you're interested in this, you might also be interested in that. Though I do think they should be pretty self contained. So part one is a little bit closer to a travel log and tech review. And this episode is gonna be really focused on the Chinese AI safety ecosystem and trying to just describe it and, as much as possible, understand it on its own terms. As we saw last time, there are a lot of similarities, Just as there are major similarities between American big tech and Chinese big tech, there are a lot of similarities between American AI safety and Chinese AI safety communities and the work that they're producing. But there are also some differences, and I think it will definitely be helpful if we have a better understanding of those. Before I really get into it, just a couple quick disclaimers again. I will be, again, following a Chatham House rule for this episode, not because I was asked to do that, but just because I wanna keep things simple for myself and make sure that I'm protecting everyone that I talk to from being misrepresented by me and put in an uncomfortable position. So I won't be naming any names or attributing things to the the organizations or institutions that people are affiliated with. The one exception to that for this particular episode will be when I'm citing published papers or reports, then those I can give you the name because those are out there in the public domain. And there are some good sources that I'll mention and we can link to in the show notes of this episode for anybody who wants to go deeper, which, of course, is always recommended. The other point of order or clarification, just in case anyone is wondering, I have no financial conflicts on this matter. I took this trip paying my own way, flights and hotels, all …

Get the full transcript (22,802 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all Cognitive Revolution transcripts →

You just read a 3-minute summary of a 134-minute episode.

Get Cognitive Revolution summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

Books, tools, and gear mentioned in this episode

SignalCast may earn commission on purchases via these links.

Tools

  • by Concordia AI

    Concordia AI's aimonitor.net tracks these evaluations publicly across all major models.
  • by Concordia AI

    Concordia AI's aisafetychina.com tracks Chinese AI safety publications, which grew from a handful of papers monthly in 2023 to 50–60 papers per month by mid-2026.

company

  • Chinese AI institutions cite US organizations including Apollo Research, Palisade, METR, and Redwood Research by name in published papers and conference presentations.
  • Chinese AI institutions cite US organizations including Apollo Research, Palisade, METR, and Redwood Research by name in published papers and conference presentations.
  • Chinese AI institutions cite US organizations including Apollo Research, Palisade, METR, and Redwood Research by name in published papers and conference presentations.
  • Chinese AI institutions cite US organizations including Apollo Research, Palisade, METR, and Redwood Research by name in published papers and conference presentations.

More from Cognitive Revolution

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Investing & Markets Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into Cognitive Revolution.

Every Monday, we deliver AI summaries of the latest episodes from Cognitive Revolution and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime