[NeurIPS Best Paper] 1000 Layer Networks for Self-Supervised RL — Kevin Wang et al, Princeton
Episode
28 min
Read time
2 min
Topics
Productivity, Startups, Fundraising & VC
AI-Generated Summary
Key Takeaways
- ✓Self-Supervised RL Objective: The breakthrough required shifting from traditional value-based RL to contrastive representation learning that classifies whether future states belong to the same trajectory, converting RL into a scalable classification problem similar to language models.
- ✓Architectural Recipe for Depth: Scaling depth alone failed initially. Success required combining residual connections, layer normalization, and specific architectural components together. Critical performance jumps occurred only when depth exceeded 50-64 layers with these modifications in place.
- ✓Parameter Efficiency Trade-offs: Scaling network depth grows parameters linearly while scaling width grows them quadratically. Depth scaling proved more sample-efficient and parameter-efficient, achieving state-of-the-art performance on goal-conditioned RL tasks with single H100 GPU training runs.
- ✓JAX GPU Acceleration Enables Scale: Using JAX-based GPU-accelerated environments allows collecting thousands of parallel trajectories simultaneously. Performance improvements only manifest after 50 million transitions, making this data throughput essential for training deep networks in RL settings.
What It Covers
Princeton researchers Kevin Wang and team achieved NeurIPS Best Paper by scaling reinforcement learning networks to 1000 layers using self-supervised learning objectives, challenging the field's conventional shallow architecture approach.
Key Questions Answered
- •Self-Supervised RL Objective: The breakthrough required shifting from traditional value-based RL to contrastive representation learning that classifies whether future states belong to the same trajectory, converting RL into a scalable classification problem similar to language models.
- •Architectural Recipe for Depth: Scaling depth alone failed initially. Success required combining residual connections, layer normalization, and specific architectural components together. Critical performance jumps occurred only when depth exceeded 50-64 layers with these modifications in place.
- •Parameter Efficiency Trade-offs: Scaling network depth grows parameters linearly while scaling width grows them quadratically. Depth scaling proved more sample-efficient and parameter-efficient, achieving state-of-the-art performance on goal-conditioned RL tasks with single H100 GPU training runs.
- •JAX GPU Acceleration Enables Scale: Using JAX-based GPU-accelerated environments allows collecting thousands of parallel trajectories simultaneously. Performance improvements only manifest after 50 million transitions, making this data throughput essential for training deep networks in RL settings.
Notable Moment
The advisor Ben initially doubted the approach would work based on prior failed attempts at deeper RL networks, but agreed to support the research bet because infrastructure improvements made experimentation low-cost and precedent from other domains suggested potential.
Episode Transcript
Welcome to Lanespace. We are basically trying to provide the best optimal sort of podcast experience of NeurIPS for people who are not here. And congrats on your paper. How's it feel? Yeah. It was very exciting. Yeah. We had a poster yester yesterday, and then today we'll have an oral talk. Were you just, like, mobbed? Oh, yeah. There was a lot of people. It's, like, three hours straight of, like, you know, like, waves of people to, like, know what we were trying to do. But So I've never received the best paper. Did you just find out on the website? Like, what, I just, like, woke up one day and, like, checked my email, and then they just tell they just Yeah. They was like, oh, like, that's hey. Like, I saw you, you know, oh, you were like, been awarded best paper. I'll let you Maybe you know from the reviews as well. Right? Sorry. Tell me Yeah. We know from the reviews that we did well, but there's a difference between, like, doing well in the reviews and getting best paper. So right before we didn't actually know. Yeah. Yeah. Okay. So I I I skipped a little bit. Maybe we can go sort of, one by one and and sort of introduce, you know, who you are and what you did on on on on the team. I'm Kevin. I was an undergrad from from Princeton, and I just graduated. And, yeah, I guess I led the project, like, started the project, and then well, we're very happy to collaborate with Ishan and Nicole and Ben also. Right. And were you in, like, the same research group? Like, how do you how do what's your idea? We're social context. So so yeah. So we're all from Princeton. Yeah. With that. Thanks to Alan for booking you guys. So this project actually started from, like, an IW seminar. So, like like, an independent work research seminar, that Ben was teaching. And this was, like, actually, like like, one of my first experiences in, like, ML research. So it was really valuable to, like, get that experience. And then Ishaan was also in that seminar and working on adjacent things, so we collaborated, a lot during that seminar. And then, yeah, the project turned out to have some pretty cool results. And then later on, also, like, the Halt working on sort of similar things also, joined it on the project and became, like, a good collaboration. Yeah. And, I I don't know if any of you guys wanna wanna chime in on, like, other elements of coming into, like, deciding on this, problem. So it's, like, probably my lab works on deep reinforcement learning. But, historically, deep meant, like, two or three or four layers. Not 1,000? When Kevin and Sean mentioned they wanted to try really deep networks, I was kinda skeptical it was gonna work. I've tried this before. It doesn't work. Other …
Get the full transcript (5,847 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 25-minute episode.
Get Latent Space summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Latent Space
🔬“We have foundation models for language, not for physics” — Anima Anandkumar, Bren Professor of Computing
Aug 26 · 83 min
How I AI
How a solo founder used Codex and ChatGPT to launch a fashion brand without engineers | Yana Welinder
Aug 17
More from Latent Space
Simulation: the new Scaling Law — Joon Sung Park, Simile AI
Aug 21 · 69 min
Lenny's Podcast
OpenAI’s Head of Design: This is the best time in history to be a designer | Ian Silber
Aug 16
More from Latent Space
We summarize every new episode. Want them in your inbox?
🔬“We have foundation models for language, not for physics” — Anima Anandkumar, Bren Professor of Computing
Simulation: the new Scaling Law — Joon Sung Park, Simile AI
🔬The BioAI Phase Shift - Matthew McPartlon & Neil Patil, Chai Discovery
The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten
Codex from 0 to 10M Users: Building ChatGPT Work — Akshay Nathan, OpenAI
Similar Episodes
Related episodes from other podcasts
How I AI
Aug 17
How a solo founder used Codex and ChatGPT to launch a fashion brand without engineers | Yana Welinder
Lenny's Podcast
Aug 16
OpenAI’s Head of Design: This is the best time in history to be a designer | Ian Silber
Acquired
Apr 13
Ferrari
The AI Breakdown
Mar 18
How to Use Agent Skills
Masters of Scale
Feb 26
Hailey Bieber, AI and fast launches: how e.l.f. Beauty is winning
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into Latent Space.
Every Monday, we deliver AI summaries of the latest episodes from Latent Space and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime