Skip to main content
Invest Like the Best with Patrick O'Shaughnessy

Sergey Levine - Building LLMs for the Physical World - [Invest Like the Best, EP.465]

66 min episode · 3 min read
·
Sergey Levine

Episode

66 min

Read time

3 min

Topics

Startups, Crypto & Web3, Psychology & Behavior

AI-Generated Summary

Key Takeaways

  • Generality over specialization: Building one robotic foundation model that handles all tasks and embodiments outperforms narrow specialists long-term, mirroring how LLMs defeated domain-specific NLP tools like machine translation systems. The key mechanism: broad data enables physical world understanding, which transfers across applications far more efficiently than rebuilding task-specific pipelines from scratch for each new robot deployment.
  • Chain-of-thought unlocks robotic common sense: Physical Intelligence's models use intermediate semantic reasoning before acting — a robot told to "clean the kitchen" first identifies which object to pick up, then moves. This chain-of-thought step activates web-scale pre-training knowledge to handle edge cases, shifting the bottleneck from low-level motor control to mid-level scene interpretation, which can be supervised with language alone.
  • Coaching replaces teleoperation data: Six months ago, Physical Intelligence discovered that labeling robot experiences with high-level semantic commands — without adding any new low-level action demonstrations — improved kitchen generalization. This means operators can improve robot performance simply by verbally coaching the system, dramatically reducing the cost and complexity of expanding a robot's capability to new environments.
  • Reinforcement learning enables superhuman throughput: After demonstrating a task via teleoperation, robots can practice autonomously and remove human-paced pauses. In cable-plugging tasks, the robot identified and eliminated all hesitation points, executing the task significantly faster than human operators. Reinforcement learning is the general mechanism; simpler speed-optimization tricks also work for throughput gains without full RL pipelines.
  • Hardware costs dropped 40x in a decade: Robot arm costs fell from roughly $400,000 for a PR2 in 2014 to approximately $3,000–$4,000 per arm today. This cost collapse, enabled by combining cheaper hardware with learning-based control that tolerates mechanical imprecision, makes broad experimentation practical. Traditional industrial control methods required high-precision hardware; foundation model approaches compensate for mechanical variability through learned adaptation.

What It Covers

Sergey Levine, cofounder of Physical Intelligence, explains why building general-purpose robotic foundation models — systems that control any robot for any task — is more tractable than narrow domain-specific approaches, drawing direct parallels to how large language models outcompeted specialized NLP systems by leveraging broad, weakly-labeled data at scale.

Key Questions Answered

  • Generality over specialization: Building one robotic foundation model that handles all tasks and embodiments outperforms narrow specialists long-term, mirroring how LLMs defeated domain-specific NLP tools like machine translation systems. The key mechanism: broad data enables physical world understanding, which transfers across applications far more efficiently than rebuilding task-specific pipelines from scratch for each new robot deployment.
  • Chain-of-thought unlocks robotic common sense: Physical Intelligence's models use intermediate semantic reasoning before acting — a robot told to "clean the kitchen" first identifies which object to pick up, then moves. This chain-of-thought step activates web-scale pre-training knowledge to handle edge cases, shifting the bottleneck from low-level motor control to mid-level scene interpretation, which can be supervised with language alone.
  • Coaching replaces teleoperation data: Six months ago, Physical Intelligence discovered that labeling robot experiences with high-level semantic commands — without adding any new low-level action demonstrations — improved kitchen generalization. This means operators can improve robot performance simply by verbally coaching the system, dramatically reducing the cost and complexity of expanding a robot's capability to new environments.
  • Reinforcement learning enables superhuman throughput: After demonstrating a task via teleoperation, robots can practice autonomously and remove human-paced pauses. In cable-plugging tasks, the robot identified and eliminated all hesitation points, executing the task significantly faster than human operators. Reinforcement learning is the general mechanism; simpler speed-optimization tricks also work for throughput gains without full RL pipelines.
  • Hardware costs dropped 40x in a decade: Robot arm costs fell from roughly $400,000 for a PR2 in 2014 to approximately $3,000–$4,000 per arm today. This cost collapse, enabled by combining cheaper hardware with learning-based control that tolerates mechanical imprecision, makes broad experimentation practical. Traditional industrial control methods required high-precision hardware; foundation model approaches compensate for mechanical variability through learned adaptation.
  • Moravec's Paradox defines the hardest remaining tasks: Tasks humans perform effortlessly — interpersonal physical assistance, elderly care, infant care — will be the last robotic capabilities achieved, not because of motor complexity but because humans are evolutionarily optimized for them. Robots will handle well-defined chaotic environments like hotel rooms or restaurant kitchens before mastering open-ended human-interaction tasks where stakes are high and edge cases are unbounded.

Notable Moment

Levine describes running the "Robot Olympics" — a blogger's list of mundane tasks no robot could do, like using a plastic bag to pick up dog waste or washing a greasy pan — as an internal stress test of their task-onboarding pipeline. The system completed nearly every task without any task-specific development, demonstrating generalization in practice.

Know someone who'd find this useful?

Episode Transcript

Most software companies try to maximize your time on their app to juice engagement. Ramp does the exact opposite. Ramp understands that no one wants to spend hours chasing receipts, reviewing expense reports, and checking for policy violations. So they built their tools to give that time back, using AI to automate 85% of expense reviews with 99% accuracy. And since Ramp saves companies 5%, it's no wonder that Shopify runs on Ramp, Stripe runs on Ramp, and my business does too. To see what happens when you eliminate the busy work, check out ramp.com/invest. Every investor should know about Rogo because Rogo AI's platform is not just another generic chatbot. Instead, it was designed to support how Wall Street bankers and investors actually work from sourcing diligence and modeling to turning analysis into deliverables. For me, three key things differentiate Rogo. First, it connects directly to your system, so it can work with your actual data. Second, it understands your workflows, how work really happens across a deal or an investment. And third, it runs end to end and produces real outputs the way the best people do, auditable spreadsheets, investment memos, diligence materials, and slide decks that match your standards. This all comes from the fact that Rogo is built by finance professionals for finance professionals, and it's already being adopted by some of the most demanding institutions in the world. To learn more, visit rogo.ai/invest. OpenAI, Cursor, Anthropic, Perplexity, and Vercel all have something in common. They all use Work OS. And here's why. To achieve enterprise adoption at scale, you have to deliver on core capabilities like SSO, SCIM, RBAC, and audit logs. That's where Work OS comes in. Instead of spending months building these mission critical capabilities yourself, you can just use WorkOS APIs to gain all of them on day zero. That's why so many of the top AI teams you hear about already run on WorkOS. WorkOS is the fastest way to become enterprise ready and stay focused on what matters most, your product. Visit workos.com to get started. Hello and welcome, everyone. I'm Patrick O'Shaughnessy, and this is Invest Like the Best. This show is an open ended exploration of markets, ideas, stories, and strategies that will help you better invest both your time and your money. If you enjoy these conversations and wanna go deeper, check out Colossus, our quarterly publication with in-depth profiles of the people shaping business and investing. You can find Colossus along with all of our podcasts at colossus.com. Patrick O'Shaughnessy is the CEO of Positive Sum. All opinions expressed by Patrick and podcast guests are solely their own opinions and do not reflect the opinion of Positive Sum. This podcast is for informational purposes only and should not be relied upon as a basis for investment decisions. Clients of positive sum may maintain positions in the securities discussed in this podcast. To learn more, visit psum.vc. My guest today is Sergei Levine, one of the cofounders …

Get the full transcript (14,227 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all Invest Like the Best with Patrick O'Shaughnessy transcripts →

You just read a 3-minute summary of a 63-minute episode.

Get Invest Like the Best with Patrick O'Shaughnessy summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

More from Invest Like the Best with Patrick O'Shaughnessy

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best Investing Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into Invest Like the Best with Patrick O'Shaughnessy.

Every Monday, we deliver AI summaries of the latest episodes from Invest Like the Best with Patrick O'Shaughnessy and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime