Skip to main content
Machine Learning Street Talk

"Vibe Coding is a Slot Machine" - Jeremy Howard

86 min episode · 3 min read
·
Jeremy Howard

Episode

86 min

Read time

3 min

Topics

Productivity, Design & UX, Artificial Intelligence

AI-Generated Summary

Key Takeaways

  • Vibe Coding Productivity Gap: A study Howard cites shows only a tiny measurable uptick in software actually shipped despite widespread AI coding adoption. The slot machine analogy applies precisely: users craft prompts, adjust MCPs, and pull the lever repeatedly, experiencing stochastic wins that feel like skill but mask the absence of genuine output growth. No organization is demonstrably producing 50x more high-quality software.
  • Coding vs. Software Engineering: LLMs perform style transfer between training data points, which constitutes coding but not software engineering. Designing novel abstractions, identifying correct component boundaries, and composing systems that have never existed before requires moving outside training distribution — something LLMs demonstrably cannot do. Howard cites Anthropic's browser and the AI-generated C compiler as empirical examples of sophisticated copying, not original design.
  • Understanding Debt in Organizations: When teams delegate cognitive tasks to LLMs, organizational knowledge erodes. Howard frames this using the call center analogy: even seemingly routine roles generate edge cases that propagate upward and keep institutional knowledge adaptive. Automating those roles removes the feedback loop that makes organizations evolvable, cutting the legs off future adaptability without any immediate visible cost.
  • The Desirable Difficulty Principle: An Anthropic study found that developers using AI coding tools experienced so little friction they retained almost nothing. Howard connects this to Ebbinghaus and spaced repetition research: memories and skills only form under effortful conditions. He recommends organizations explicitly prioritize employee learning slope over output intercept, measuring how fast individuals grow rather than how many pull requests they close.
  • Interactive Notebook Environments Outperform Terminal-Based AI: Howard's nbdev framework embeds tests, documentation, implementation, and examples inside Jupyter notebooks, enabling CI integration without sacrificing exploratory feedback loops. Placing AI inside a live Python interpreter rather than a bash terminal gives both human and model richer real-time feedback. Howard reports feeling energized after notebook sessions versus drained after 14-hour Claude Code marathons.

What It Covers

Deep learning pioneer Jeremy Howard joins Machine Learning Street Talk to argue that vibe coding functions like a slot machine, creating an illusion of control while eroding genuine software engineering competence. He draws on ULMFiT's origins, transfer learning history, and his own Claude Code experiments to distinguish coding from software engineering, warning that organizations betting on AI productivity gains face measurable, documented risks.

Key Questions Answered

  • Vibe Coding Productivity Gap: A study Howard cites shows only a tiny measurable uptick in software actually shipped despite widespread AI coding adoption. The slot machine analogy applies precisely: users craft prompts, adjust MCPs, and pull the lever repeatedly, experiencing stochastic wins that feel like skill but mask the absence of genuine output growth. No organization is demonstrably producing 50x more high-quality software.
  • Coding vs. Software Engineering: LLMs perform style transfer between training data points, which constitutes coding but not software engineering. Designing novel abstractions, identifying correct component boundaries, and composing systems that have never existed before requires moving outside training distribution — something LLMs demonstrably cannot do. Howard cites Anthropic's browser and the AI-generated C compiler as empirical examples of sophisticated copying, not original design.
  • Understanding Debt in Organizations: When teams delegate cognitive tasks to LLMs, organizational knowledge erodes. Howard frames this using the call center analogy: even seemingly routine roles generate edge cases that propagate upward and keep institutional knowledge adaptive. Automating those roles removes the feedback loop that makes organizations evolvable, cutting the legs off future adaptability without any immediate visible cost.
  • The Desirable Difficulty Principle: An Anthropic study found that developers using AI coding tools experienced so little friction they retained almost nothing. Howard connects this to Ebbinghaus and spaced repetition research: memories and skills only form under effortful conditions. He recommends organizations explicitly prioritize employee learning slope over output intercept, measuring how fast individuals grow rather than how many pull requests they close.
  • Interactive Notebook Environments Outperform Terminal-Based AI: Howard's nbdev framework embeds tests, documentation, implementation, and examples inside Jupyter notebooks, enabling CI integration without sacrificing exploratory feedback loops. Placing AI inside a live Python interpreter rather than a bash terminal gives both human and model richer real-time feedback. Howard reports feeling energized after notebook sessions versus drained after 14-hour Claude Code marathons.
  • ULMFiT's Three-Stage Architecture Predicted Modern LLM Training: Howard's 2018 model used general-purpose Wikipedia pretraining on an AWD-LSTM with five regularization types, followed by domain-specific supervised fine-tuning, then classifier fine-tuning — matching today's pretraining, mid-training, and post-training pipeline. Discriminative learning rates assigned per layer and sequential unfreezing from last to first layer were the key fine-tuning innovations, achieving state-of-the-art sentiment classification in minutes on a single gaming GPU.

Notable Moment

Howard spent two weeks using GPT-4 and Codex to fix a crashing IPython kernel upgrade, ultimately producing what he believes is the only working implementation of the new protocol. He then faced a genuine dilemma: whether to build a company product on code nobody, including him, actually understands.

Know someone who'd find this useful?

Episode Transcript

This episode is brought to you by Indeed. Stop waiting around for the perfect candidate. Instead, use Indeed sponsored jobs to find the right people with the right skills fast. It's a simple way to make sure your listing is the first candidate see. According to Indeed data, sponsored jobs have four times more applicants than non sponsored jobs. So go build your dream team today with Indeed. Get a $75 sponsored job credit at indeed.com/podcast. Terms and conditions apply. It it it literally disgusts me. Like, I literally think it's it's inhumane. My mission remains the same as it has been for, like, twenty years, which is to stop people working like this. Jeremy Howard, a deep learning pioneer, a Kaggle grandmaster. He is a huge advocate for actually understanding what we are building through an interactive loop, a notebook, a REPL, the act of poking at a problem until it pushes back. He argues this is where the real insight happens. And the funny thing is they're both right. LLMs cosplay understanding things. Like, they pretend to understand things. No one's actually creating 50 times more high quality software than they were before. So we've actually just done a study of this, and there's a tiny uptick, tiny uptick in what people are actually shipping. The thing about AI based coding is that it's like a slot machine and that you you have an illusion of control, you know, you could get to craft your prompt and your list of MCPs and your skills and whatever and then but in the end, you pull the lever. Right? Here's a piece of code that no one understands. Yep. And am I going to bet my company's product on it? And I the answer is I don't know because, like, I I don't I don't like, I I don't know what to do now because no one's, like, been in this situation. They're they're really bad at software engineering. And then I think that's possibly always gonna be true. The idea that a human can do a lot more with a computer when the human can, like, manipulate the objects in inside that computer in real time, study them, and move them around, and combine them together. Whoever you listen to, you know, whether it be Feynman or whatever, like, you you always hear from the great scientists how they build deeper intuition by by building mental models, which they get over time by interacting with the things that they're learning about. A machine could kind of build an effective hierarchy of abstractions about what the world is and how it works entirely through looking at the statistical correlations of a huge corpus of text using a deep learning model. That was my premise. This video is brought to you by NVIDIA GTC. It's running March 16 until the nineteenth in San Jose and streaming free online. The key topics this year are agentic AI and reasoning, high performance inference …

Get the full transcript (14,845 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all Machine Learning Street Talk transcripts →

You just read a 3-minute summary of a 83-minute episode.

Get Machine Learning Street Talk summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

Books, tools, and gear mentioned in this episode

SignalCast may earn commission on purchases via these links. As an Amazon Associate, SignalCast earns from qualifying purchases.

Books

  • ULMFiTBy guest

    by Jeremy Howard

    He draws on ULMFiT's origins, transfer learning history, and his own Claude Code experiments to distinguish coding from software engineering. Howard's 2018 model used general-purpose Wikipedia pretraining on an AWD-LSTM with five regularization types, followed by domain-specific supervised fine-tuning.

Tools

  • by Anthropic

    his own Claude Code experiments to distinguish coding from software engineering. Howard reports feeling energized after notebook sessions versus drained after 14-hour Claude Code marathons.
  • nbdevRecommendedBy guest

    by Jeremy Howard

    Howard's nbdev framework embeds tests, documentation, implementation, and examples inside Jupyter notebooks, enabling CI integration without sacrificing exploratory feedback loops.
  • Howard's nbdev framework embeds tests, documentation, implementation, and examples inside Jupyter notebooks, enabling CI integration without sacrificing exploratory feedback loops.
  • by OpenAI

    Howard spent two weeks using GPT-4 and Codex to fix a crashing IPython kernel upgrade, ultimately producing what he believes is the only working implementation of the new protocol.
  • by OpenAI

    Howard spent two weeks using GPT-4 and Codex to fix a crashing IPython kernel upgrade, ultimately producing what he believes is the only working implementation of the new protocol.
  • Howard spent two weeks using GPT-4 and Codex to fix a crashing IPython kernel upgrade, ultimately producing what he believes is the only working implementation of the new protocol.

Gear

  • Howard's 2018 model used general-purpose Wikipedia pretraining on an AWD-LSTM with five regularization types, followed by domain-specific supervised fine-tuning.

More from Machine Learning Street Talk

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's AI & Machine Learning Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into Machine Learning Street Talk.

Every Monday, we deliver AI summaries of the latest episodes from Machine Learning Street Talk and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime