Jonathan Ross, Founder of Groq
Episode
71 min
Read time
3 min
Topics
Career Growth, Relationships, Startups
AI-Generated Summary
Key Takeaways
- ✓GPU + LPU Architecture: Groq's NVIDIA deal emerged from discovering that LPUs handle memory-throughput-constrained matrix operations while GPUs handle compute-constrained ones within LLM decoder layers. Combining both eliminates bottlenecks across all matrix multiplications, similar to using both 18-wheelers and delivery vans in a logistics network. This insight went from concept to wired funds in three weeks after Ross presented to Jensen Huang.
- ✓Minimal Constraints Leadership: Ross distilled Groq's entire company direction onto a challenge coin: 25 million tokens per second. The fewer constraints given to highly autonomous people, the more freedom they have to surprise you with solutions. Over-constraining kills innovation; under-constraining without a crisp objective causes paralysis. The goal must be simple enough to engrave on a coin yet open enough to allow creative problem-solving.
- ✓Intentional Leadership Phrasing: Replacing "should we do this?" with "I intend to do this" eliminates unnecessary pessimistic pushback while preserving genuine safety-critical objections. Drawn from David Marquet's *Turn the Ship Around*, this shift helped Ross stop being talked out of three consecutive LLM deployment opportunities. People volunteer objections only when something is genuinely wrong, not reflexively negative.
- ✓Hiring for Loss Bias: Groq's people spec screened for "book the win early" personalities — people who, upon hearing something is achievable, immediately treat not doing it as a loss. Ross observed that in architecture meetings, the same information produced opposite reactions: some heard "chip could be twice as fast next version," while loss-biased thinkers heard "chip is half as fast if we skip this now." Hire the latter for innovation-driven organizations.
- ✓Groq Bonds Survival Mechanism: Three weeks from insolvency, Ross rejected a layoff plan that would have eliminated critical compiler engineers needed to reach product viability. Instead, he offered salary-for-equity exchanges called Groq Bonds. Roughly 80% of employees participated, with approximately half reducing salaries to statutory minimums around $50,000–$60,000. Attrition dropped below 10%, possibly to 5%, extending runway by nearly two months.
What It Covers
Groq founder Jonathan Ross discusses the $20B NVIDIA partnership, how LPU and GPU architectures complement each other to accelerate AI inference, and lessons from a decade building Groq — covering leadership philosophy, hiring frameworks, near-bankruptcy survival, and why fast inference was contrarian before becoming consensus.
Key Questions Answered
- •GPU + LPU Architecture: Groq's NVIDIA deal emerged from discovering that LPUs handle memory-throughput-constrained matrix operations while GPUs handle compute-constrained ones within LLM decoder layers. Combining both eliminates bottlenecks across all matrix multiplications, similar to using both 18-wheelers and delivery vans in a logistics network. This insight went from concept to wired funds in three weeks after Ross presented to Jensen Huang.
- •Minimal Constraints Leadership: Ross distilled Groq's entire company direction onto a challenge coin: 25 million tokens per second. The fewer constraints given to highly autonomous people, the more freedom they have to surprise you with solutions. Over-constraining kills innovation; under-constraining without a crisp objective causes paralysis. The goal must be simple enough to engrave on a coin yet open enough to allow creative problem-solving.
- •Intentional Leadership Phrasing: Replacing "should we do this?" with "I intend to do this" eliminates unnecessary pessimistic pushback while preserving genuine safety-critical objections. Drawn from David Marquet's *Turn the Ship Around*, this shift helped Ross stop being talked out of three consecutive LLM deployment opportunities. People volunteer objections only when something is genuinely wrong, not reflexively negative.
- •Hiring for Loss Bias: Groq's people spec screened for "book the win early" personalities — people who, upon hearing something is achievable, immediately treat not doing it as a loss. Ross observed that in architecture meetings, the same information produced opposite reactions: some heard "chip could be twice as fast next version," while loss-biased thinkers heard "chip is half as fast if we skip this now." Hire the latter for innovation-driven organizations.
- •Groq Bonds Survival Mechanism: Three weeks from insolvency, Ross rejected a layoff plan that would have eliminated critical compiler engineers needed to reach product viability. Instead, he offered salary-for-equity exchanges called Groq Bonds. Roughly 80% of employees participated, with approximately half reducing salaries to statutory minimums around $50,000–$60,000. Attrition dropped below 10%, possibly to 5%, extending runway by nearly two months.
- •Reality Quotient Over IQ: Ross hired explicitly for reality quotient — the ability to identify the dominant game being played rather than optimizing a subordinate metric. Myspace maximized account signups; Facebook maximized monthly active users and won by playing the superior game. Founders must connect every team member's daily work to the dominant metric, then practice full-time change management by making necessary pivots feel like continuations, not disruptions.
Notable Moment
West Coast VCs passed on Groq repeatedly due to lemming dynamics — one firm's rejection cascaded across the entire ecosystem. The NVIDIA partnership, described as the largest deal in NVIDIA's history by nearly three times, was ultimately funded by East Coast crossover funds who ran independent analysis regardless of what other investors decided.
Episode Transcript
Let's start with this $20,000,000,000 rumored $20,000,000,000 partnership that you have with NVIDIA. Can you talk about the structure of the deal and how it came about? So the most interesting part about it is the call, where the idea was first floated was about three weeks before money was in the bank. That was a chance it moves fast. Of course. That's how you stay ahead. Yeah. So how'd it come about? So we had been working on, integrating GPU and LPUs together. And the best way to describe why this helps is if you're building out a logistics network for The United States, and I told you, you could have either 18 wheelers or, you know, vans for last mile delivery, which one would you pick? And the answer is both, right? And so GPUs and LPUs combined ended up giving better performance across the performance curves. We had implemented it, and we had gone to Jensen asking, could we buy about a 100,000 GPUs because we were gonna deploy them ourselves? And Jensen saw what we had done and thought maybe it would be better to make this available to all of their customers. You and I had this conversation, at NVIDIA GTC, and you were talking about the fact that these technologies are very complementary. Can you explain a little bit more about that? When you're processing an LLM token, what's happening is you're doing all of these different matrix multiplies. And some of them are more compute constrained, and some of them are more memory throughput constrained. And the ones that are more compute constrained, we put on the GPU, and the ones that are more memory throughput constrained, we put on the LPU. And the bottlenecks are all over the place. There's all sorts of different bottlenecks. There is no one strategy. There is no one perfect architecture. So the realization was you put these two things together and you defeat the bottlenecks across all of the different matmuls. There's another thing that I love that you said when we had this conversation that when AI is talking to other AI, speed is becoming more and more important. Yeah. I mean, a human can wait a second or two to get a response when they type a command into a computer. AI is just sitting there waiting because it produces these tokens so much fast it thinks so much faster. Now you bring in LPUs, and the speed is so much faster that it just becomes all about how do you move as fast as you can. AI is really good at using AI, and that's really what Agentic is, right? So humans benefit from using AI. So does AI. Just like you would do research about me before I show up on your show, AI is gonna kick off a job doing research on different tools it's gonna use while it's using this other tool. And so kicks it off to another AI. …
Get the full transcript (14,010 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 68-minute episode.
Get David Senra summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from David Senra
Travis Kalanick, Founder of Uber & Atoms
Aug 16 · 108 min
20VC (20 Minute VC)
20VC: OpenAI and Anthropic Will Build Their Own Chips | NVIDIA Will Be Worth $10TRN | How to Solve the Energy Required for AI... Nuclear | Why China is Behind the US in the Race for AGI with Jonathan Ross, Groq Founder
Sep 29
More from David Senra
Lulu Cheng Meservey, Founder of Rostra
Aug 12 · 77 min
Masters of Scale
Serena Williams on winning in business
Jul 28
Books, tools, and gear mentioned in this episode
SignalCast may earn commission on purchases via these links. As an Amazon Associate, SignalCast earns from qualifying purchases.
Books
Turn the Ship AroundRecommendedby David Marquet
“Intentional Leadership Phrasing: Replacing "should we do this?" with "I intend to do this" eliminates unnecessary pessimistic pushback while preserving genuine safety-critical objections. Drawn from David Marquet's *Turn the Ship Around*, this shift helped Ross stop being talked out of three consecutive LLM deployment opportunities.”
More from David Senra
We summarize every new episode. Want them in your inbox?
Travis Kalanick, Founder of Uber & Atoms
Lulu Cheng Meservey, Founder of Rostra
Michael Ovitz, Co-founder of CAA
Micky Malka, Founder of Ribbit Capital
David Heinemeier Hansson (DHH), Co-founder of 37signals
Similar Episodes
Related episodes from other podcasts
20VC (20 Minute VC)
Sep 29
20VC: OpenAI and Anthropic Will Build Their Own Chips | NVIDIA Will Be Worth $10TRN | How to Solve the Energy Required for AI... Nuclear | Why China is Behind the US in the Race for AGI with Jonathan Ross, Groq Founder
Masters of Scale
Jul 28
Serena Williams on winning in business
Business Breakdowns
Jul 27
Applied Intuition: A Billion Intelligent Machines - [Business Breakdowns, EP.248]
a16z Podcast
Jul 2
Outsmarting Uber: Why Bolt Wins in Europe
Masters of Scale
Jun 11
The future of EVs, with Rivian’s RJ Scaringe
Explore Related Topics
This podcast is featured in Best Business Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into David Senra.
Every Monday, we deliver AI summaries of the latest episodes from David Senra and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime