Skip to main content
David Senra

Jonathan Ross, Founder of Groq

71 min episode · 3 min read
·
Jonathan Ross

Episode

71 min

Read time

3 min

Topics

Career Growth, Relationships, Startups

AI-Generated Summary

Key Takeaways

  • GPU + LPU Architecture: Groq's NVIDIA deal emerged from discovering that LPUs handle memory-throughput-constrained matrix operations while GPUs handle compute-constrained ones within LLM decoder layers. Combining both eliminates bottlenecks across all matrix multiplications, similar to using both 18-wheelers and delivery vans in a logistics network. This insight went from concept to wired funds in three weeks after Ross presented to Jensen Huang.
  • Minimal Constraints Leadership: Ross distilled Groq's entire company direction onto a challenge coin: 25 million tokens per second. The fewer constraints given to highly autonomous people, the more freedom they have to surprise you with solutions. Over-constraining kills innovation; under-constraining without a crisp objective causes paralysis. The goal must be simple enough to engrave on a coin yet open enough to allow creative problem-solving.
  • Intentional Leadership Phrasing: Replacing "should we do this?" with "I intend to do this" eliminates unnecessary pessimistic pushback while preserving genuine safety-critical objections. Drawn from David Marquet's *Turn the Ship Around*, this shift helped Ross stop being talked out of three consecutive LLM deployment opportunities. People volunteer objections only when something is genuinely wrong, not reflexively negative.
  • Hiring for Loss Bias: Groq's people spec screened for "book the win early" personalities — people who, upon hearing something is achievable, immediately treat not doing it as a loss. Ross observed that in architecture meetings, the same information produced opposite reactions: some heard "chip could be twice as fast next version," while loss-biased thinkers heard "chip is half as fast if we skip this now." Hire the latter for innovation-driven organizations.
  • Groq Bonds Survival Mechanism: Three weeks from insolvency, Ross rejected a layoff plan that would have eliminated critical compiler engineers needed to reach product viability. Instead, he offered salary-for-equity exchanges called Groq Bonds. Roughly 80% of employees participated, with approximately half reducing salaries to statutory minimums around $50,000–$60,000. Attrition dropped below 10%, possibly to 5%, extending runway by nearly two months.

What It Covers

Groq founder Jonathan Ross discusses the $20B NVIDIA partnership, how LPU and GPU architectures complement each other to accelerate AI inference, and lessons from a decade building Groq — covering leadership philosophy, hiring frameworks, near-bankruptcy survival, and why fast inference was contrarian before becoming consensus.

Key Questions Answered

  • GPU + LPU Architecture: Groq's NVIDIA deal emerged from discovering that LPUs handle memory-throughput-constrained matrix operations while GPUs handle compute-constrained ones within LLM decoder layers. Combining both eliminates bottlenecks across all matrix multiplications, similar to using both 18-wheelers and delivery vans in a logistics network. This insight went from concept to wired funds in three weeks after Ross presented to Jensen Huang.
  • Minimal Constraints Leadership: Ross distilled Groq's entire company direction onto a challenge coin: 25 million tokens per second. The fewer constraints given to highly autonomous people, the more freedom they have to surprise you with solutions. Over-constraining kills innovation; under-constraining without a crisp objective causes paralysis. The goal must be simple enough to engrave on a coin yet open enough to allow creative problem-solving.
  • Intentional Leadership Phrasing: Replacing "should we do this?" with "I intend to do this" eliminates unnecessary pessimistic pushback while preserving genuine safety-critical objections. Drawn from David Marquet's *Turn the Ship Around*, this shift helped Ross stop being talked out of three consecutive LLM deployment opportunities. People volunteer objections only when something is genuinely wrong, not reflexively negative.
  • Hiring for Loss Bias: Groq's people spec screened for "book the win early" personalities — people who, upon hearing something is achievable, immediately treat not doing it as a loss. Ross observed that in architecture meetings, the same information produced opposite reactions: some heard "chip could be twice as fast next version," while loss-biased thinkers heard "chip is half as fast if we skip this now." Hire the latter for innovation-driven organizations.
  • Groq Bonds Survival Mechanism: Three weeks from insolvency, Ross rejected a layoff plan that would have eliminated critical compiler engineers needed to reach product viability. Instead, he offered salary-for-equity exchanges called Groq Bonds. Roughly 80% of employees participated, with approximately half reducing salaries to statutory minimums around $50,000–$60,000. Attrition dropped below 10%, possibly to 5%, extending runway by nearly two months.
  • Reality Quotient Over IQ: Ross hired explicitly for reality quotient — the ability to identify the dominant game being played rather than optimizing a subordinate metric. Myspace maximized account signups; Facebook maximized monthly active users and won by playing the superior game. Founders must connect every team member's daily work to the dominant metric, then practice full-time change management by making necessary pivots feel like continuations, not disruptions.

Notable Moment

West Coast VCs passed on Groq repeatedly due to lemming dynamics — one firm's rejection cascaded across the entire ecosystem. The NVIDIA partnership, described as the largest deal in NVIDIA's history by nearly three times, was ultimately funded by East Coast crossover funds who ran independent analysis regardless of what other investors decided.

Know someone who'd find this useful?

Episode Transcript

Let's start with this $20,000,000,000 rumored $20,000,000,000 partnership that you have with NVIDIA. Can you talk about the structure of the deal and how it came about? So the most interesting part about it is the call, where the idea was first floated was about three weeks before money was in the bank. That was a chance it moves fast. Of course. That's how you stay ahead. Yeah. So how'd it come about? So we had been working on, integrating GPU and LPUs together. And the best way to describe why this helps is if you're building out a logistics network for The United States, and I told you, you could have either 18 wheelers or, you know, vans for last mile delivery, which one would you pick? And the answer is both, right? And so GPUs and LPUs combined ended up giving better performance across the performance curves. We had implemented it, and we had gone to Jensen asking, could we buy about a 100,000 GPUs because we were gonna deploy them ourselves? And Jensen saw what we had done and thought maybe it would be better to make this available to all of their customers. You and I had this conversation, at NVIDIA GTC, and you were talking about the fact that these technologies are very complementary. Can you explain a little bit more about that? When you're processing an LLM token, what's happening is you're doing all of these different matrix multiplies. And some of them are more compute constrained, and some of them are more memory throughput constrained. And the ones that are more compute constrained, we put on the GPU, and the ones that are more memory throughput constrained, we put on the LPU. And the bottlenecks are all over the place. There's all sorts of different bottlenecks. There is no one strategy. There is no one perfect architecture. So the realization was you put these two things together and you defeat the bottlenecks across all of the different matmuls. There's another thing that I love that you said when we had this conversation that when AI is talking to other AI, speed is becoming more and more important. Yeah. I mean, a human can wait a second or two to get a response when they type a command into a computer. AI is just sitting there waiting because it produces these tokens so much fast it thinks so much faster. Now you bring in LPUs, and the speed is so much faster that it just becomes all about how do you move as fast as you can. AI is really good at using AI, and that's really what Agentic is, right? So humans benefit from using AI. So does AI. Just like you would do research about me before I show up on your show, AI is gonna kick off a job doing research on different tools it's gonna use while it's using this other tool. And so kicks it off to another AI. …

Get the full transcript (14,010 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all David Senra transcripts →

You just read a 3-minute summary of a 68-minute episode.

Get David Senra summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

Books, tools, and gear mentioned in this episode

SignalCast may earn commission on purchases via these links. As an Amazon Associate, SignalCast earns from qualifying purchases.

Books

  • by David Marquet

    Intentional Leadership Phrasing: Replacing "should we do this?" with "I intend to do this" eliminates unnecessary pessimistic pushback while preserving genuine safety-critical objections. Drawn from David Marquet's *Turn the Ship Around*, this shift helped Ross stop being talked out of three consecutive LLM deployment opportunities.

More from David Senra

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best Business Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into David Senra.

Every Monday, we deliver AI summaries of the latest episodes from David Senra and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime