Skip to main content
Software Engineering Daily

The Gap Between AI Spending and AI Value

54 min episode · 2 min read
·
Emily Hsu,Kevin Ball

Episode

54 min

Read time

2 min

Topics

Relationships, Investing, Fundraising & VC

AI-Generated Summary

Key Takeaways

  • Three-layer failure model: Enterprise AI breaks down across three distinct layers: foundation models lack enterprise-specific benchmarks, legacy systems contain fragmented multimodal data that agents cannot reliably ingest, and leadership fails to define measurable ROI targets before deployment. Addressing all three simultaneously, rather than treating them as sequential problems, separates successful deployments from stalled pilots.
  • Benchmark misalignment: Frontier model development teams have minimal exposure to enterprise workflows, so evaluation benchmarks prioritize math, code, and conversational ability over professional operational requirements. Enterprises should build their own domain-specific evaluation datasets covering happy paths, edge cases, hard negatives, and ambiguous inputs before selecting or fine-tuning any foundation model for production use.
  • Data foundation prerequisite: Scale AI's research on the 6% of successful enterprises identifies data infrastructure as the strongest predictor of deployment success. Before deploying agents, organizations must resolve entity resolution across fragmented tables, establish data access governance policies, and implement feedback collection pipelines — using AI-assisted entity resolution tools to accelerate the normalization process itself.
  • Co-development over buy-or-build: The 6% of successful enterprises combine internal domain expertise with external AI specialists rather than choosing purely between purchasing a product or building internally. Internal teams contribute workflow knowledge and acceptable-outcome definitions; external partners contribute model evaluation best practices, agent architecture patterns, and known failure modes from cross-industry deployments.
  • Pilot selection as leverage point: The highest-leverage action for any engineer or line manager is identifying the correct first pilot — one that is genuinely valuable to business metrics, has accessible and clean enough data to execute, and falls within the team's technical reach. A successful pilot-to-production launch compounds organizational trust and resources for subsequent deployments.

What It Covers

Emily Hsu, Head of Enterprise AI at Scale AI, examines why only 6% of large enterprises successfully deploy AI at scale. The episode breaks down three failure layers — model capability gaps, fragmented data infrastructure, and organizational change management — and identifies patterns shared by the companies succeeding.

Key Questions Answered

  • Three-layer failure model: Enterprise AI breaks down across three distinct layers: foundation models lack enterprise-specific benchmarks, legacy systems contain fragmented multimodal data that agents cannot reliably ingest, and leadership fails to define measurable ROI targets before deployment. Addressing all three simultaneously, rather than treating them as sequential problems, separates successful deployments from stalled pilots.
  • Benchmark misalignment: Frontier model development teams have minimal exposure to enterprise workflows, so evaluation benchmarks prioritize math, code, and conversational ability over professional operational requirements. Enterprises should build their own domain-specific evaluation datasets covering happy paths, edge cases, hard negatives, and ambiguous inputs before selecting or fine-tuning any foundation model for production use.
  • Data foundation prerequisite: Scale AI's research on the 6% of successful enterprises identifies data infrastructure as the strongest predictor of deployment success. Before deploying agents, organizations must resolve entity resolution across fragmented tables, establish data access governance policies, and implement feedback collection pipelines — using AI-assisted entity resolution tools to accelerate the normalization process itself.
  • Co-development over buy-or-build: The 6% of successful enterprises combine internal domain expertise with external AI specialists rather than choosing purely between purchasing a product or building internally. Internal teams contribute workflow knowledge and acceptable-outcome definitions; external partners contribute model evaluation best practices, agent architecture patterns, and known failure modes from cross-industry deployments.
  • Pilot selection as leverage point: The highest-leverage action for any engineer or line manager is identifying the correct first pilot — one that is genuinely valuable to business metrics, has accessible and clean enough data to execute, and falls within the team's technical reach. A successful pilot-to-production launch compounds organizational trust and resources for subsequent deployments.

Notable Moment

When discussing individual AI tool adoption, Hsu points out that employees using tools like ChatGPT or coding assistants are inadvertently leaking enterprise IP — not just proprietary data, but the implicit judgment and institutional knowledge embedded in how they phrase follow-up questions, which never gets captured organizationally.

Know someone who'd find this useful?

Episode Transcript

It is widely reported that a gap has emerged between enterprise spending on AI and the durable value captured from that spend. Individual employees have enthusiastically adopted coding assistance and chatbots, yet those gains do not seem to be transforming businesses at an organizational level. One of the most important questions in the tech industry today is understanding why AI is not yet delivering returns that match the investment, And what separates the small number of enterprises succeeding from the many that are not? Scale AI is known for supplying the human labeled data behind many frontier models. It now also builds AI applications and agents for large enterprises. That combination of working alongside Frontier Labs and inside enterprise deployments gives the company a rare view of why enterprise AI may be stalling. Emily Hsu is the head of enterprise AI at Scale AI and previously spent over a decade at Google. In this episode, Emily joins Kevin Ball to discuss the three layers where enterprise AI breaks down, why frontier model benchmarks miss what enterprises actually need, the data foundation problem, how the most successful companies combine internal domain expertise with outside AI specialists, and more. Kevin Ball or Kate Ball is the vice president of engineering at Mento and an independent coach for engineers and engineering leaders. He cofounded and served as CTO for two companies, founded the San Diego JavaScript meetup, and organizes the AI in action discussion group through Latent Space. Check out the show notes to follow Keball on Twitter or LinkedIn, or visit his website, keball.llc. Emily, welcome to the show. Thank you so much for inviting me, Calvin. It's my pleasure to be here. Yeah. I'm really excited to dig in with you on this. Let's start with a bit of your background. So can you kind of give us a little bit of your career to date and how you ended up at Scale? Yeah. I absolutely love to share that. A lot of people were saying, like, my career is kind of a reverse of many people. So currently, I'm the head of enterprise AI and Scale AI. It's kind of a company, a startup company, but it's not, like, as small as a startup anymore that people kind of note from the perspective supplying the data to many of the frontier labs and being active participation to the Gen AI large launch model to where we are. So, it's very interesting work from both understanding the customer requirements on enterprise side as well as driving the innovations to research for pushing enterprise side to land with revenues. Then before that, I was actually worked for a really large company, Google, one of the hyperscalers and also the Frontier Labs, two in one. I worked there for eleven years being very hands on engineers. I'm also early member of Google Brain team. So I was a researcher, hands on model training, publishing papers, and I was engineering lead for several of …

Get the full transcript (9,018 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all Software Engineering Daily transcripts →

You just read a 3-minute summary of a 51-minute episode.

Get Software Engineering Daily summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

More from Software Engineering Daily

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best Cybersecurity Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Investing & Markets Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into Software Engineering Daily.

Every Monday, we deliver AI summaries of the latest episodes from Software Engineering Daily and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime