Skip to main content
The TWIML AI Podcast

Proactive Agents for the Web with Devi Parikh - #756

56 min episode · 2 min read
·
Devi Parikh

Episode

56 min

Read time

2 min

Topics

Startups, Leadership, Artificial Intelligence

AI-Generated Summary

Key Takeaways

  • Visual-based browser navigation: Training models on website screenshots rather than DOM information proves more reliable and generalizable across different sites, solving challenges like date pickers that plagued DOM-based approaches with constant edge cases requiring site-specific solutions.
  • Scouts architecture combines APIs and browser automation: The system uses 80-90 MCP servers for structured data access but spins up remote browsers with custom-trained navigator models for information behind forms, optimizing for coverage first then precision in user-facing reports.
  • Post-training progression maximizes model capability: Utori trains QwQ models through supervised fine-tuning, then rejection sampling, then reinforcement learning to achieve reliable browser automation while keeping costs lower than using third-party API providers for their production workloads.
  • Background agents require hierarchical tool management: When orchestrating 80-90 tools, reliability breaks down if all tools are available simultaneously. Sub-agents with access to specific tool subsets enable scalable multi-agent workflows that adapt based on real-time web information.

What It Covers

Devi Parikh, co-founder of Utori, explains how AI browser agents will replace manual web interactions through proactive monitoring and automation, starting with Scouts, their product that monitors websites for user-specified information changes.

Key Questions Answered

  • Visual-based browser navigation: Training models on website screenshots rather than DOM information proves more reliable and generalizable across different sites, solving challenges like date pickers that plagued DOM-based approaches with constant edge cases requiring site-specific solutions.
  • Scouts architecture combines APIs and browser automation: The system uses 80-90 MCP servers for structured data access but spins up remote browsers with custom-trained navigator models for information behind forms, optimizing for coverage first then precision in user-facing reports.
  • Post-training progression maximizes model capability: Utori trains QwQ models through supervised fine-tuning, then rejection sampling, then reinforcement learning to achieve reliable browser automation while keeping costs lower than using third-party API providers for their production workloads.
  • Background agents require hierarchical tool management: When orchestrating 80-90 tools, reliability breaks down if all tools are available simultaneously. Sub-agents with access to specific tool subsets enable scalable multi-agent workflows that adapt based on real-time web information.

Notable Moment

Parikh reveals that despite initial assumptions, consuming web pages visually like humans rather than parsing underlying code proved essential for building reliable browser agents, as identical-looking pages often have completely different underlying structures.

Know someone who'd find this useful?

Episode Transcript

We will no longer be interacting with the web in the same way that we do right now. We won't be clicking buttons, fiddling with forms on websites and browsers. We'll be interacting with the web one level higher in the abstraction where we're describing what needs to be done. Maybe our assistant is proactively noticing what needs to be done and sort of agents in the background are starting to execute these workflows, on the web on on your behalf. Alright, everyone. Welcome to another episode of the TwinWell AI podcast. I am your host, Sam Charrington. Today, I'm joined by Davey Parikh. Davey is cofounder and co CEO of Uturi. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Davey, it has been a while since we caught up last. Welcome back to the podcast. Thank you. Thank you for having me again. Yeah. Five years later, not much has happened at all. Right? In in some ways, a lot has happened, but in some ways, I'm like, wow. It's been five years. So yeah. I know. I know. I know. So we're going to be talking a bit about AI browsers and browser use agents, and what you're building at YouTory, but I'd love to have you take a few minutes and catch us all up on what you've been up to recently. Yeah. Yeah. And I can I can go, a little bit further back than than five years, just to talk about my background a little bit? So I've been working in AI for about twenty years now. Originally, my PhD thesis was in computer vision. And then over time, I got interested in seeing if we can find ways in which people can interact with these systems more naturally. And that's how I move towards multimodal problems at the intersection of vision and language. So things like given an image, can you describe it in a sentence? Can you answer questions about it? Can you have a conversation going back and forth, about the content of an image? And this was back in 2014 or so. So it was after that initial excitement of deep learning models, but it was starting to feel like, wait. These models are doing something. Stuff is actually starting to work. But it was well before all of the current excitement around Gen AI and other lens and so on. So these models weren't really as good, as they are today. And yeah. So it was it was kind of fun to tinker on the boundaries of of what's possible. And then I started getting interested in seeing if we can find ways in which we can use AI as a tool for creative expression. And that's how I got involved with generative models for images and videos and music and and other modalities, like that. I was in academia for a while, faculty at Virginia Tech …

Get the full transcript (10,042 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all The TWIML AI Podcast transcripts →

You just read a 3-minute summary of a 53-minute episode.

Get The TWIML AI Podcast summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

Books, tools, and gear mentioned in this episode

SignalCast may earn commission on purchases via these links. As an Amazon Associate, SignalCast earns from qualifying purchases.

Tools

  • Utori trains QwQ models through supervised fine-tuning, then rejection sampling, then reinforcement learning to achieve reliable browser automation
  • The system uses 80-90 MCP servers for structured data access but spins up remote browsers with custom-trained navigator models

Products

  • ScoutsBy guest

    by Utori

    Devi Parikh, co-founder of Utori, explains how AI browser agents will replace manual web interactions through proactive monitoring and automation, starting with Scouts, their product that monitors websites for user-specified information changes.

company

  • UtoriBy guest
    Devi Parikh, co-founder of Utori, explains how AI browser agents will replace manual web interactions through proactive monitoring and automation

More from The TWIML AI Podcast

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into The TWIML AI Podcast.

Every Monday, we deliver AI summaries of the latest episodes from The TWIML AI Podcast and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime