Skip to main content
The TWIML AI Podcast

High-Efficiency Diffusion Models for On-Device Image Generation and Editing with Hung Bui - #753

52 min episode · 2 min read
·
Hung Bui

Episode

52 min

Read time

2 min

Topics

Productivity, Startups, Fundraising & VC

AI-Generated Summary

Key Takeaways

  • Model Size Reduction: A sub-4-billion parameter Vietnamese language model outperformed the original 7-billion parameter version by iterating over the same dataset multiple times during training and applying minor optimization adjustments, proving smaller can be better with proper training techniques.
  • One-Step Diffusion: Swift Brush eliminates the typical 50-100 denoising steps in diffusion models by distilling multi-step knowledge into a single-step student network, achieving image generation in under 0.25 seconds while maintaining quality scores equal to or better than the original teacher model.
  • Image Editing Architecture: Swift Edit enables one-step image editing by training an inverted network that converts images to noise, then applies the one-step generation model. Training uses both real data and synthetic data generated by the efficient one-step model, creating highly intuitive loss functions.
  • Test-Time Scaling Advantage: Small models with inference-time scaling can outperform significantly larger models on specific tasks like math, making them viable for on-device deployment despite the increased compute requirements. This transforms the constraint of limited device resources into an opportunity for efficient specialized performance.

What It Covers

Hung Bui explains how VinAI Research achieved efficient on-device AI by training smaller models that match larger model performance, developing one-step diffusion for real-time image generation, and building Vietnam's top AI research lab.

Key Questions Answered

  • Model Size Reduction: A sub-4-billion parameter Vietnamese language model outperformed the original 7-billion parameter version by iterating over the same dataset multiple times during training and applying minor optimization adjustments, proving smaller can be better with proper training techniques.
  • One-Step Diffusion: Swift Brush eliminates the typical 50-100 denoising steps in diffusion models by distilling multi-step knowledge into a single-step student network, achieving image generation in under 0.25 seconds while maintaining quality scores equal to or better than the original teacher model.
  • Image Editing Architecture: Swift Edit enables one-step image editing by training an inverted network that converts images to noise, then applies the one-step generation model. Training uses both real data and synthetic data generated by the efficient one-step model, creating highly intuitive loss functions.
  • Test-Time Scaling Advantage: Small models with inference-time scaling can outperform significantly larger models on specific tasks like math, making them viable for on-device deployment despite the increased compute requirements. This transforms the constraint of limited device resources into an opportunity for efficient specialized performance.

Notable Moment

Vietnamese users complained that even the 7-billion parameter open-weight model was too large for their GPUs, prompting the team to halve the model size. The resulting sub-4-billion parameter version unexpectedly performed better than the original larger model.

Know someone who'd find this useful?

Episode Transcript

So we're realizing this open weight models, 7,000,000,000 parameter model. We're still getting complaint from the community in Vietnam that, oh, this model is too big. Right? We can't fit in in on on on our GPU. Right? And we said, okay. Fine. You know? Let let let let us go, you know, one more step. Try to, like, reduce the the number of parameters. Try to have the size of the model to, less than 4,000,000,000 parameter. And with the couple of improvement over the way we train it, we noticed that this model, less than 4,000,000,000 parameters, actually perform, you know, even better than the 7,000,000,000 parameter model. Alright, everyone. Welcome to another episode of the TwiML AI podcast. I am your host, Sam Charrington. Today, I'm joined by Hung Bui. Hung recently joined Qualcomm as VP of technology through the recent acquisition of VinAI Research, which ranked in the world's top 25 industrial AI labs based on research output in top conferences like ICML and NeurIPS. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Hung, welcome to the podcast. Thank you so much, Sam, and it's my great pleasure to be here. So we've got a bunch of really interesting topics to dig into, including your research into topics like diffusion models, image generation and editing, and more, and, of course, how to make all of that efficient on mobile devices. To get us started, I'd love to have you share a little bit about your background, which includes time at places like Google DeepMind and Adobe Research, among others, and tell us a little bit about how you got into AI. Actually, let's see. I I did my PhD almost thirty years ago. And a PhD is actually on this topic of multi agent system. And back then, it was also a very interesting AI topic. And the the the reason I got into AI is is is just by curiosity. During my undergrad, classes in Australia, I I learned about things like Turing test, and, I was, you know, very curious to myself that, you know, how we'll be able to book on a machine to pass the Turing test. And back then, I I I have to be honest that I don't think I'm gonna, live to see the machine passing the Turing test, which, you know, is kind of I take it for granted now. Tell us a little bit more about some of the work that you've done in your career. So, I think my first job, first real job after academia, was at a place called the AI Center at, SRI International. SRI is Stanford Research Institute? That that's right. Yeah. It it is is formally known as Stanford Research Institute. And back then, you know, we're talking about twenty years ago, I got a chance to work on some project called Calo. And back then, we tried …

Get the full transcript (8,623 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all The TWIML AI Podcast transcripts →

You just read a 3-minute summary of a 49-minute episode.

Get The TWIML AI Podcast summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

Books, tools, and gear mentioned in this episode

SignalCast may earn commission on purchases via these links.

Tools

  • Swift EditBy guest

    by VinAI Research

    Swift Edit enables one-step image editing by training an inverted network that converts images to noise, then applies the one-step generation model.
  • Swift BrushBy guest

    by VinAI Research

    Swift Brush eliminates the typical 50-100 denoising steps in diffusion models by distilling multi-step knowledge into a single-step student network, achieving image generation in under 0.25 seconds while maintaining quality scores equal to or better than the original teacher model.

More from The TWIML AI Podcast

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into The TWIML AI Podcast.

Every Monday, we deliver AI summaries of the latest episodes from The TWIML AI Podcast and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime