Why Vision Language Models Ignore What They See with Munawar Hayat - #758
Episode
57 min
Read time
2 min
Topics
Productivity, Fundraising & VC, Artificial Intelligence
AI-Generated Summary
Key Takeaways
- ✓Vision Token Attention Failure: Vision language models attend poorly to visual tokens despite having images as input. Cross-attention modules injected every fourth transformer block with auxiliary loss maximizing attention on relevant segmentation masks improves visual grounding while reducing compute complexity from m+n squared to n squared plus mn.
- ✓Physics Understanding Gap: Foundation models fail simple physical reasoning tasks like unstacking boxes, generating deformed objects with changed sizes and properties. Prompt expansion describing physical constraints during training helps, but models trained on trillions of text tokens versus billions of image-text pairs struggle with spatial correspondence and physical world simulation.
- ✓Generalized Contrastive Learning: Standard CLIP training uses image-text pairs only, failing at composed queries mixing modalities. Training on all permutations of image, text, and fused embeddings without collecting new triplet data enables cross-modal retrieval and generalizes to video benchmarks despite no video training data, maintaining CLIP's 300-400 million parameter efficiency.
- ✓Multi-Person Generation Solution: Models lose facial identity when generating multiple people and fail accurate person counts beyond three to four subjects. Defining attention masks preventing tokens of one person from attending to another person's tokens reduces identity leakage, enabling inference-only personalization without fine-tuning adapters for each face.
What It Covers
Munawar Hayat from Qualcomm AI Research discusses three NeurIPS papers addressing critical failures in vision language models: why they ignore visual input, physics-based generation limitations, and multi-person image generation challenges with proposed solutions.
Key Questions Answered
- •Vision Token Attention Failure: Vision language models attend poorly to visual tokens despite having images as input. Cross-attention modules injected every fourth transformer block with auxiliary loss maximizing attention on relevant segmentation masks improves visual grounding while reducing compute complexity from m+n squared to n squared plus mn.
- •Physics Understanding Gap: Foundation models fail simple physical reasoning tasks like unstacking boxes, generating deformed objects with changed sizes and properties. Prompt expansion describing physical constraints during training helps, but models trained on trillions of text tokens versus billions of image-text pairs struggle with spatial correspondence and physical world simulation.
- •Generalized Contrastive Learning: Standard CLIP training uses image-text pairs only, failing at composed queries mixing modalities. Training on all permutations of image, text, and fused embeddings without collecting new triplet data enables cross-modal retrieval and generalizes to video benchmarks despite no video training data, maintaining CLIP's 300-400 million parameter efficiency.
- •Multi-Person Generation Solution: Models lose facial identity when generating multiple people and fail accurate person counts beyond three to four subjects. Defining attention masks preventing tokens of one person from attending to another person's tokens reduces identity leakage, enabling inference-only personalization without fine-tuning adapters for each face.
Notable Moment
When testing proprietary foundation models on simple box unstacking tasks, researchers found models that generate intricate visual details fail basic physics, producing deformed boxes with altered sizes, revealing a fundamental gap in spatial reasoning despite impressive general capabilities.
Episode Transcript
If you take these foundation models, proprietary models, and you should give it very simple images, like an image of two boxes. Two cardboard boxes, very plain. One is bigger, one is smaller. And you ask it to generate an image where we unstack these boxes. So once it unstacks, physical properties of the boxes change. They are not exactly the same boxes. Their shapes might be deformed. Their sizes might be different. And this is problematic. Alright, everyone. Welcome to another episode of the TwiML AI podcast. I am your host, Sam Charrington. Today, I'm joined by Munawar Hayat. Munawar is a researcher at Qualcomm AI Research. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Munawar, welcome to the podcast. Hey, Sam. Thanks a lot. Thanks for having me. I'm really looking forward to digging into our conversation. We're going to be talking about some of Qualcomm's papers on multimodal AI and visual AI from the recent NeurIPS conference. To get us started, I'd love to have you share a little bit about your background. How'd you get into the field? So I did my PhD in 2015 from Australia. I, got interested in computer vision, and I took some courses in my undergrad, on image processing. That's what fascinated me. And, during my PhD, I worked on, visual, analysis of facial data. And then afterwards, I have been contributing to mostly computer different subfields of computer vision. And, yeah. And then in '20, after my PhD, I stayed in academia. I've had a faculty position, and then I moved to Qualcomm in 2023. What's your research focus at Qualcomm? At Qualcomm, I have been, mostly working on, multimodal generative AI. So on vision language models for understanding generation and retrieval. And, those are the three areas where I have my, papers at NeurIPS that I believe we are gonna, discuss deeply. So, when I joined Qualcomm, I've the first project they worked on was, making diffusion models run efficiently on Qualcomm hardware. That's a mobile, phone. And, we're generating images in under half a second. And then, we also work on visual understanding. That's visual question answering. Given an image and a question, a text question, we want to respond to the question. And for real question answering, we also had a model running on Qualcomm hardware. So, basically, here in Qualcomm, I have been working primarily on multimodal generative AI on, generating visual content, understanding visual content, and also retrieving cross model information. When you think about how the field has evolved over the past couple years since you started there, what's most exciting for you about this moment right now? I guess we have come a long way. Like, for decades, we were focused on solving niche, very specific problems. Think of we spend so much time in just solving, classifying different objects from visual images. And then, we moved on quite a …
Get the full transcript (7,840 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 54-minute episode.
Get The TWIML AI Podcast summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from The TWIML AI Podcast
World Models and the Future of Spatial AI with Justin Johnson - #775
Sep 1 · 66 min
Planet Money
Don't hate the replicator, hate the game
Feb 27
More from The TWIML AI Podcast
Why the Next AI Breakthrough May Come from Physics with Max Welling - #774
Aug 25 · 55 min
Dwarkesh Podcast
The rise and fall of agent civilizations
Aug 31
More from The TWIML AI Podcast
We summarize every new episode. Want them in your inbox?
World Models and the Future of Spatial AI with Justin Johnson - #775
Why the Next AI Breakthrough May Come from Physics with Max Welling - #774
Why Image Generation Needs More Than Bigger Models with Fatih Porikli - #773
Why Models Are AI’s Next Training Dataset with Damian Borth - #772
How AI Learns to Smell with Alex Wiltschko - #771
Similar Episodes
Related episodes from other podcasts
Planet Money
Feb 27
Don't hate the replicator, hate the game
Dwarkesh Podcast
Aug 31
The rise and fall of agent civilizations
Deep Questions with Cal Newport
Aug 31
Rethinking the Deep Life Stack (Again!) | Monday Advice
Modern Wisdom
Aug 31
WW3 DEBATE: “We’re On the Brink of Global Collapse” - #1144
The Joe Rogan Experience
Aug 26
#2546 - Michael Button
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's AI & Machine Learning Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into The TWIML AI Podcast.
Every Monday, we deliver AI summaries of the latest episodes from The TWIML AI Podcast and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime