In 5 Years, 90% of What You Use AI For Will Run on Your Smartphone | Paolo Ardoino, Tether
Episode
58 min
Read time
2 min
Topics
Productivity, Investing, Fundraising & VC
AI-Generated Summary
Key Takeaways
- ✓Edge AI Timeline: By 2029, smartphones will handle 70% of common AI tasks locally; by 2031, that figure reaches 90%. Tether's medical model with 4 billion parameters already outperforms Google's MedGemma at 27 billion parameters, and a 1.7 billion parameter version runs on an average $80 African smartphone today.
- ✓QVAC Architecture: Tether's open-source QVAC platform combines a modified llama.cpp inference engine with a restructured BitNet one-bit weight model, making it compatible with consumer GPUs including Snapdragon, Adreno, and Apple chips. Developers integrate it as an SDK, adding text, voice, image, and fine-tuning capabilities directly into existing applications.
- ✓On-Device Fine-Tuning via LoRA: Low Rank Adaptation allows users to customize AI models locally by adjusting only a subset of model weights, enabling a personalized assistant trained on private emails and documents without sending data externally. QVAC standardizes this LoRA fine-tuning framework across hundreds of supported models on any consumer GPU.
- ✓AI ASIC Efficiency Shift: Purpose-built AI ASICs already run Llama 3.2 at 17,000 tokens per second versus 150 tokens per second on a good GPU — a 113x improvement. Within five years, enterprises and banks could deploy $50,000–$100,000 on-premise ASIC clusters, eliminating cloud dependency while reducing energy consumption by up to 98%.
- ✓Data Center Financial Risk: Major AI companies currently subsidize $200 subscriptions that cost $1,000–$5,000 to deliver, masking losses while private to inflate valuations ahead of planned 2026–2027 IPOs. Once public, this subsidy becomes unsustainable, and third-party-built data center capacity — contracted via offtake agreements — may be left significantly underutilized.
What It Covers
Paolo Ardoino, CEO of Tether, argues that within five years, 90% of everyday AI use cases will run locally on smartphones, making massive centralized data centers largely redundant for consumer applications, while explaining Tether's QVAC platform and its open-source approach to edge AI deployment.
Key Questions Answered
- •Edge AI Timeline: By 2029, smartphones will handle 70% of common AI tasks locally; by 2031, that figure reaches 90%. Tether's medical model with 4 billion parameters already outperforms Google's MedGemma at 27 billion parameters, and a 1.7 billion parameter version runs on an average $80 African smartphone today.
- •QVAC Architecture: Tether's open-source QVAC platform combines a modified llama.cpp inference engine with a restructured BitNet one-bit weight model, making it compatible with consumer GPUs including Snapdragon, Adreno, and Apple chips. Developers integrate it as an SDK, adding text, voice, image, and fine-tuning capabilities directly into existing applications.
- •On-Device Fine-Tuning via LoRA: Low Rank Adaptation allows users to customize AI models locally by adjusting only a subset of model weights, enabling a personalized assistant trained on private emails and documents without sending data externally. QVAC standardizes this LoRA fine-tuning framework across hundreds of supported models on any consumer GPU.
- •AI ASIC Efficiency Shift: Purpose-built AI ASICs already run Llama 3.2 at 17,000 tokens per second versus 150 tokens per second on a good GPU — a 113x improvement. Within five years, enterprises and banks could deploy $50,000–$100,000 on-premise ASIC clusters, eliminating cloud dependency while reducing energy consumption by up to 98%.
- •Data Center Financial Risk: Major AI companies currently subsidize $200 subscriptions that cost $1,000–$5,000 to deliver, masking losses while private to inflate valuations ahead of planned 2026–2027 IPOs. Once public, this subsidy becomes unsustainable, and third-party-built data center capacity — contracted via offtake agreements — may be left significantly underutilized.
Notable Moment
Ardoino points out that the human brain operates on roughly the same electricity as a potato clock, yet current AI infrastructure consumes tens of kilowatts to approximate its function — suggesting the entire architectural approach to AI compute may be fundamentally misaligned with long-term efficiency goals.
Episode Transcript
If in 2026, people can already run local models on their smartphones to solve only 50% of the use cases of AI, in five years, 95% of the population are able to run 90% of the use cases of AI on a smartphone. They're building all of these data centers, and it's not at all clear that the compute paradigm that we're currently operating under is gonna continue, in which case maybe you don't need all of those data centers. So if you don't understand how your AI works, it's not you becoming more intelligent. It's someone else becoming more intelligent with your data. The big transformational market is enterprise and government use, and those will likely run on data centers because problems that they're solving, the speed at which they have to solve them are gonna require massive compute. Is that possible that there'll be both? Why don't we start by having you introduce yourself to listeners about how you came to to be involved with Tethr and what Tethr is and why it's important, which a lot of people, including me, don't quite understand. Look. Tether was a company is a company born in 2014. Simple idea. Digital dollar. There are plenty of digital dollars, but only USD, our digital dollar, was able to achieve the holy grail of distribution and financial inclusion and impact in the world. Over the last twelve years, our company built the most used digital dollar in the world with five seventy three million users growing by 30 plus million users per quarter. I call it the biggest financial inclusion success story in the history of humanity. Sometimes my blood boils because how is possible that a little company like Tether was able to achieve more for financial inclusion than all NGOs and charities and whatnot for the last fifty years. And we did it in a very simple way. We used new technologies like blockchain to make the dollar accessible. And you you asked me before, well, you know, why not just use the regular dollar? Well, the reality is that for four, five billion people, so more than half of the population of the world, they cannot have access to that regular dollar. They don't have access to basic financial services. The people that are unbanked, the number of people that are unbanked, the world is just enormous. And these are not it's not like they are are unbanked because they are bad people. They are unbanked because they are too poor for being of interest of the banking system. They live in countries where their national currency is devaluating so fast against the US dollar. Think about Argentina. The the Argentinian peso lost 94.5% of its value against the US dollar in the last five years. The Turkish lira lost 81% of its value against the US dollar in the last five years. The Venezuelan Boulevard lost 99.8% of its value against the US dollar Nasdaq. It could go on. …
Get the full transcript (8,464 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 55-minute episode.
Get Eye on AI summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Eye on AI
Inside Ukraine's Azov Drone R&D: The Engineer Building AI Weapons 18 km From the Front Line | Alexander Palamarchuk
Aug 27 · 41 min
Hard Fork
Tim Cook’s Legacy + The Future of U.B.I. With Andrew Yang + HatGPT
Apr 24
More from Eye on AI
95% of AI Agent Projects Fail to Reach Production. Here's Why | Manoj Saxena, TrustWise
Aug 24 · 61 min
The AI Breakdown
Something Big Is Happening
Feb 15
More from Eye on AI
We summarize every new episode. Want them in your inbox?
Inside Ukraine's Azov Drone R&D: The Engineer Building AI Weapons 18 km From the Front Line | Alexander Palamarchuk
95% of AI Agent Projects Fail to Reach Production. Here's Why | Manoj Saxena, TrustWise
From Zero to 150 Robots in Just 20 Months | Mike LeBlanc, Foundation Future Industries
Why People Are Paying 10x More for AI - and What That Means for the Chip Market | Sid Sheth, d-Matrix
American Companies Have 36 Months to Go AI-Native or Get Left Behind | Drew Cukor, TWG AI
Similar Episodes
Related episodes from other podcasts
Hard Fork
Apr 24
Tim Cook’s Legacy + The Future of U.B.I. With Andrew Yang + HatGPT
The AI Breakdown
Feb 15
Something Big Is Happening
Techmeme Ride Home
Feb 11
The “Covid Moment” For AI?
What Bitcoin Did
Feb 10
#147 - Andrea Miotti - The War Against AI Has Begun
10% Happier with Dan Harris
Aug 24
Longevity Secrets (And Controversies) From The Blue Zones | Dan Buettner
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Investing & Markets Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into Eye on AI.
Every Monday, we deliver AI summaries of the latest episodes from Eye on AI and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime