Skip to main content
Bankless

AI Finds 70% of Smart Contract Exploits | Alpin Yukseloglu

61 min episode · 3 min read
·
Alpin Yukseloglu

Episode

61 min

Read time

3 min

Topics

Relationships, Investing, Fundraising & VC

AI-Generated Summary

Key Takeaways

  • AI exploit capability trajectory: Frontier models went from finding roughly 12–13% of critical smart contract bugs to over 70% within six months — a jump that occurred partly between drafting and publishing the EVM Bench paper. GPT-5.3 Codex now matches or approaches the collective output of human auditors on fund-draining vulnerabilities sourced from Code Arena audit contests. Expect a superhuman AI auditor within six to eight months.
  • False positive elimination via verifiability: Previous AI auditing tools produced high false positive rates, making them impractical. EVM Bench solves this by running exploits against a production-grade EVM environment loaded with real chain state. If the agent claims a bug exists, it must produce a working proof-of-concept that drains funds from the contract — reducing false positives to near zero and making AI audit results actionable.
  • Long-tail contract risk: Low-TVL protocols on EVM-compatible chains like Binance Smart Chain face the highest near-term exploit risk. These contracts were historically sheltered because the maximum extractable value was too small to attract skilled attackers. As inference costs drop below the value of exploiting even small contracts, AI agents will systematically collect this long tail — making security investment non-optional regardless of protocol size.
  • Crypto's verifiability accelerates AI training: Crypto code is among the most verifiable software in existence — agents can deploy contracts, assert state changes, and confirm exploits without human labelers. This creates a tight training signal that accelerates model improvement faster than most software domains. Paradigm expects models to develop strong crypto capabilities with less direct training data than initially anticipated, compressing the timeline to superhuman performance.
  • Offense-defense arms race framing: The near-term security outcome depends on whether white-hat or black-hat actors access frontier AI capabilities first. Paradigm's strategic response is embedding crypto benchmarks directly inside model labs — EVM Bench is now running inside OpenAI — to ensure defensive tooling develops alongside offensive capability. Protocols housing significant TVL should begin proactive AI-assisted auditing now rather than waiting for the first AI-attributed exploit.

What It Covers

Alpin Yukseloglu, investment and research partner at Paradigm, presents findings from EVM Bench — a benchmark co-authored with OpenAI measuring AI agents' ability to detect, patch, and exploit smart contract vulnerabilities. Top models jumped from under 20% to over 70% exploit detection in six months, reshaping crypto security assumptions.

Key Questions Answered

  • AI exploit capability trajectory: Frontier models went from finding roughly 12–13% of critical smart contract bugs to over 70% within six months — a jump that occurred partly between drafting and publishing the EVM Bench paper. GPT-5.3 Codex now matches or approaches the collective output of human auditors on fund-draining vulnerabilities sourced from Code Arena audit contests. Expect a superhuman AI auditor within six to eight months.
  • False positive elimination via verifiability: Previous AI auditing tools produced high false positive rates, making them impractical. EVM Bench solves this by running exploits against a production-grade EVM environment loaded with real chain state. If the agent claims a bug exists, it must produce a working proof-of-concept that drains funds from the contract — reducing false positives to near zero and making AI audit results actionable.
  • Long-tail contract risk: Low-TVL protocols on EVM-compatible chains like Binance Smart Chain face the highest near-term exploit risk. These contracts were historically sheltered because the maximum extractable value was too small to attract skilled attackers. As inference costs drop below the value of exploiting even small contracts, AI agents will systematically collect this long tail — making security investment non-optional regardless of protocol size.
  • Crypto's verifiability accelerates AI training: Crypto code is among the most verifiable software in existence — agents can deploy contracts, assert state changes, and confirm exploits without human labelers. This creates a tight training signal that accelerates model improvement faster than most software domains. Paradigm expects models to develop strong crypto capabilities with less direct training data than initially anticipated, compressing the timeline to superhuman performance.
  • Offense-defense arms race framing: The near-term security outcome depends on whether white-hat or black-hat actors access frontier AI capabilities first. Paradigm's strategic response is embedding crypto benchmarks directly inside model labs — EVM Bench is now running inside OpenAI — to ensure defensive tooling develops alongside offensive capability. Protocols housing significant TVL should begin proactive AI-assisted auditing now rather than waiting for the first AI-attributed exploit.
  • Agency over singularity anxiety: When facing uncertainty about AI's trajectory, Yukseloglu recommends replacing speculative theorizing with direct experimentation at the frontier. Both full acceptance and full denial of AI risk produce the same passive outcome. The practical alternative is running experiments, engaging model labs directly, and shipping within 24 hours of inception — speed over cohesion is the correct operating mode when the frontier remains experimentally unknowable.

Notable Moment

Yukseloglu describes a counterintuitive dynamic: Solana's prevalence of closed-source contracts, typically seen as a disadvantage, may actually accelerate AI model development on that stack. Contracts absent from public training data provide cleaner, uncontaminated evaluation signals — potentially giving closed-source ecosystems an unexpected edge in AI capability benchmarking.

Know someone who'd find this useful?

Episode Transcript

Bank of Station, we are here with Alpen Yupsololu. He is an investment and research partner at Paradigm, also the co author of a pay paper titled EVM Bench, an open benchmark for smart contract security agents written in collaboration with OpenAI, to measure the ability of AI agents to just detect or patch or exploit smart contract vulnerabilities. We're gonna talk about the way that AI and AI capabilities are going to impact our crypto ecosystem, our smart contracts. Alpen, welcome to Bankless. Hi. Thanks for having me. I wanna start off the question with a very big this podcast with a very big question. How at risk are we from AI? How large of a threat does AI smart contract capabilities pose to our industry? Yeah. I mean, in the long term, it's now increasingly clear that AI is going to be extremely, extremely good for crypto because especially on the security front because we're going to get to a world where because everything is much more secure, the ceiling on the industry is much higher. So our partner, Matt, talks about how, if you have a grocery store that's run by mom and pop, because they can't see everything in the store, there's a limit to how big they can get. But the moment you add security cameras in, security has this effect of increasing the capacity, the carrying capacity of an industry. I think in the short term, it's up to us because the models are getting extremely good, like, strikingly good. When we started working on EVM Bench, which is a benchmark that consists entirely of fund draining critical bugs, around six months ago, the models were able to find less than 20% of the bugs, like around 12 to 13%. And just over the course of while we were working on the benchmark, this number went up to over 50%. And in between when I drafted the launch tweet and when I had to actually hit send with the release of Jeep 5.3 codecs, it jumped up to over 70%. So these things are just growing at a blistering pace, and it's very important that we position the industry in a way that we can defensively protect against attacks. But in the long term, I think it massively increases the carrying capacity of crypto. Yeah. I think what you're saying is in the long term, we get something approaching perfect security. Yeah. Right now, we do not have perfect security. Let me ask you the same question, but a little bit differently. Say only bad actors, only black hat hacked, actors have access to AI capabilities. In that context, how at risk is our industry? Like, how exploitable are our smart contracts given the increase in AI capabilities? Yeah. I mean, I think it's really hard to say when we approach superintelligence levels. I do think until we hit the like, right now, the models are quite good, but they're not better than the …

Get the full transcript (11,560 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all Bankless transcripts →

You just read a 3-minute summary of a 58-minute episode.

Get Bankless summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

Books, tools, and gear mentioned in this episode

SignalCast may earn commission on purchases via these links.

Tools

  • GPT-5.3 Codex now matches or approaches the collective output of human auditors on fund-draining vulnerabilities sourced from Code Arena audit contests.
  • EVM BenchRecommended

    by Paradigm

    Alpin Yukseloglu, investment and research partner at Paradigm, presents findings from EVM Bench — a benchmark co-authored with OpenAI measuring AI agents' ability to detect, patch, and exploit smart contract vulnerabilities.

company

  • EVM Bench is now running inside OpenAI — to ensure defensive tooling develops alongside offensive capability.
  • Alpin Yukseloglu, investment and research partner at Paradigm, presents findings from EVM Bench — a benchmark co-authored with OpenAI

More from Bankless

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best Crypto Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Investing & Markets Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into Bankless.

Every Monday, we deliver AI summaries of the latest episodes from Bankless and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime