Skip to main content
Software Engineering Daily

Context-Aware SQL and Metadata with Shinji Kim

41 min episode · 2 min read
·
Shinji Kim

Episode

41 min

Read time

2 min

Topics

Relationships, Startups, Artificial Intelligence

AI-Generated Summary

Key Takeaways

  • Automated metadata collection: SelectStar parses SQL query logs to track which tables join together, join conditions, and usage frequency across users, creating a knowledge graph without manual documentation that reveals actual data relationships and trust signals through behavioral patterns.
  • Three-layer metadata architecture: Physical assets form layer one, usage signals like popularity and lineage comprise layer two, and business context including semantic models and metrics definitions make layer three. This structure enables AI to find correct datasets and generate accurate queries.
  • Cost optimization through usage tracking: Organizations reduce cloud warehouse billing by identifying unused tables and unviewed BI dashboards through popularity metrics. Combining lineage with usage data reveals which data models consume resources without delivering value to end users or downstream systems.
  • MCP server for AI workflows: SelectStar's Model Context Protocol server provides four tools—metadata search, asset details, lineage traversal, and impact analysis—that enable AI agents in Claude and Cursor to generate queries with higher accuracy by accessing popularity scores and example queries.

What It Covers

SelectStar founder Shinji Kim explains how automated metadata platforms solve data discovery challenges by analyzing query logs to build knowledge graphs, enabling AI agents to generate accurate SQL through popularity scores, lineage tracking, and semantic models.

Key Questions Answered

  • Automated metadata collection: SelectStar parses SQL query logs to track which tables join together, join conditions, and usage frequency across users, creating a knowledge graph without manual documentation that reveals actual data relationships and trust signals through behavioral patterns.
  • Three-layer metadata architecture: Physical assets form layer one, usage signals like popularity and lineage comprise layer two, and business context including semantic models and metrics definitions make layer three. This structure enables AI to find correct datasets and generate accurate queries.
  • Cost optimization through usage tracking: Organizations reduce cloud warehouse billing by identifying unused tables and unviewed BI dashboards through popularity metrics. Combining lineage with usage data reveals which data models consume resources without delivering value to end users or downstream systems.
  • MCP server for AI workflows: SelectStar's Model Context Protocol server provides four tools—metadata search, asset details, lineage traversal, and impact analysis—that enable AI agents in Claude and Cursor to generate queries with higher accuracy by accessing popularity scores and example queries.

Notable Moment

Kim reveals that foundation models trained on world data fail against real enterprise databases because messy data with similar table names, denormalized structures, and multi-level calculations causes hallucinations that example queries and popularity context prevent.

Know someone who'd find this useful?

Episode Transcript

A common challenge in data rich organizations is that critical context about the data is often hard to capture and even harder to keep up to date. As more people across the organization use data and data models get more complex, simply finding the right dataset can be slow and create bottlenecks. Select Star is a data discovery and metadata platform that builds a continuously updated knowledge graph of an organization's data by analyzing both its structure and how it's actually used. It enriches data with context such as popularity, lineage, and semantic models, making it easier for AI and teams to discover, trust, and use the right data. These enriched metadata layers are also highly valuable for large language models, significantly improving the accuracy of generated SQL queries. Shinji Kim is the founder and CEO of SelectStar, and she joined Sean Falconer to discuss solving metadata curation challenges, managing data context at scale, using LLMs for SQL generation, emerging trends in metadata management, and more. This episode is hosted by Shawn Falconer. Check the show notes for more information on Shawn's work and where to find him. Shinji, welcome to the show. Thanks, Sean. Great to be here. Yeah. I probably should have said welcome back since you've been here before, although it's been a couple years. Yeah. More than three years ago to introduce Select Star, but I am really excited to be back and software engineering daily has always been, yeah, also morphing and changing a lot. So Yeah. Well, it's been three years. So why don't you catch us up? I mean, three years, especially in the world of tech, the world of startups, and now what's increasingly becoming the world of AI is a lot of time. A lot could happen in three years. So what's happening with Select Star today? Maybe go back even to the beginning. Sort of what's the story behind where you guys started and where are you today? Amazing. Sure. Yes. So much changed. I started Select Star five years ago after noticing time and time that a lot of enterprises collect, store, and process data. But to try to use the data, it takes days or weeks to find the right data and actually use it properly. You have to rely on outdated documentation. Usually, you need to just find somebody else, rely on tribal knowledge to understand how to use the data. I mean, this is something that I saw firsthand at Akamai when I was running the product for their IoT data processing, partnering with consumer electronics and automotive enterprises building their next consumer applications. They were looking to pull a lot more telematics data. And especially in enterprise perspective, this was an issue and, hence, there are solutions like traditional enterprise data catalogs that are trying to solve this issue. At the same time, I've noticed that there was a lot more demand around this also as more companies are adopting quote unquote modern data stack …

Get the full transcript (6,655 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all Software Engineering Daily transcripts →

You just read a 3-minute summary of a 38-minute episode.

Get Software Engineering Daily summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

Books, tools, and gear mentioned in this episode

SignalCast may earn commission on purchases via these links.

Tools

  • by SelectStar

    SelectStar's Model Context Protocol server provides four tools—metadata search, asset details, lineage traversal, and impact analysis—that enable AI agents in Claude and Cursor to generate queries with higher accuracy.
  • SelectStar's Model Context Protocol server provides four tools—metadata search, asset details, lineage traversal, and impact analysis—that enable AI agents in Claude and Cursor to generate queries with higher accuracy.
  • by Anthropic

    SelectStar's Model Context Protocol server provides four tools—metadata search, asset details, lineage traversal, and impact analysis—that enable AI agents in Claude and Cursor to generate queries with higher accuracy.

company

  • SelectStarBy guest
    SelectStar founder Shinji Kim explains how automated metadata platforms solve data discovery challenges by analyzing query logs to build knowledge graphs, enabling AI agents to generate accurate SQL.

More from Software Engineering Daily

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best Cybersecurity Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into Software Engineering Daily.

Every Monday, we deliver AI summaries of the latest episodes from Software Engineering Daily and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime