Sovereign AI in Poland: Language Adaptation, Local Control & Cost Advantages with Marek Kozlowski
Episode
89 min
Read time
3 min
Topics
Fundraising & VC, Artificial Intelligence, Software Development
AI-Generated Summary
Key Takeaways
- ✓Language Adaptation vs. Full Pretraining: Rather than training from scratch — which requires at least 1 trillion tokens for stable results — PLUM continues pretraining Llama and Mistral base models on ~200 billion curated Polish tokens. This "language adaptation" injects local linguistic and cultural knowledge while preserving existing multilingual capabilities, achieving competitive Polish-language performance without the compute cost of full pretraining runs.
- ✓Frontier Model Quality Degrades for Niche Languages Over Generations: Benchmarking on Poland's PLCC (Polish Holistic Cultural Competency) benchmark reveals that successive Claude and GPT model releases show declining Polish language and cultural performance. As frontier labs prioritize coding and reasoning benchmarks, niche language quality becomes a trade-off casualty — meaning organizations relying on cloud APIs risk worsening performance over time without any warning or recourse.
- ✓Small Fine-Tuned Models Match Large Cloud Models for Specific Tasks: When a business has 10–20 defined use cases and prepares at least 1,000 supervised fine-tuning instructions per task, a smaller on-premise model matches zero-shot or few-shot performance from large cloud LLMs. This approach reduces energy costs, eliminates cloud dependency, and enables deployment in regulated sectors where data cannot leave the organization's infrastructure.
- ✓Domain Adaptation Requires ~10 Billion Clean Tokens to Be Worthwhile: PLUM's work with Central Eastern Europe's largest bank demonstrates that domain-specific continued pretraining delivers measurable quality gains — but only when the organization can supply roughly 10 billion tokens post-deduplication and filtering. Since raw data shrinks by 3–4x through curation, organizations need 30–40 billion raw tokens, a threshold fewer than 100 European companies realistically meet.
- ✓EU Regulation Eliminates ~80% of Usable Training Data: The EU AI Act combined with local authorship rights legislation creates constraints far stricter than any voluntary model constitution. These regulations prevent large-scale web scraping and require detailed model cards disclosing training data, compute, and security measures. PLUM compensates by securing bilateral agreements with publishers and libraries, and by building internal human annotation pipelines producing organic instruction and preference datasets.
What It Covers
Marek Kozlowski, head of Poland's National Information Processing Institute AI Lab, explains how Project PLUM (Polish Large Language Models) builds locally controlled AI by performing language adaptation on Llama and Mistral base models using ~200 billion curated Polish tokens, targeting performance parity with models 10x larger for language and cultural tasks.
Key Questions Answered
- •Language Adaptation vs. Full Pretraining: Rather than training from scratch — which requires at least 1 trillion tokens for stable results — PLUM continues pretraining Llama and Mistral base models on ~200 billion curated Polish tokens. This "language adaptation" injects local linguistic and cultural knowledge while preserving existing multilingual capabilities, achieving competitive Polish-language performance without the compute cost of full pretraining runs.
- •Frontier Model Quality Degrades for Niche Languages Over Generations: Benchmarking on Poland's PLCC (Polish Holistic Cultural Competency) benchmark reveals that successive Claude and GPT model releases show declining Polish language and cultural performance. As frontier labs prioritize coding and reasoning benchmarks, niche language quality becomes a trade-off casualty — meaning organizations relying on cloud APIs risk worsening performance over time without any warning or recourse.
- •Small Fine-Tuned Models Match Large Cloud Models for Specific Tasks: When a business has 10–20 defined use cases and prepares at least 1,000 supervised fine-tuning instructions per task, a smaller on-premise model matches zero-shot or few-shot performance from large cloud LLMs. This approach reduces energy costs, eliminates cloud dependency, and enables deployment in regulated sectors where data cannot leave the organization's infrastructure.
- •Domain Adaptation Requires ~10 Billion Clean Tokens to Be Worthwhile: PLUM's work with Central Eastern Europe's largest bank demonstrates that domain-specific continued pretraining delivers measurable quality gains — but only when the organization can supply roughly 10 billion tokens post-deduplication and filtering. Since raw data shrinks by 3–4x through curation, organizations need 30–40 billion raw tokens, a threshold fewer than 100 European companies realistically meet.
- •EU Regulation Eliminates ~80% of Usable Training Data: The EU AI Act combined with local authorship rights legislation creates constraints far stricter than any voluntary model constitution. These regulations prevent large-scale web scraping and require detailed model cards disclosing training data, compute, and security measures. PLUM compensates by securing bilateral agreements with publishers and libraries, and by building internal human annotation pipelines producing organic instruction and preference datasets.
- •Organic Human-Annotated Data Drives Quality at the SFT Stage: Synthetically generated instruction data from other LLMs degrades model output quality when those synthetic examples contain poor linguistic structure. PLUM employs dozens to hundreds of human annotators to create and review instructions and preference pairs manually. This organic data pipeline — combined with publishing dataset samples and a ~100-page technical cookbook on Hugging Face — differentiates PLUM from open-weight-only releases that share no training data transparency.
Notable Moment
Kozlowski reveals that when his team analyzed successive Claude model releases against the PLCC benchmark, Polish cultural and linguistic performance measurably declined across versions. This means organizations that deeply integrate a cloud LLM into Polish-language workflows could find their vendor's next release quietly performs worse on their core use case with no rollback option available.
Episode Transcript
This podcast is sponsored by Google. Hey, folks. I'm Omar, product and design lead at Google DeepMind. We just launched a revamped Vibe Coding experience in AI Studio that lets you mix and match AI capabilities to turn your ideas into reality faster than ever. Just describe your app, and Gemini will automatically wire up the right models and APIs for you. And if you need a spark, hit I'm feeling lucky, and we'll help you get started. Head to a i.studio/build to create your first app. Hello, and welcome back to the Cognitive Revolution. While we often discuss Sovereign AI in the Silicon Valley AI bubble, we rarely hear directly from the technical leaders who are actually leading national AI projects. And so today, I'm very glad to share my conversation with Merit Koslowski, who's leading project PLUM, which stands for Polish large language models, in his role as head of the AI lab at the National Information Processing Institute of Poland. Poland, with a population of 38,000,000 and GDP of roughly 1,000,000,000,000, roughly 103% of The United States respectively, is an interesting and in some ways a representative case study. It clearly doesn't have the resources required to compete with The US and China at the AI frontier, but it does have strong technical talent, a real sense of pride in its language and culture, and a deep desire to control its own technological destiny and avoid domination by global superpowers. So what does that mean in practice? As you'll hear, Merrick's strategy relies on the core belief that by training small models for a particular local language and cultural context, countries like Poland and projects like PLUM can compete with the latest frontier models, all while retaining control, preserving data privacy, and achieving a major cost advantage. In this conversation, we dig into the strategic realities that motivate projects like Plum and the technical challenges that they have to overcome to succeed, including how today's frontier models, which are trained on overwhelmingly English and Chinese data, fall short in other languages, why this problem is actually getting worse from one generation to the next as frontier model developers prioritize things like coding performance above support for niche languages, how EU regulation prevents European AI builders from conducting massive web scrapes and instead forces them to rely on more focused data curation projects, how the Polish government is thinking about investing its finite resources across data, compute, and talent, the language adaptation techniques that Merrick's team layers on top of LAMA and Mistral based models so as to inject local knowledge without needing to start from scratch, why they haven't yet had to worry about developing a constitution or other explicit articulation of values for Polish AI systems, and why government agencies and national champion companies are often better served by smaller models fine tuned for specific tasks and served locally than by massive generalist models served from the cloud. Overall, Merrick's mix of realism about the challenges …
Get the full transcript (16,346 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 86-minute episode.
Get Cognitive Revolution summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Cognitive Revolution
AI:AM Highlights: Recursive Self-Improvement, Rushed and Vibe-Coded?
Aug 28 · 131 min
Everything Everywhere Daily
The Lincoln-Douglas Debates
Jun 6
More from Cognitive Revolution
RL's a Hell of a Drug: Metagaming, Reward Seeking & Motivated CoT Reasoning – Bronson Schoen, Apollo
Aug 26 · 134 min
Odd Lots
Why Susquehanna Is Building a Prediction Markets Business
Jun 6
Books, tools, and gear mentioned in this episode
SignalCast may earn commission on purchases via these links.
Tools
by Hugging Face
“This organic data pipeline — combined with publishing dataset samples and a ~100-page technical cookbook on Hugging Face — differentiates PLUM from open-weight-only releases”
by National Information Processing Institute
“Benchmarking on Poland's PLCC (Polish Holistic Cultural Competency) benchmark reveals that successive Claude and GPT model releases show declining Polish language and cultural performance.”
by Meta
“PLUM (Polish Large Language Models) builds locally controlled AI by performing language adaptation on Llama and Mistral base models using ~200 billion curated Polish tokens”
by Mistral AI
“PLUM (Polish Large Language Models) builds locally controlled AI by performing language adaptation on Llama and Mistral base models using ~200 billion curated Polish tokens”
More from Cognitive Revolution
We summarize every new episode. Want them in your inbox?
AI:AM Highlights: Recursive Self-Improvement, Rushed and Vibe-Coded?
RL's a Hell of a Drug: Metagaming, Reward Seeking & Motivated CoT Reasoning – Bronson Schoen, Apollo
AI in the AM — Weekly Highlights: Relaunch Week (Aug 17–20, 2026)
Let There Be Germicidal Light: This $500 Fixture Could Stop the Next Pandemic, from Complex Systems
Lindy Teammate: Flo Crivello on Multiplayer Agents, Memory & Why He'd Ban the Chinese Models He Uses
Similar Episodes
Related episodes from other podcasts
Everything Everywhere Daily
Jun 6
The Lincoln-Douglas Debates
Odd Lots
Jun 6
Why Susquehanna Is Building a Prediction Markets Business
The Daily (NYT)
Jun 1
Inside Trump’s Mad Dash to Renovate Washington
Lenny's Podcast
Apr 19
Why half of product managers are in trouble | Nikhyl Singhal (Meta, Google)
Eye on AI
Apr 15
#334 Abhishek Singh: The $1.2 Billion Plan to Turn India Into an AI Superpower
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's AI & Machine Learning Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into Cognitive Revolution.
Every Monday, we deliver AI summaries of the latest episodes from Cognitive Revolution and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime