← Article directory

A New National Revival: Czechia Has No AI of Its Own — and Its Military Pays the Price

22. 2. 2026
A New National Revival: Czechia Has No AI of Its Own — and Its Military Pays the Price
Image from the original article on Médium.cz

The Czech Republic has no language model of its own usable for national-security purposes, which represents a strategic gap across defense, the intelligence services, and the police. While France, Germany, and the United Kingdom invest heavily in the military deployment of AI — and language models perform 10–30% worse in Czech than in English — Czech security forces are not addressing the problem publicly. The article compares the approaches of Iceland, Finland, and Estonia and warns of the risks of depending on American and Chinese technologies for processing sensitive data.

Forty volunteers on an island of 370,000 people took matters into their own hands. In 2023, a group of linguists, programmers and enthusiasts came together in Iceland and began manually training GPT-4 in Icelandic using reinforcement learning from human feedback. They provided OpenAI with over 4 billion words of high-quality Icelandic text. The result: inflection accuracy in GPT-4 Turbo rose from 25% to 66%. The total government investment over five years was approximately 14 million euros.

At the same time, the Czech Republic — with a population of ten million, one of the strongest traditions in computational linguistics in Europe, and the company Seznam.cz, which operates its own search engine — had, and still has, no language model usable for national security purposes. Neither the military, the intelligence services, nor the police have publicly stated any intention to build or deploy such a model.

This article is not about technical performance benchmarks. It is about a strategic gap that concerns state security, cultural identity and technological sovereignty. The benchmarks do exist, and they say something clear: language-model performance in Czech is 10–30% worse than in English. Not a single standalone model has passed the Czech bar exams. Models confuse the Velvet Revolution with the collapse of Austria-Hungary and invent Czech words that do not exist. And while France in January 2026 concluded a framework contract with the company Mistral AI for its entire armed forces, and Germany is investing billions of euros in the defense company Helsing, the Czech security agencies are not publicly addressing this topic at all.

The situation resembles the linguistic dependence of the era before the Czech National Revival. This time, however, it is not only about culture — it is about who controls the information infrastructure on which the state depends.

The Czech Republic's dependence on foreign artificial-intelligence technology is multi-layered. All the specialized chips (NVIDIA, AMD) come from the USA. The dominant language models — GPT-4o, Claude, Gemini, Llama, DeepSeek — are developed outside Europe. Most computational processing takes place on American cloud platforms. And in the training data from which the models learn, the less-represented half of all the EU's official languages put together makes up a mere 2.4% of the data.

The most advanced Czech commercial initiative is SeLLMa by Seznam.cz — a model with 70 billion parameters built by continued fine-tuning of the open-source Llama 3.1 model. Development was begun by a team of roughly fifty people in the second half of 2023 with an investment of over 100 million crowns, including the purchase of hundreds of NVIDIA H100 chips. SeLLMa powers the chatbot Seznam Asistent, launched on 25 November 2025, and other products — search-result summaries, ad creation, discussion moderation. The key point is the strategic intent: according to Diana Hlaváčová, product manager for large language models, Seznam does not want to depend on third-party technology and wants full control over a model running on its own servers.

But SeLLMa is not a model built from scratch. It is a fine-tuned version of American Llama. Seznam also temporarily uses OpenAI models via the Azure platform. And any continued fine-tuning inherits the limitations of the base model — its architecture, the way it splits text into units, and its implicit cultural framework.

At the European level there is a groundbreaking project, OpenEuroLLM, coordinated by Professor Jan Hajič of the Institute of Formal and Applied Linguistics at the Faculty of Mathematics and Physics, Charles University. The project, with a budget of 37.4 million euros (of which 20.6 million is from the Digital Europe programme), brings together more than 20 partners from across Europe. The goal is to develop a family of fully open multilingual models covering all the EU's official languages and other strategically important languages. The first version of the models is expected by mid-2026, the final version by the start of 2028.

As the coordinating country, the Czech Republic is in a privileged position. The Institute of Formal and Applied Linguistics brings decades of experience with Czech language resources — the Prague Dependency Treebank, the LINDAT/CLARIAH-CZ infrastructure, the RobeCzech model. Brno University of Technology created BenCzechMark, the first comprehensive Czech benchmark suite for language models. The academic ecosystem is strong.

What is missing is a link to the security sector. OpenEuroLLM is a research-and-commercial project, not a security one. The Czech Republic's National Artificial Intelligence Strategy 2030, approved by the government on 24 July 2024 with planned investments of 19 billion crowns, contains no separate chapter on language sovereignty. The forthcoming KarolAIna supercomputer (850 petaflops, going into operation at the end of 2026) could in theory also serve to train a security model — but no such intention has been stated.

A state security analyst — whether a criminal investigator, an intelligence officer, a cybersecurity specialist, a military officer, or a tax-administration official — today faces the same dilemma. They need to process large volumes of text in Czech, but have no language model they can trust.

Imagine an investigator at the National Organized Crime Agency examining large-scale economic crime. On the desk are thousands of pages — accounting documents, contracts, email correspondence, registry extracts, expert reports. A language model could identify patterns, link entities across documents, summarize key passages. An American model can do this — but in English. In Czech it understands context worse: it cannot distinguish a "jednatel" (managing director) from a "jednatel v likvidaci" (managing director in liquidation), it does not understand the specific terminology of Czech insolvency law, it does not know what a "věřitelský výbor" (creditors' committee) is in the context of Act No. 182/2006 Sb. And above all: the investigator must not upload the contents of a criminal case file onto the servers of an American corporation.

Imagine an analyst at the Security Information Service monitoring disinformation operations targeting the Czech public. Every day they need to process hundreds of social-media posts, articles from alternative media, and Telegram channels. A model could sort narratives, recognize coordinated behavior, and map the spread. But a model trained on English data does not grasp why the phrase "patriotic opposition" is a different signal in the Czech context of 2026 than in the American context. It does not understand that "alternative media" in the Czech Republic span a spectrum from legitimate criticism to coordinated Kremlin propaganda. It does not know who's who on the Czech disinformation scene — and benchmarks show that it fails on factual knowledge of Czech realities.

Imagine a specialist at the National Cyber and Information Security Agency analyzing a cyberattack on Czech critical infrastructure. The logs are partly in Czech, the internal communications of the targeted organization are in Czech, and the incident report must meet the requirements of the Czech Cybersecurity Act. A model would help with triaging the incident, drafting the report, and analyzing patterns in the logs. But every query sent through a commercial interface travels to the servers of an American company subject to the CLOUD Act, which allows U.S. authorities to access data regardless of where it is physically located. For processing information about an attack on Czech critical infrastructure, that is an unacceptable risk.

Imagine an officer of Military Intelligence preparing a situational analysis. They need to process open sources — Czech, Slovak, Polish, German, Russian. A model could summarize, translate, and identify key information. But not one of the available models has been vetted by the Czech security agencies. None has been fine-tuned on Czech military doctrines, operational procedures, or intelligence terminology. And the officer has no assurance that a model — whether American or Chinese — does not carry hidden biases that would influence the analytical output.

And imagine an official at the Financial Analytical Office tracking suspicious transactions linked to sanctioned entities. Every day they process reports from banks, exchange offices, and crypto exchanges — in Czech, with Czech legal terminology, with ties to specifically Czech corporate structures. A model could link entities, identify patterns, and flag anomalies. Without a high-quality Czech model, they do it by hand.

The common denominator of all five scenarios: an open-source alternative exists — running Llama, Mistral, or another open model on Czech infrastructure. But none of these models has been vetted for Czech security purposes, fine-tuned on Czech domain documents, or certified. Llama is American, Mistral is French, Qwen is Chinese, DeepSeek is Chinese. The continued fine-tuning by Seznam (SeLLMa) improves the quality of Czech, but it is designed for commercial purposes — not for processing classified documents.

The BIS annual report for 2023 describes artificial intelligence as a "huge challenge" for the security agencies — but with an emphasis on the risks of fake videos and images, not on the absence of in-house tools. In its National Cybersecurity Strategy 2026–2030, NÚKIB identifies artificial intelligence as both an opportunity and a threat, without formulating any concrete intention to build or deploy an independent language model. As for Military Intelligence, the Office for Foreign Relations and Information, the Financial Analytical Office, and the Police of the Czech Republic, no publicly available information exists in the context of their own language models.

The contrast with allies is striking.

France has set the benchmark in Europe. In January 2026 the Ministry of the Armed Forces concluded a framework contract with the company Mistral AI covering all branches of the military. The models will be deployed exclusively on French infrastructure; sensitive data will not leave the jurisdiction. The ministerial agency AMIAD oversees the rollout, and Mistral will fine-tune the models on defense-specific data for logistics, intelligence analysis, and simulations. France, moreover, has its own world-class company — Mistral is valued at around 14 billion dollars.

Germany is investing, through the company Helsing, in autonomous combat systems. Helsing, valued at approximately 5.4 billion euros and with more than 500 employees, develops unmanned aircraft, aerial combat systems, and underwater surveillance systems. The CA-1 Europa project with the company HENSOLDT targets an autonomous combat aircraft with purely European technology.

The United Kingdom, in its Strategic Defence Review 2025, designated artificial intelligence as a central element of the transformation of its armed forces, with an investment of one billion pounds in digital combat infrastructure. In December 2025 the USA launched the GenAI.mil platform for 3 million Pentagon employees, with contracts with the companies Anthropic, OpenAI, xAI, and Google.

NATO, in its revised artificial-intelligence strategy of July 2024, urgently calls for the adoption of generative artificial intelligence "as soon as possible."

On the adversaries' side: China's Xian Technological University demonstrated a military simulation based on the DeepSeek model, in which the system generates 10,000 combat scenarios in 48 seconds — compared with 48 hours of human planning. In February 2026 OpenAI warned the U.S. Congress that Chinese labs are continuing the so-called adversarial distillation of America's most advanced models — the systematic copying of capabilities through artificially generated data.

In the Czech Republic there is no equivalent of France's AMIAD, the British Defence AI Centre, or the American GenAI.mil. According to the Oxford Insights Government AI Readiness Index, the Czech Republic ranks around 30th globally (28th in 2024), and an absolute minimum of defense-budget funds goes into artificial-intelligence projects. In April 2023 the military signed a memorandum of cooperation with CTU in the area of breakthrough technologies, and students of the University of Defence experimented with software for processing satellite imagery — but no systematic program for deploying language models in defense exists publicly.

The gulf between English and Czech in language models is measurable and systematic.

BenCzechMark, created by the CZLC consortium (Brno University of Technology, FI MUNI, CIIRC CTU), tests over 50 models on 50 tasks, 90% of which are in original, untranslated Czech. The results confirm persistent 10–30% performance drops between English and Czech across all the tested models. On the Umímeto-Czech task, 19 of the 50 evaluated models score zero.

The test by the law firm HAVEL & PARTNERS on 1,840 questions from the bar exams of the Czech Bar Association produced an unambiguous result: not one standalone model passed the exam in any of the five rounds (the threshold being 85 out of 100 points). Claude-3-Opus performed best, but even it did not cross the threshold. Only a system combining GPT-4 with a Czech legal database (the retrieval-augmented generation method, known as RAG) passed the exam — which shows that without specific Czech sources, models are not sufficient for Czech law.

Cultural knowledge is even more problematic. Martin Richter of the University of Hradec Králové (Historia Aperta, 2025) experimentally demonstrated that, when describing Czech history, language models confused the Velvet Revolution with the collapse of Austria-Hungary and invented Czech words that do not exist, including their inflected forms. The models tend toward a semantic framing of conflict and power imbalance when describing Czech-German relations, thereby producing a narrowed interpretation.

The case of Jára Cimrman represents a natural litmus test of an artificial intelligence's Czech cultural literacy. As a fictional character deliberately presented within a "playing at reality" framework, Cimrman tests whether a model grasps the meta-level of Czech humor — or whether it presents him as a real historical figure. No formal study yet exists, but anyone who has attempted it knows that the results tend to be both comical and disquieting at once.

The research by Tao et al. (2024, PNAS Nexus), testing five versions of GPT against data from 107 countries, demonstrated that the cultural values of the GPT models are closest to those of the Anglo-Saxon and Protestant European countries and most distant from the values of Eastern Europe. In 2025 the Ada Lovelace Institute identified three sources of cultural bias: algorithmic monoculture in the language-model market, the predominance of Western training data, and post-training fine-tuning with predominantly Western evaluators.

For the security agencies this has concrete consequences: a model that does not understand the Czech legal system, distorts Czech history, and applies an Anglo-Saxon cultural framework is not a suitable tool for intelligence analysis of the Czech environment.

The Czech Republic is not the only country grappling with language sovereignty in artificial intelligence. A comparison with countries in a similar situation shows that solutions exist — and that they differ fundamentally in both ambition and cost.

Iceland (370,000 speakers) bet on cooperation with the giants. Forty volunteers training GPT-4 and 4 billion words of high-quality text delivered directly to OpenAI brought a measurable result at a fraction of the cost. Total investment: approximately 14 million euros over five years.

Finland (5.5 million speakers) chose the most ambitious path. The LUMI supercomputer (380 petaflops, an investment of 144.5 million euros) serves as the base for the Viking 7B/13B/33B models developed by the company SiLo AI. The company AMD bought that firm in August 2024 for 665 million dollars. The Finnish government added a further 250 million euros for expansion and is seeking to host a pan-European artificial-intelligence factory.

Estonia (1.1 million speakers) is integrating artificial intelligence directly into state administration. The "Kratt" strategy identified over 200 uses in central government; 37% of officials already use artificial intelligence. Investment of approximately 130 million euros, with the goal of saving 21 million hours of administrative work per year by 2030. Less emphasis on an in-house cutting-edge model, more on practical deployment.

Norway (5.4 million speakers) brought in media companies as providers of training data and the National Library as the administrator of the data infrastructure for artificial intelligence. The government commitment: 1 billion Norwegian crowns (approximately 90 million euros) for research over five years.

Hungary (13 million speakers, the closest comparison with the Czech Republic) chose an academic approach. The PULI models (6.7 billion parameters) and adaptations of Llama are developed by HUN-REN NYTK, mostly on research grants. Smaller, efficiently adapted models show that results can be achieved even with limited means. Hungary, however, lacks an industrial partner of the Seznam.cz type.

Israel represents a paradox: a world leader in artificial-intelligence talent (1st place in citations per paper), yet without a national Hebrew language model. In August 2025 the Nagel Commission warned that Israel "is neither in the right place nor on the desired trajectory" — a direct parallel with the Czech Republic, where talent is high but infrastructure insufficient.

In September 2025 Switzerland released the Apertus model, with 70 billion parameters trained on 15 trillion text units on the Alps supercomputer. A fully open approach with an investment of 20 million Swiss francs — but it requires a massive supercomputing backbone that the Czech Republic does not yet have.

For the Czech Republic, a mixed path suggests itself: Icelandic-style cooperation with global providers to improve Czech, the Finnish model of using European supercomputers for training, and — this is the key missing piece — the French security framework with dedicated infrastructure for defense purposes.

Language models produce Czech text that a Czech can spot. Not because it contains gross errors, but because it sounds "different." Excessive formality, an absence of colloquial expressions, a tendency toward long compound sentences with connectors like "nicméně" (nevertheless) and "nadto" (moreover), insufficient use of Czech idiom, imitation of English sentence structures, and a complete absence of regional variants and common Czech — that is the accent of artificial intelligence.

The cause is structural. As Czech sources confirm, the models first internally process the content of the Czech input in English and then formulate the answer back in Czech. This is not native Czech thinking, but a hidden translation. A study with 48 paired mathematical problems (6 models, 576 answers) showed that the English answers are consistently longer and more detailed than the Czech ones across all the tested models — even though Czech is a morphologically richer language that should, in theory, allow for more concise expression.

For everyday conversation this may be immaterial. For producing official texts, analyses, or intelligence outputs it is a problem. A model that "thinks in English" will formulate Czech texts according to Anglo-Saxon stylistic conventions — and, along with them, adopt the Anglo-Saxon cultural and analytical framework.

Three objections deserve honest consideration.

First: is an independent model realistic? Training a cutting-edge model from scratch costs hundreds of millions of dollars and requires thousands of compute units. The Czech Republic does not have, and will not have in the foreseeable future, the means to compete with companies like OpenAI or Google. This objection is valid — but no one is claiming that the Czech Republic needs to compete with GPT-5. It needs a model of sufficient quality for specific Czech security and administrative purposes, fine-tuned on Czech data and operated on Czech infrastructure. Iceland showed that even an investment on the order of millions of euros can deliver measurable results.

Second: won't OpenEuroLLM solve this? The Prague-coordinated project is ambitious and important. But it covers dozens of languages, its initial focus is academic-commercial, and the first models will be available no earlier than mid-2026. The security dimension — vetting the models for intelligence purposes, fine-tuning on classified documents, certification for defense deployment — is not contained within the project.

Third: isn't continued fine-tuning of existing models, as Seznam does, sufficient? For commercial purposes, probably yes. For security purposes it is more complicated. Continued fine-tuning improves linguistic quality, but it does not provide full control over the model's architecture, its security properties, or what data went into the base training. The French approach — deploying a vetted model on dedicated infrastructure — represents a higher level of control.

Despite these reservations, the key fact remains: France, Germany, the United Kingdom, and the USA are actively addressing the security dimension of artificial intelligence. The Czech Republic is not even publicly discussing whether it should.

In the first decades of the 19th century, the situation of the Czech language looked hopeless. German dominated science, administration, and higher education. Czech survived as a "low" language, spoken in the countryside and in pubs. The revivalists — Jungmann, Palacký, Šafařík — consciously built linguistic infrastructure: dictionaries, grammars, encyclopedias, scientific terminology. Not because it was economically rational, but because they understood language sovereignty as a precondition of political emancipation.

The parallel with the present is analytically substantive, not merely rhetorical. Today English dominates the infrastructure of artificial intelligence just as German dominated the science and administration of that era. Czech security documents, intelligence analyses, and military doctrines are peripheral to the global models, processed through an English filter. And just as then, it holds true that linguistic dependence is not merely a cultural problem — it is a problem of power.

The Czech Republic coordinates the largest European project for open language models, has a strong academic tradition, and the only European technology company operating its own model for the Czech market. It has a supercomputer that by the end of the year will be capable of training models. It possesses talent, data, and institutions.

What it lacks is the will to connect these capabilities with the security sector. A realistic investment in a national artificial-intelligence language program — 50 to 100 million euros over five years — represents a fraction of the planned expenditure on large infrastructure projects. Jungmann created his work without state support, under difficult conditions, because he understood what was at stake. Today the Czech state has the means, the infrastructure, and the knowledge. The question is not whether it can afford to do this. The question is whether it can afford not to.

This analysis was prepared with the help of artificial intelligence. Data and sources were verified as of the date of processing (22 February 2026). Before use for decision-making, we recommend independent verification of the key findings.

BenCzechMark: Fejgl et al. (2024), Transactions of the Association for Computational Linguistics, MIT Press. 50 tasks, 90% native Czech data. arxiv.org/html/2412.17933v1

Czech bar exams: HAVEL & PARTNERS (March 2024), 1,840 questions from the Czech Bar Association, 5 rounds. en.havelpartners.blog

OpenEuroLLM: over 20 partners, budget 37.4 million euros, coordinator ÚFAL MFF UK. openeurollm.eu

SeLLMa / Seznam Asistent: continued fine-tuning of Llama 3.1, 70 billion parameters, investment over 100 million CZK. Lupa.cz (25 November 2025), e15.cz

NAIS 2030: National Artificial Intelligence Strategy of the Czech Republic 2030, Ministry of Industry and Trade, approved 24 July 2024. mpo.gov.cz

France/Mistral — defense contract: January 2026, framework contract for all branches of the armed forces. Reuters (8 January 2026)

Helsing: valued at approximately 5.4 billion euros, autonomous defense systems. tech-now.io

USA GenAI.mil: December 2025, deployment for 3 million Department of Defense employees. DefenseScoop (9 December 2025)

NATO — artificial-intelligence strategy: revised in July 2024. nato.int

China/DeepSeek — military simulations: South China Morning Post

Cultural bias of language models: Tao et al. (2024), PNAS Nexus. Oxford Academic

Gender bias in Czech: Martinková, Stanczak, Augenstein (2023), SlavicNLP workshop at the ACL conference. aclanthology.org

Czech history in language models: Richter (2025), Historia Aperta vol. 54. journals.uhk.cz

Iceland/Icelandic in GPT-4: Miðeind, cooperation with OpenAI. mideind.is, ArcticToday

Finland/LUMI: 380 petaflops, investment 144.5 million euros. CSC (csc.fi); SiLo AI, AMD acquisition for 665 million USD. amd.com (12 August 2024)

Estonia/Kratt: over 200 uses in state administration, 37% of officials. The Mandarin

Hungary/PULI: HUN-REN NYTK, Racka-4B. arxiv.org (2601.01244)

Israel — Nagel Commission: August 2025. Israel Tech Insider

Switzerland/Apertus: ETH Zurich, 70 billion parameters, Apache 2.0 license. ethz.ch (September 2025)

NÚKIB — cybersecurity: National Strategy 2026–2030. nukib.gov.cz

Czech Army — CTU memorandum: April 2023. acr.mo.gov.cz

KarolAIna / Czech AI Factory: 850 petaflops, going into operation at the end of 2026. Innovate Moravia

Transparency of creation:

The concept, structure, and editorial line of the article are the work of the author, who developed the content outline, established the key theses, and directed the entire creative process. Generative AI (Claude, Anthropic) was used as a technical tool for research, fact-checking, and fleshing out the author's draft.

The author edited the outputs throughout, verified the key findings, and approved the final wording. No part of the text was published without human review. All factual data were verified against the publicly available sources cited in the text.

The procedure complies with the requirements of Art. 50 of EU Regulation 2024/1689 (the AI Act) on transparency of AI-generated content. #poweredByAI

Read the Czech original on Médium.cz.

AI · Claude — machine translation, may contain inaccuracies.