Czech in the Gears of the Models: A Loss of Style Nobody Is Watching

Studies show that large language models produce stylistically flatter and more monotonous Czech texts than human authors, while Czech makes up less than one percent of training data. The article maps the impact on Czech education, where 69% of pupils use AI for schoolwork, and analyzes the inadequacy of detection tools for the Czech language as well as the inconsistent responses of universities — from annulling bachelor's theses to introducing their own rules for working with AI.
When the team of quantitative linguist Jiří Milička of Charles University had sixteen of the most advanced language models write Czech texts in 2025, modeled on samples from the Czech National Corpus, the oldest of the tested models — OpenAI's davinci-002 — was unable to produce a single coherent Czech paragraph. It collapsed into an unintelligible jumble of words. Newer models fared better, but in a different way: they wrote grammatically acceptable Czech that differed from human texts in a manner the study's authors were able to measure systematically. The Czech produced by a language model was flatter, more monotonous, and stylistically poorer than the human original.
This finding might have remained an academic curiosity were it not for one fact: according to a survey by the company Scio from the turn of 2024 and 2025, 69 percent of Czech pupils use artificial intelligence for school purposes. The texts the models generate — with their specific limitations — thus become part of the linguistic input of an entire generation.
A paradox arises that no one has yet examined systematically. There are dozens of studies on how Czech affects the performance of language models. The opposite direction — how the models affect Czech — is a research blind spot. Yet the first data suggest that the consequences may be far-reaching: from the infiltration of English syntactic patterns, through the impoverishment of stylistic variety, to a reduction in students' cognitive engagement while writing.
Czech makes up less than one percent of the training data of large language models. English accounts for over 54 percent of the web from which the models learn. This imbalance has measurable consequences.
The study "English-Czech Output Bias in LLMs," presented at the ICAIR conference in 2025, tested language models on 48 paired mathematical problems in Czech and English. The result: the models consistently produced more extensive and detailed answers in English. Even though the Czech question had the same content, the English version received a more elaborate explanation — as if the models devoted more "attention" to English. The authors interpret this as evidence that the architecture of the models favors English even when they answer in Czech.
On this point, the Czech Wikipedia entry on large language models states that the models "often present the Anglo-American perspective as the only correct one." No one has yet published a direct study on the increase of anglicisms in Czech texts generated by language models — this is a fundamental research gap. But the broader context exists: Czech linguistics has long tracked the penetration of anglicisms into Czech, from direct loanwords through calques to so-called pseudo-anglicisms, as documented by the 2021 study "Pseudo-Anglicisms in Czech," produced within the Global Anglicism Database Network project.
Artificial intelligence is probably accelerating this tendency — but "probably" is the key word. Direct empirical evidence is so far lacking.
While anglicization remains a hypothesis, stylistic flattening is a measurable fact. In September 2025, Milička and his team published the study "Benchmark of Stylistic Variation in LLM-Generated Texts" in the arXiv repository, in which they applied so-called Biber multidimensional analysis to sixteen top-tier models. They deliberately chose Czech as a representative of mid-sized languages because — as they write — English occupies a wholly exceptional position and results from it cannot be generalized.
The key finding: base models trained on diverse texts lose the ability to imitate stylistically and intellectually varied personalities once they undergo instruction tuning and learning from human feedback. In other words: the process that turns models into "safe" and "useful" conversational partners simultaneously turns them into worse writers.
International research confirms this from several angles. A study from Cornell University for the CHI 2025 conference demonstrated, with 118 participants, that suggestions from language models lead to the adoption of Western and American stylistic patterns. The authors measured how well a classification algorithm could distinguish Indian from American writing: without machine assistance it achieved 90.6 percent accuracy; with it, accuracy fell to 83.5 percent. They named their finding "AI colonialism."
A text published on the UNESCO website under the title "AI and the Great Linguistic Flattening" (2025) described the so-called Default Approach Problem: language models privilege certain expressions and syntactic structures and suppress others. Paradoxically, it suggested that the impact on smaller languages may be inversely proportional to the number of speakers — less institutionalized languages may be more resilient, because they have fewer formal corpora in the training data.
In English academic texts, the flattening is already quantifiable. The word "delve" rose by 1,500 percent in scholarly articles, "underscore" by 1,000 percent, and "intricate" by 700 percent between 2022 and 2024, as shown by review analyses based on studies by Kobak, Liang, and others. At least ten percent of biomedical abstracts in 2024 were, according to Kobak's study, processed by a language model.
For Czech, analogous figures do not exist. The AI-Koditex corpus — the 21.5 million tokens that Milička's team assembled — creates the conditions for obtaining them. But the analyses have not yet been published.
The most direct empirical evidence of what language models do to Czech comes from MiniCzechBenchmark, whose results Petr Šimeček presented at the NeurIPS 2025 workshop. He tested over fifty models and noted two things. First: 10–30 percent performance gaps persist between English and Czech capabilities. Second: accuracy on multiple-choice test questions and the quality of generated text represent distinct skills. A model can answer Czech questions correctly while producing low-quality Czech text.
This dissociation between comprehension and text production has direct implications for education. A student who has a language model write a term paper gets a text that is acceptable in content — the model "understands" the assignment — but stylistically impoverished, because the model "speaks" Czech worse than it "understands" it.
The causes are structural. Czech text requires more tokens than English because of its morphological complexity — seven cases, four genders including the animacy distinction, rich conjugation and derivation. The tokenization penalty means the model expends more computational space to encode the same content, leaving it less capacity for stylistic quality.
How exactly do language models "flatten" Czech? Kamil Kopecký, a Czech researcher focused on digital literacy, documented as early as 2023 systematic errors of GPT-3 in categories that have no English equivalent: confusions between animate and inanimate masculine declensions, errors in vyjmenovaná slova (the set of Czech words spelled with a hard "y"), incorrect distinction between "mě" and "mně," and erroneous case forms. Yet the model excelled at reading comprehension — at identifying key ideas and verifying claims. The dissociation in practice.
At the lexical level, Czech sources — Fakticky.cz, Seznam Zprávy, UIW.cz — identify typical hallmarks of machine text: overuse of the words "klíčové" (key), "výzva" (challenge), and "fascinující" (fascinating); monotonous constructions built on "což" (which); the absence of synonymic variation; and evasive phrasings of the type "It is necessary to bear in mind that…" At the structural level, machine text is characterized by uniform paragraph length, predictable organization, and formulaic transitional phrases.
The data on how many Czech students work with language models are surprisingly rich — and the numbers grow with every survey.
The Scio "AI Kompas" survey from the turn of 2024 and 2025, which polled 3,406 pupils, found that 72 percent of them use artificial intelligence regularly for leisure and 69 percent for school purposes. A STEM survey for the Nekrachni platform reports that 45 percent of pupils aged 11–19 have tried some artificial-intelligence tool, with 89 percent of them using ChatGPT. A survey by the One World in Schools program from the end of 2023 among 1,200 secondary-school students found that 86 percent have experience with artificial intelligence; the most common uses are creating and editing texts (51 percent) and translation (46 percent).
What does this mean for learning? The SYRI National Institute produced the most alarming Czech figure to date: a third of primary-school pupils who use artificial intelligence when preparing for school also believe they no longer need to memorize things. Pupils using these tools showed a tendency toward worse grades in both mathematics and Czech.
This is a correlation, not a proven causal relationship — that must be emphasized. Pupils with worse grades may reach for artificial intelligence as a compensatory tool, not the other way around. A longitudinal study that would prove or refute causation does not exist in the Czech setting.
International research, however, signals that the concerns are not unfounded. A study published in Education Week in June 2025 demonstrated lower brain activity in writers using language models. Professor Steve Graham of Arizona State University pointed out that students miss out on practicing important writing skills. Bai, Liu, and Su, in a 2025 review article for PubMed Central, document that excessive reliance on artificial intelligence reduces cognitive engagement and long-term retention.
On the other hand: a study in the International Journal of AI in Education from the same year showed that generative artificial intelligence can reduce writing time by 56.7 percent and increase output quality, with the greatest benefit recorded among non-native speakers. Artificial intelligence is not only a threat — it is a tool with real potential, whose effects depend on how it is used.
Czech universities responded to artificial intelligence proactively, but inconsistently. Each creates its own rules — sometimes even by faculty.
Masaryk University was first: in April 2023 it issued a position on the use of artificial intelligence in teaching. It views it as an opportunity, requires transparency — classifying undisclosed use as plagiarism — and warns against uploading student work into external detection tools for privacy reasons. Its Faculty of Economics and Administration abolished classic bachelor's theses and replaced them with practical projects.
The most radical step was taken by VŠE (Prague University of Economics and Business). Its Faculty of Business Administration abolished bachelor's theses entirely as of the 2024/25 academic year. Dean Jiří Hnilica justified this by arguing that, with well-crafted prompts, artificial intelligence can produce a substantial part of the thesis and that this realistically cannot be detected. The replacement: practical projects — internships, research projects, business plans.
CTU (Czech Technical University) issued a methodological guideline in September 2023 that distinguishes grammatical correction (requiring no disclosure) from significant textual changes (mandatory declaration). BUT Brno (Brno University of Technology), in the same month, established eight principles and classifies artificial intelligence exclusively as an auxiliary and consultative tool. Charles University set up a working group with prg.ai and runs the central website ai.cuni.cz; its Faculty of Law issued detailed rules including a five-level scale of artificial-intelligence use.
The Ministry of Education took a supportive stance. In 2025 it launched the Základka.ai program with the aim of training all primary-school teachers in working with artificial intelligence. According to the TALIS 2024 survey, 46 percent of Czech teachers use artificial intelligence in teaching — compared with the European average of 32 percent.
A unified nationwide policy does not exist. Coordination is being attempted by the working group at Charles University and prg.ai, which submitted recommendations to the Czech Rectors Conference, but implementation in practice depends on individual institutions.
When VŠE tried detection tools, it ultimately abandoned them. They were not reliable enough. And VŠE is not alone.
Seznam Zprávy tested several detectors on Czech texts. The best result was achieved by Copyleaks, with 75 percent correct identification out of twenty varied texts — which the article's authors described as "markedly unreliable." A test by the Fakticky.cz server from February 2025 found even lower figures for most tools: ZeroGPT reached 52 percent, Writer AI Content Detector 50 percent.
The only purpose-built Czech tool, DetekceGPT.cz by Kryštof Olík and Filip Petroušek, claims over 90 percent accuracy on machine text and 98 percent on human text. But these figures come from the creators and have not been independently verified. Among international tools, GPTZero reports a 96.4 percent detection rate for Czech at a 0.1 percent false-positive rate — roughly 3 percentage points lower than for English.
A fundamental problem was described by the team of Debora Weber-Wulff and Czech researcher Tomáš Foltýnek of Mendel University in a 2023 study: they tested fourteen detection tools and concluded that they are neither accurate nor reliable, with a bias toward classifying output as human-written. They tested only English — for Czech the results would probably be even worse.
The Stanford study by Liang and colleagues from 2023 revealed another layer of the problem: 61 percent of essays by non-native English speakers on the TOEFL exam were classified by detectors as language-model text, whereas for native speakers the detectors worked almost flawlessly. The statistical properties of texts by less experienced writers — shorter sentences, simpler vocabulary — overlap with the properties of machine text. Detectors can thus systematically penalize students who did not use artificial intelligence.
The existing Czech infrastructure — the Theses.cz and Odevzdej.cz systems operated by Masaryk University — focuses on detecting plagiarism, not machine text. Turnitin, used at Charles University since 2019, offers machine-text detection with limitations for non-English languages. The National Library of Technology recommends using detectors only as auxiliary tools, not as evidence.
Poland offers the closest comparison. Since 2019 it has operated the Unified Anti-Plagiarism System — a nationwide tool mandatory for all universities. In February 2024 a machine-text detection module was added to it, available free of charge. Every thesis in Poland is automatically checked for both plagiarism and machine content — a unique solution by European standards.
Germany emphasizes legal compliance under the EU AI Act, which since February 2025 classifies artificial intelligence in education as potentially high-risk. As of November 2024, however, only about 30 percent of German universities had published formal guidelines for working with artificial intelligence, as reported by Weßels and Lindner on the Forschung und Lehre portal. A 2023 legal opinion from Ruhr-Universität Bochum confirmed that a blanket ban on artificial intelligence makes no sense.
France underwent a rapid transformation. Sciences Po Paris formally banned the submission of work created by artificial intelligence in January 2023, but today it is rethinking its assessment criteria on three levels: assessing skills rather than the product, adapting instruction, and incorporating teaching about artificial intelligence. Université Gustave Eiffel formulated a principle that also applies to the Czech context: previously, what the student produced was assessed; now it is necessary to assess how the student manages the creative process.
Common patterns: no blanket bans, mandatory transparency, a shift from assessing the product to assessing skills.
It would be convenient to tell a story about how artificial intelligence is destroying Czech and dulling students. But the data are more nuanced.
First: the 10–30 percent performance gap between English and Czech in language models is shrinking with each generation. MiniCzechBenchmark tested models from 2023 to 2025, and newer models perform markedly better. Claude 3 Opus, in CzechBench tests from CTU, demonstrated the highest grammatical competence in Czech of all the tested models. The problems Kopecký documented in GPT-3 are less frequent in GPT-4 and newer models.
Second: the text on the UNESCO website itself admits that smaller languages may paradoxically be more resilient to homogenization. Fewer formal corpora in the training data means that artificial intelligence has less "leverage" to influence the language norm. Czech also has a strong linguistic tradition — the Institute of the Czech Language, the Czech National Corpus, a living language-advisory service — that serves as a counterweight.
Third: the causal relationship between the use of artificial intelligence and worse academic results is not proven. The correlation from the SYRI study may equally well reflect the fact that weaker students reach for artificial intelligence more often.
And fourth: the OpenEuroLLM project, led by Jan Hajič of Charles University, with a total budget of 34 million euros — of which over 20 million is from the Digital Europe programme — for a multilingual model for 32 European languages, shows that Czech research institutions are not merely analyzing the problem but actively working on a solution.
Jiří Milička and his team created a corpus of 21.5 million tokens of Czech generated by language models — a tool that, for the first time, makes it possible to measure how machine Czech differs from human Czech. Fourteen hundred kilometers to the north, a Polish nationwide system operates that automatically checks every thesis for machine content. In Prague, Brno, and Ostrava, meanwhile, five universities are tackling the same problem in five different ways.
Research into the influence of artificial intelligence on Czech is in its early stages. We know that the models understand Czech better than they generate it. We know that generated Czech is stylistically poorer. We know that most Czech students work with artificial intelligence and that detection tools are unreliable for Czech. We do not know whether, and how, artificial intelligence is changing the vocabulary and syntax of the generation growing up with it.
The answer to that question depends on whether someone asks it early enough — and precisely enough.
Underlying data and sources as of the date of preparation: February 2026. This analysis was prepared with the assistance of artificial intelligence. Key studies: Milička et al. (2025), AI-Brown and AI-Koditex, arXiv:2509.22996; Milička et al. (2025), Benchmark of Stylistic Variation, arXiv:2509.10179; Šimeček (2025), MiniCzechBenchmark, NeurIPS LLM Evaluation Workshop; Fajčík et al. (2024), BenCzechMark, TACL; Agarwal, Naaman & Vashistha (2025), CHI; Weber-Wulff, Foltýnek et al. (2023), IJEI; Scio AI Kompas, STEM/Nekrachni, and JSNS surveys.
Transparency of creation:
The concept, structure, and editorial line of the article are the work of the author, who developed the content outline, established the key theses, and directed the entire creative process. Generative AI (Claude, Anthropic) was used as a technical tool for research, fact-checking, and fleshing out the author's draft.
The author continuously edited the outputs, verified the key findings, and approved the final wording. No part of the text was published without human review. All factual data were verified against the publicly available sources cited in the text.
This procedure complies with the requirements of Article 50 of EU Regulation 2024/1689 (the AI Act) on the transparency of AI-generated content. #poweredByAI
Read the Czech original on Médium.cz.
AI · Claude — machine translation, may contain inaccuracies.