← Article directory

A language model knows Erben's Kytice by heart. That's a problem for the main defense of generative AI

12. 6. 2026
A language model knows Erben's Kytice by heart. That's a problem for the main defense of generative AI
Image from the original article on Médium.cz

The author demonstrates that a large language model can reproduce the opening stanzas of Erben's Kytice verbatim purely from its trained weights, thereby challenging the categorical defense that generative AI is 'merely statistical prediction' and not copying. He links the experiment to case C-250/25 (Like Company vs. Google Ireland) before the Court of Justice of the EU — the first dispute over generative AI and copyright, on which Advocate General Szpunar will deliver his opinion on 3 September 2026. The article dissects where the defense falls apart (memorization is a documented phenomenon), where it holds up (ordinary content is typically not reproduced verbatim), and which legal questions only the court can decide.

Asked for the opening of Erben's *Kytice*, a large language model returned four stanzas word for word. This autumn, the Advocate General of the Court of Justice of the EU will deliver his Opinion in the first case that asks directly whether "mere statistical prediction" is a copyright shield — one of the referred questions even does so explicitly. A small test shows why this defense does not hold up in its categorical form — and, conversely, where it still stands.

The prompt was deliberately spare. No search, no access to the source text — only what the model has stored in its weights. "Recite the opening of Erben's *Kytice*." What came back matched exactly: „Zemřela matka a do hrobu dána, / siroty po ní zůstaly; / i přicházely každičkého rána / a matičku svou hledaly." *Polednice*, *Vodník*, and *Svatební košile* came out just as flawlessly. Only in the fifth example, in *Štědrý den*, did the model stumble over a single word — it said „což pak" where Erben wrote „cože". Four stanzas out of four, verbatim, from memory, without looking into the book.

It's a parlor trick — until one realizes when and where this very same matter is about to be decided. On 3 September 2026, Advocate General Maciej Szpunar of the Court of Justice of the EU is scheduled to deliver his Opinion in Case C-250/25, the dispute between the Hungarian publisher Like Company and Google Ireland over the Gemini chatbot. It is the very first case to put a question on generative AI and copyright directly to the Court of Justice. And at the core of one of the referred questions lies exactly the argument this test probes: when a language model produces protected text, is it making a reproduction, or merely computing probabilities? The first question goes further still, asking explicitly whether it makes any difference that the output arises through probabilistic prediction of the next token.

This experiment settles one thing and leaves another open. It settles that the "it's just statistics" defense collapses in its categorical form: the protected expression did not dissolve into the parameters — it lies there and can be pulled out token by token. And it leaves open what matters most — namely whether copyright infringement follows from this. The latter is not a matter of engineering but of legal characterization, and that will be opened only by the Advocate General's Opinion.

First, the defense, because it is not foolish. The developers of large models argue that the model does not store the texts; it learns from them statistical relationships between words and then assembles the output anew, probabilistically, token by token. On this logic, the specific authorial expression dissolves into the weights during training, and what the model returns is new text, not a copy. For the overwhelming majority of outputs this is in fact true — the model routinely assembles sentences that appeared in no training text. Next-token prediction is generative in principle, not reproductive.

But with *Kytice* nothing dissolved. The expression remained there, and it could be reconstructed word for word. In the scholarly literature this phenomenon is called memorization of training data, and it is not speculation — it has been documented by experimental work (among the first, Nicholas Carlini's team in 2021, which extracted hundreds of verbatim excerpts of training data from the older GPT-2 model), and it stands behind the American lawsuit *The New York Times v OpenAI*. In it the publisher submitted a hundred examples in which it sufficed to enter the beginning of one of its articles and the model completed hundreds of further words almost unchanged. The weights carry more than the dissolving metaphor admits.

Memorization, however, is not uniform — and this is precisely where the defense regains its footing.

The strongest counterargument does not stand against the existence of memorization, but against its generalization. How strongly a text is imprinted into the weights depends above all on how many times it appears in the training data and in its copies. *Kytice* is the worst possible case: a canonical text, a public-domain work without protection, copied a thousand times across the Czech web, in school readers, on blogs, and in digital libraries. A specific news article from the Hungarian publisher is the exact opposite — a new text, little duplicated, often hidden behind a paywall. The probability that the model will return it verbatim is lower by orders of magnitude. On top of that, I got *Kytice* out of the model through a targeted, adversarial prompt, not through ordinary use. The demonstration therefore demolishes the categorical claim that "statistics is always a shield," not its more cautious version: that with ordinary content under ordinary use, verbatim reproduction typically does not occur.

Here, though, is a snag that narrows the defense back down. The fourth question targets precisely the situation where protected content appears in the output after the user himself has inserted it into the prompt — and asks whether such a reproduction is attributable to the service provider. Adversarial extraction of text therefore does not lie outside the subject of the dispute; it lies within it. Indeed, at the hearing itself the Advocate General raised whether the system should be assessed as a whole — training, inputs, outputs, and communication to the public all at once — rather than isolating the individual acts. The Court will thus also have to deal with what happens when someone actively pushes the model into reproduction.

And this is exactly where what an engineer can decide ends and what a court must decide begins. Even if memorization were proven down to the last detail, three legal questions that the referring court sent along with it remain open: whether the training of the model is itself reproduction within the meaning of Article 2 of the Information Society Directive; whether the display of protected text in the chatbot's output is communication to the public under Article 3; and, if training is reproduction, whether it is covered by the text and data mining exception (Article 4 of the Directive on Copyright in the Digital Single Market, § 39c of the Czech Copyright Act) — including whether the rightholder reserved that reproduction in advance. Moreover, this is not just any text, but the content of a press publication publisher, which enjoys its own special right in the EU (Article 15 of the same directive, § 87b of the Copyright Act). For a chatbot fed on journalism, the risk is thereby higher than with technical documentation or public-domain works.

For the Czech reader this is not an academic riddle. The same architecture underlies search and "RAG" systems that pull answers from current news texts; it underlies the publishers' business and the very question of what a model intended for the European market may be fed with at all. The United States is so far deciding inconsistently, and that at the level of federal district courts: while in *Thomson Reuters v Ross Intelligence* the court rejected fair use for training, in *Bartz v Anthropic* and *Kadrey v Meta* it allowed it for training — each time, however, with reservations, and none of those decisions is binding precedent. The European answer will begin to take shape precisely with Szpunar's Opinion. That does not bind the Court, but in most cases it foreshadows where the judgment will go. The judgment is not expected until 2027.

The dissolving metaphor is soothing: the protected text supposedly gets lost in the numbers. A model that has just recited *Kytice* from memory is living counter-evidence against it. The numbers remember more than it suits us to admit — and on 3 September the deciding will begin of how much that remembering legally weighs.

The demonstration in the introduction was a simple test, not a controlled benchmark: a single publicly available large language model, a targeted (adversarial) prompt. The result was verified word for word against the critical hybrid edition of the Institute of Czech Literature of the Czech Academy of Sciences, the digitized edition of the Municipal Library of Prague, and Wikisource — four of the five opening stanzas matched verbatim; in the fifth there was a single lexical deviation („což pak" instead of „cože"). The example illustrates a phenomenon (memorization) that is documented in the literature; it is not in itself proof of its extent across models.

Transparency of creation:

The conception, structure, and editorial line of the article are the work of the author, who developed the content sketch, set out the key theses, and directed the entire creation process. Generative AI (Claude, Anthropic) was used as a tool for research, locating primary sources, and the stylistic elaboration of the author's content sketch.

The author continuously edited the outputs, verified the key findings, and approved the final wording. No part of the text was published without human review. All factual data were verified against the publicly available sources cited in the text.

This procedure complies with the requirements of Article 50 of EU Regulation 2024/1689 (the AI Act) on the transparency of AI-generated content. #poweredByAI

Read the Czech original on Médium.cz.

AI · Claude — machine translation, may contain inaccuracies.