May Czech Authorities Block AI Bots? The robots.txt Legal Grey Zone

The article analyses the legal questions surrounding the blocking of AI bots by Czech public institutions through the robots.txt file. It examines the clash of three legal regimes — copyright law (official works without protection vs. reservation of rights), the Act on Free Access to Information, and EU directives on open data. It concludes that blanket blocking is legally questionable, particularly for official works, which are not subject to copyright protection, and for scientific research, where the law does not permit a reservation of rights.
When, in September 2024, the Hamburg district court — in the case of Kneschke v. LAION — became the first court in Europe to rule on the exception for the automated analysis of text and data in the context of training artificial intelligence, it opened a debate that concerns the entire European Union. In December 2025, the court of appeal essentially upheld its conclusions and clarified the rules for the reservation of rights. In the Czech Republic, however, this debate has a specific dimension: a number of public institutions — from universities to municipal authorities — block access by artificial-intelligence crawlers to their websites in the robots.txt file. Are they doing so in accordance with the law?
The answer is not simple, and the definitive word will rest with case law. At the intersection, three legal regimes meet that contradict one another, and the outcome depends on what content specifically is at stake.
Czech law creates three distinct regulatory frameworks for the content of public institutions, which overlap and at times contradict each other.
The Copyright Act (No. 121/2000 Coll.) distinguishes two key categories. Section 3 excludes official works from copyright protection — legal regulations, decisions, measures of a general nature, public instruments, publicly accessible registers and the collections of their documents, publications of the Chamber of Deputies and the Senate, municipal chronicles, state symbols, and other works in which there is a public interest in their exclusion from protection. The 2022 amendment (Act No. 429/2022 Coll., effective from 5 January 2023) then introduced, in Section 39c, a general exception for the automated analysis of text and data, under which anyone is entitled to make a reproduction of a work for the purpose of such analysis. The author may, however, expressly reserve this use in an appropriate manner — for works made available to the public via the internet, by machine-readable means. It is precisely the robots.txt file that tends to serve as the instrument of such a reservation.
The Freedom of Information Act (No. 106/1999 Coll.) imposes on obliged entities — state bodies, territorial self-governing units, public institutions — the duty to provide information. Since September 2022, moreover, through the amendment No. 241/2022 Coll., it transposes EU Directive 2019/1024 on open data, which introduces the principle of open data as the default standard for publicly accessible registers maintained on the basis of law.
The Directive on copyright in the Digital Single Market (2019/790), in Article 4, establishes an exception for the automated analysis of text and data with the possibility of a reservation of rights by rightholders, but its premise is that a rightholder exists at all — that is, that the content is protected by copyright.
The core of the problem is Section 3 of the Copyright Act. Official works are not subject to copyright protection. No subjective copyright arises in respect of them for the author at all — the public interest in their free dissemination excludes them from protection.
From this it can be inferred that if an official work is not protected by copyright, its creator (a public-administration body) should not be able to assert a reservation of rights under Section 39c(2) of the Copyright Act. This reservation, after all, presupposes the existence of copyright protection from which the exception departs. Where there is nothing to protect, the reservation loses any basis in law. The definitive word, however, should belong to case law or to expert legal analysis.
If this interpretation holds, a robots.txt file blocking artificial-intelligence crawlers should have no legal effect with respect to the following content:
This enumeration is not closed — Section 3(a) contains the wording "and other such works in which there is a public interest in their exclusion from protection," which opens space for a broadening interpretation. The boundary between an official and a non-official work is not always sharp, however, and in practice may be the subject of dispute.
Public institutions also produce content that is not an official work. Press releases, editorial articles, photo galleries, educational materials, expert analyses — all of these may be subject to copyright protection. For this content, institutions may assert a reservation of rights through robots.txt or another machine-readable means.
The question, however, is whether they should.
Several arguments speak against a reservation of rights for the content of public institutions. Act No. 106/1999 Coll. establishes the general principle of the openness of public administration — no one has to demonstrate a legal interest in order to access information, and the obliged entity must provide the requested information. EU Directive 2019/1024 on open data reinforces this principle and introduces open data as the default standard for the registers envisaged by law. And finally, content created with public funds should — as a matter of principle — be publicly available.
In favour of a reservation of rights, one can argue protection against misinterpretation (hallucinations of generative artificial intelligence), the risk of being taken out of context, and the protection of personal data in certain materials. These arguments are legitimate, but they are addressed through other legal instruments (the General Data Protection Regulation, the Personal Data Protection Act), not through blanket blocking of crawlers.
A special category is formed by the exception for the scientific automated analysis of text and data under Section 39d of the Copyright Act. Universities, cultural heritage institutions, and research organisations may carry out automated analysis of text or data for the purposes of scientific research, without the rightholder being able to prohibit this manner of use. Here the law expressly does not permit a reservation of rights. Moreover, reproductions may be retained permanently for the purposes of verifying research results.
This leads to a remarkable situation: a public university, which itself enjoys the benefits of the scientific exception without the possibility of a reservation of rights, at the same time blocks crawlers on its own website. From the standpoint of the internal consistency of the legal order, this is at the very least debatable — although it cannot be ruled out that arguments might be found for distinguishing these situations.
The robots.txt file is technically a voluntary standard dating from 1994. It acquired the formal status of an internet standard only in September 2022, when it was published as RFC 9309. Legal bindingness in the context of the automated analysis of text and data was conferred on it only by the implementation of the Directive on copyright in the Digital Single Market — but only as one of the forms of "reservation of rights in an appropriate manner." Among the principal artificial-intelligence crawlers that institutions block are GPTBot (OpenAI), ClaudeBot (Anthropic), CCBot (Common Crawl), Google-Extended (Google/Gemini), and PerplexityBot.
The problem lies in the fact that robots.txt is a binary tool — it blocks the entire website, not just copyright-protected content. If a municipal authority prohibits access for GPTBot across its whole website in robots.txt, it thereby blocks both press releases (where a reservation of rights may be legitimate) and ordinances and decisions (where the effectiveness of a reservation is at the very least questionable).
Official works (Section 3 of the Copyright Act): Not subject to copyright protection. A reservation of rights in robots.txt is probably without legal effect.
Other copyrighted works of the institution: Subject to copyright protection. A reservation of rights is formally possible under Section 39c, but questionable for publicly financed content.
Content for scientific research: Subject to copyright protection, but the law does not permit a reservation of rights (Section 39d). The eligible persons are universities, research organisations, and cultural heritage institutions.
Personal data: Addressed through the General Data Protection Regulation, not the Copyright Act. The robots.txt file is irrelevant to its protection.
Open data (Section 5a of Act No. 106/1999 Coll.): For official works, not subject to copyright protection. For other content, protection is limited. The principle of open data as the default standard applies.
EU Regulation 2024/1689 (the so-called Artificial Intelligence Act) brings a further dimension. Article 53(1)(c) imposes on providers of general-purpose artificial-intelligence models the duty to put in place a policy for respecting reservations of rights under EU copyright law. At the same time, Article 53(1)(d) requires these providers to publish a sufficiently detailed summary of the content used for training.
For public institutions, this means that the Artificial Intelligence Act reinforces the reservation-of-rights mechanism — but only where it can be effectively asserted, that is, for copyright-protected content. For official works, the Artificial Intelligence Act changes nothing.
The current practice, in which public institutions blanket-block artificial-intelligence crawlers in the robots.txt file, raises a number of legal questions to which there is as yet no unequivocal answer from case law or from settled legal doctrine. Nevertheless, several lines of argument can be identified suggesting that this practice may be problematic.
For official works, an interpretation suggests itself that a reservation of rights lacks legal effect, because these works are not subject to copyright protection. In such a case, the robots.txt file could not restrict access to content over which no exclusive proprietary rights exist. This interpretation has so far been neither confirmed nor refuted by the courts.
Blanket blocking is difficult to reconcile with the principle of the openness of public administration and with the principle of open data as the default standard, introduced by Directive 2019/1024 and transposed into Act No. 106/1999 Coll. Whether this tension is strong enough to render a blanket reservation of rights through robots.txt unlawful is a question for legal theory.
For the scientific automated analysis of text and data, the law expressly does not permit a reservation of rights (Section 39d), so blocking crawlers could be contrary to the law as against the eligible persons — universities, research organisations, and cultural heritage institutions. Yet here too an authoritative confirmation would be desirable.
A practical solution could be a detailed robots.txt configuration that would block crawlers only for those sections of the website containing copyright-protected, non-official content. This, however, requires institutions first to carry out an audit of their content and distinguish official works from other content — a task for which most of them are not yet prepared.
The topic deserves the attention of the legal community. As long as case law is lacking, we are operating in the realm of interpretation — and that interpretation may turn out variously depending on whether the emphasis on the transparency of public administration prevails, or the fear of uncontrolled use of content by artificial-intelligence systems.
Methodological note: The analysis is based on the wording of the Copyright Act (No. 121/2000 Coll.) in force as of February 2026, Act No. 106/1999 Coll. as amended by amendment No. 241/2022 Coll., the Directive on copyright in the Digital Single Market (2019/790), the Open Data Directive (2019/1024), and the Artificial Intelligence Regulation (2024/1689). From case law, the relevant decisions are that of the Hamburg District Court in the matter of Kneschke v. LAION of 27 September 2024 (case no. 310 O 227/23) and that of the Hamburg Higher Regional Court of 10 December 2025 (case no. 5 U 104/24), which dealt with the exception for the automated analysis of text and data in the context of training artificial intelligence. The article presents an analytical examination from the position of a technical expert, not a legal opinion.
Before making any decision on the basis of the information presented here, I recommend consulting an attorney specialising in intellectual property law.
Transparency of creation
The concept, structure, and editorial line of the article are the work of the author, who prepared the content outline, established the key theses, and directed the entire creative process. Generative AI (Claude Opus 4.6, Anthropic) was used as a tool for research, fact-checking, and elaborating the author's draft.
The author verified the key findings and approved the final wording. No part of the text was published without conscious authorial control. The factual data were verified against the publicly available sources cited in the text.
This procedure complies with the transparency principles of EU Regulation 2024/1689 (the AI Act). #poweredByAI
Read the Czech original on Médium.cz.
AI · Claude — machine translation, may contain inaccuracies.