Poisoned Web: 250 Documents Are Enough to Manipulate Artificial Intelligence

The article documents how the outputs of language models can be manipulated using a minimal amount of poisoned data – a mere 250 documents are enough to plant a backdoor in a model with 13 billion parameters. It describes the real-world consequences of contaminated data, from false court citations generated by ChatGPT (fines of up to $31,100), through the degradation of legal databases' free content, to the targeted cloaking of websites for AI crawlers that serves them disinformation.
In June 2023, New York attorney Steven Schwartz stood before federal judge P. Kevin Castel and explained why his filing in the case of Mata v. Avianca cited six court decisions that had never existed. Schwartz had generated them using ChatGPT. In court he testified that it had not even occurred to him that the chatbot might be "making up cases on its own" — he had assumed it found them in legal databases. That was not true. ChatGPT had fabricated them from beginning to end, including the names of judges, the dates of the decisions, and the legal arguments. The sanction was $5,000, but the real damage lay elsewhere: in trust in artificial intelligence as a research tool.
Schwartz was neither the first nor the last. The French lawyer and researcher Damien Charlotin, who works as a research fellow in legal data at HEC Paris, recorded in his database, as of early 2026, more than a thousand court decisions from around the world in which fabricated outputs of artificial intelligence caused procedural complications. Since mid-2025, new cases have been added practically every day.
But fabricated citations are merely a symptom. The real problem is deeper and more systemic: the information environment from which language models draw is increasingly contaminated — by degraded free content from commercial portals, by machine-generated text passed off as human work, and by targeted manipulation aimed directly at AI crawlers. Academic research shows that just 250 poisoned documents — 0.00016% of the training data — are enough to create a backdoor in a language model with 13 billion parameters. And commercial knowledge portals systematically serve bots and non-paying users measurably worse content than they serve paying customers.
This is not an academic curiosity. For anyone building a search engine or a knowledge system with artificial-intelligence components, it is a fundamental challenge.
The notion that poisoning the training data of a large language model requires a massive intervention does not hold up. The study by Nicholas Carlini and colleagues, "Poisoning Web-Scale Training Datasets is Practical" (IEEE Symposium on Security and Privacy, 2024), demonstrated that for just $60 one can contaminate 0.01% of the LAION-400M or COYO-700M datasets. The researchers used two techniques: so-called split-view poisoning, which exploits the variability of internet content over time, and frontrunning poisoning — timing malicious edits just before a snapshot of the dataset is taken.
Even more alarming results came from what is so far the largest study of data poisoning in language models. Alexandra Souly and colleagues from the UK AI Security Institute, Anthropic, and the Alan Turing Institute (arXiv, October 2025) tested models ranging from 600 million to 13 billion parameters and found that roughly 250 poisoned documents — about 420,000 tokens, that is, 0.00016% of the training data — are enough to create a functioning backdoor. The key finding: this number practically does not change with the size of the model or the dataset. A larger model does not require proportionally more poisoned data. The study's authors concluded that language models may be more vulnerable to data poisoning than previously assumed.
In the medical context, the problem is even more pressing. A study published in Nature Medicine (January 2025) showed that replacing just 0.001% of the training tokens with medical misinformation — hidden in invisible HTML tags — led to the production of potentially harmful answers. Moreover, the poisoned models exhibited comparable performance on standard benchmarks. The poisoning was practically invisible to ordinary evaluations.
In parallel with direct poisoning, there exists a subtler but equally devastating phenomenon: model collapse. Ilia Shumailov and colleagues, in Nature (July 2024), formally described how repeated training on synthetic data — that is, on the outputs of previous generations of models — causes irreversible defects. Minority distributions disappear first, while overall performance can seemingly even improve. The Stanford AI Index 2025 documents that the share of restricted tokens in the C4 (Common Crawl) dataset rose from 5–7% to 20–33% between 2023 and 2024 — a reflection of growing restrictions on access to quality content.
The problem of poisoned sources manifests most sharply where an error has legal consequences. And it is precisely in legal databases that a measurable, empirically proven gap exists between paid and free content.
The most thorough comparison was provided by Professor Susan Nevelow Mart of the University of Colorado Law School (ABA Journal, March 2018). Westlaw achieved 67% relevance in the first ten results, with roughly a third of the results being relevant and also unique. Lexis Advance achieved 57% relevance with about 20% relevant uniques. The free alternatives — Casetext, Fastcase, Google Scholar, and Ravel — averaged around 40% relevance with about 12% relevant unique results.
But the difference does not lie only in search accuracy. The crucial difference is in the expert layers that free databases lack entirely. The West Key Number System encompasses over 450 topics and more than 100,000 subtopics created by human editors for every published case. Legal summaries of key points (headnotes) exist exclusively on paid platforms. And the most fundamental difference is represented by citation tools: Shepard's Citations (LexisNexis) and KeyCite (Westlaw) make it possible to verify whether a cited case is still good law — a red flag means an overturned precedent, yellow a distinguished one. On Google Scholar, Justia, and CourtListener, no such tool exists.
CourtListener, the most comprehensive free database, operated by the nonprofit Free Law Project with more than 10 million decisions, likewise has no legal summaries, no systematic classification, and no formal citation tool. Google Scholar covers appellate and supreme courts of the U.S. states from 1950 and federal courts from 1923, but most trial-court decisions, unpublished decisions, and all secondary sources are missing.
For an artificial-intelligence system or a knowledge architecture, this means that legal research based exclusively on free sources systematically lacks information about whether cited precedents have been overturned, distinguished, or criticized. It is precisely this type of error that lies behind a series of court sanctions.
Since Schwartz's $5,000 in 2023, sanctions for fabricated AI citations have been escalating. In the case of Park v. Kim (2nd Cir., January 2024), the appellate court referred attorney Jae S. Lee to a disciplinary panel. In California, the penalty for 21 of 23 fabricated citations reached $10,000 (Noland v. Land of the Free, September 2025).
The highest documented sanction so far — $31,100 — was handed down in May 2025 in the case of Lacey v. State Farm, where attorneys from the firms Ellis George Cipollone and K&L Gates used a combination of the tools CoCounsel, Westlaw Precision, and Google Gemini. Not even three different AI-using tools prevented 9 of 27 citations from being incorrect and two cases from being entirely nonexistent.
The most serious consequence to date was a default judgment in the case of Flycatcher Corp. v. Affable Avenue LLC (S.D.N.Y., February 2026). The filing by attorney Steven Feldman contained at least 13 nonexistent cases and 8 real cases with fabricated quotations. The client lost the dispute not on the merits, but because of his lawyer's false citations.
The problem, moreover, is not confined to American courts. In December 2025 the Czech Constitutional Court imposed a procedural fine of CZK 25,000 on the Prague attorney Pavol Kehl for a constitutional complaint evidently drafted by artificial intelligence (file no. I. ÚS 3004/25). The filing contained a total of 12 decisions of the Constitutional Court and the European Court of Human Rights, a substantial part of which did not exist at all and the rest of which were grossly misinterpreted. The presiding judge, Tomáš Langášek, stated in the reasoning that attorneys, as professionals versed in the law, assume full responsibility for a filing, and thus are also responsible for any "hallucinating argumentation of artificial intelligence." Nor was this an isolated case — the Supreme Administrative Court had already, in October 2025 (file no. 3 As 34/2025), cut the reimbursement of litigation costs for the procedurally successful party's attorney precisely with reference to nonexistent judgments: "Whoever or whatever wrote the submissions in question, the court did not consider them, taken as a whole, to be of any benefit to the matter."
The Stanford RegLab/HAI study — the first pre-registered empirical evaluation of AI legal tools (22 J. Empirical Legal Stud. 216, 2025) — measured the rates of fabricated outputs even in specialized legal systems. The tested tools from LexisNexis and Thomson Reuters generated false information in 17 to 33% of queries, with Lexis+ AI representing the lower and Westlaw AI-Assisted Research the upper bound of this range. An earlier Stanford study on general chatbots found fabrication rates of 58–82% for legal queries. The providers' marketing claims of outputs entirely free of fabrications proved to be exaggerated.
The platforms whose content shaped the first generation of language models are undergoing their own quality crisis.
Stack Overflow has recorded a decline in monthly questions of roughly 76% since the launch of ChatGPT — from a peak of 200,000 down to levels close to those of 2009, when the platform started. A peer-reviewed study in Nature/Scientific Reports (2024) estimated, using synthetic-control methods, a decline in daily traffic of about a million (roughly 12%). In June 2023, a moderators' strike broke out after management introduced a policy that effectively forbade them from deleting machine-generated content. The moderators argued that this quietly enabled the spread of incorrect information and unlimited plagiarism.
Wikipedia faces similar pressure. A study from Princeton (October 2024) estimated that approximately 5% of newly created articles in August 2024 were generated by artificial intelligence. The WikiProject AI Cleanup project flagged nearly 3,000 suspicious articles. In August 2025, Wikipedia created a speedy-deletion criterion (G15) targeting machine-generated articles. Editors identify such content by fabricated citations, the excessive use of words like "moreover" and the em dash, and phrases of the type "Here is your Wikipedia article on."
Reddit has become the most-cited source in artificial-intelligence models — three times more often than Wikipedia (according to a Profound AI analysis). The platform has concluded licensing agreements with Google ($60 million per year) and OpenAI (roughly $70 million per year). Total disclosed licensing contracts for AI data reached $203 million before its stock-market listing. At the same time, services such as ReplyGuy use AI bots to automatically post promotional content in relevant discussions — a parasitic way of exploiting Reddit's high ranking in Google's results. Steve Huffman, Reddit's chief executive, confirmed to the Financial Times in June 2025 that companies use AI bots to create fake posts with the goal of having their content picked up by chatbots.
The outlet 404 Media documented this phenomenon as a closed loop: companies create machine-generated content on Reddit, Google indexes it highly, chatbots cite it, and users regard it as authentic.
Cases involving Google AI Overviews demonstrate the direct line from a degraded source to a mass-distributed error. The recommendation to add "an eighth of a cup of non-toxic glue" to pizza dough came from an 11-year-old joke comment on Reddit. The claim that "geologists recommend eating at least one small rock per day" came from a satirical article in The Onion. And then a self-referential loop set in: after these errors were written about in the media, AI Overviews began citing articles about its own mistakes as source material.
Perplexity AI represents a special case. In 2025 Cloudflare documented that Perplexity uses undeclared crawlers posing as the Google Chrome browser on macOS to circumvent robots.txt files. The service GPTZero identified a growing number of sources cited by Perplexity that are themselves machine-generated — including medical blogs with contradictory pharmaceutical information. NewsGuard (August 2025) measured that the misinformation rate in Perplexity's answers on current events rose from 0% to 46.67%.
In October 2025 the security firm SPLX demonstrated a technique that shifts the problem from passive degradation to active manipulation. They created a website for a fictitious designer, Zerphina Quortane: a human visitor saw a professional portfolio, but an artificial-intelligence crawler — identified by the User-Agent header — received entirely different content, which described her as a "notorious product saboteur." ChatGPT and other tools faithfully reproduced the poisoned story.
In a second experiment, a manipulated candidate résumé served the purpose: after detecting an artificial-intelligence crawler, the server served a version with inflated credentials and invented achievements, which completely changed the ranking of the candidates in the assessment. SPLX's conclusion: a single conditional rule on a web server — "if the requester is an artificial-intelligence crawler, display a different page" — can shape what millions of users see as a trustworthy output.
The technique is called agent-aware cloaking, and its execution is trivial. The web server compares the request's User-Agent header against known artificial-intelligence crawlers (GPTBot, ClaudeBot, ChatGPT-User, PerplexityBot) and serves them different content. As of early 2026, the technical defenses of companies developing artificial intelligence focus on the opposite problem — how to prevent content operators from blocking their crawlers — while the problem of deliberately serving false content specifically to artificial intelligence remains largely unaddressed.
The blocking of artificial-intelligence crawlers has meanwhile reached mass adoption. As of December 2025, GPTBot is blocked by approximately 5.6 million websites (an increase of about 70% since July 2025), and ClaudeBot by 5.8 million. According to a BuzzStream study (2025), 79% of leading news websites block at least one AI training bot, and 71% block bots used for search. Traffic from GPTBot grew by 305% between May 2024 and 2025 (Cloudflare data).
In March 2025 Cloudflare deployed the so-called AI Labyrinth — instead of blocking unauthorized crawlers, it serves them machine-generated fake pages designed to waste their resources. Any visitor who passes through four levels of generated links is almost certainly a bot. In July 2025 Cloudflare switched to blocking artificial-intelligence crawlers by default for new domains — the first infrastructure provider to take such a step.
The disproportion between the volume of downloading and the return traffic is telling: Google downloads content at a ratio of roughly 14 pages for every link back to the source. OpenAI operates at 1,700:1 and Anthropic at 73,000:1 (Cloudflare data, June 2025). Companies developing artificial intelligence consume orders of magnitude more content than the traffic they return to publishers.
The llms.txt standard, proposed by Jeremy Howard (Answer.AI) in September 2024 as a curated Markdown-format file for more efficient use of content by language models, is in an early stage of adoption. Cloudflare, Hugging Face, and Anthropic support it, but no major language-model provider has officially confirmed using it. Google's John Mueller likened it to the keywords meta tag — a historically unsuccessful attempt to control indexing.
The EU Regulation on Artificial Intelligence (the AI Act, Regulation 2024/1689, in force since 1 August 2024) constitutes the most detailed global regulatory framework for the quality of training data. Article 10 requires that datasets for high-risk artificial-intelligence systems be "relevant, sufficiently representative, and to the best extent possible free of errors and complete." Datasets must take into account the geographical, contextual, and behavioral characteristics of the deployment environment. For general-purpose AI models, Article 53(1)(d) requires the publication of a public summary of the training data — the European Commission's mandatory template took effect on 2 August 2025. Penalties for non-compliance: up to EUR 15 million or 3% of global annual turnover.
The General Data Protection Regulation (GDPR) complements the AI Regulation with the principle of data accuracy (Art. 5(1)(d)) — personal data must be factually correct and up to date, which also feeds into the requirements on training-data quality. The academic Philipp Hacker (European University Viadrina) argues that AI training data is currently inadequately captured by EU law, and identifies three risks: qualitative, discriminatory, and innovation-related.
European legal databases suffer from asymmetries similar to the American ones. BAILII (the British and Irish Legal Information Institute), with 14.4 million visits per year, is the most popular British free source, but it lacks legal summaries, a citation tool, and — crucially — its terms of service explicitly prohibit bulk downloading and data mining. Westlaw UK covers British case law from 1220. The content collections of Westlaw UK and Lexis+ UK largely do not overlap, so full coverage requires access to both.
In Germany, the paid databases Juris (a collaboration with the Ministry of Justice) and Beck-online (C.H. Beck) dominate. The free Gesetze im Internet provides almost all federal statutes in force with daily updates, but with an important caveat: the consolidated legal texts available online are not the official version. In France, the Légifrance portal offers free access to constitutional texts and decisions of the highest courts, but the official gazette (Journal officiel) is not complete before 1990, and the legislation on Légifrance is not annotated — for annotated versions, Dalloz or LexisNexis France is required.
All the individual phenomena — degraded free sources, machine-generated content, cloaking, crawler blocking — together create a self-reinforcing cycle.
Scientific publishers — Elsevier, Springer Nature, Wiley, Taylor & Francis, and SAGE control roughly half of the peer-reviewed literature — began in November 2024 to restrict access to abstracts on platforms such as OpenAlex. Abstract coverage fell from about 80% to below 40%, and for Elsevier articles from 2022–2024 even to 22.5%. The publishers recognized the value of abstracts for AI training and launched negotiations on licensing agreements. Elsevier launched LeapSpace (2025) — a paid AI tool accessing 18 million full-text paywalled articles. Free AI tools are restricted predominantly to openly accessible articles.
The German study "Is Google Getting Worse?" (Leipzig University, Bauhaus-Universität Weimar, 2024), in a year-long investigation of 7,392 unique product queries, proved that higher-ranked pages are on average more optimized for search engines, more saturated with affiliate marketing, and show signs of lower textual quality. Algorithm updates have an "observable but short-lived effect" — it is an arms race. NewsGuard tracks the growth of AI-generated "news" websites from 49 to 1,271 between May 2023 and 2025.
The study "Retrieval Collapses When AI Pollutes the Web" (arXiv 2602.16136, 2025) formalized this phenomenon as a two-stage degradation. In the first stage, synthetic content optimized for search engines dominates and homogenizes search results. In the second stage, the integrity of information degrades. At 67% contamination of the source pool, users' exposure to contaminated content reaches over 80%.
The feedback loop is accelerating. Surveys by Originality.ai show that the share of machine-generated content in Google search results is continuously rising — from approximately 11% in the summer of 2024 to nearly 19% by mid-2025.
For knowledge systems with artificial-intelligence components, the study by Barnett and colleagues, "Seven Failure Points When Engineering a RAG System" (2024), identifies seven failure modes. The most relevant to the problem of degraded sources: missing content in the knowledge base, documents out of context, and incomplete answers. The PoisonedRAG framework demonstrated that just 5 poisoned texts in a database of millions can achieve a 90% attack success rate.
For architects of knowledge systems and search engines, several converging approaches exist.
The corrective approach CRAG (Corrective RAG) adds a lightweight classifier — a fine-tuned T5-large model with 770 million parameters — that rates documents as reliable, unreliable, or ambiguous before passing them to the generator. When quality is low, it automatically switches to alternative sources. The RA-RAG approach (Reliability-Aware RAG, presented at ICLR 2025) proposes iterative estimation of source reliability and selective retrieval with weighted majority voting — without the need for manually labeled data. The RAGuard method (2025) uses an expanded retrieval range to dilute poisoned texts, filtering by span-level probability volatility and filtering by textual similarity.
The differential-downloading technique — downloading the same content with different user-agent identifiers and IP addresses and then comparing the results — can reveal cloaking. Multi-agent verification, where one agent downloads and another independently verifies, addresses the problem of dependence on a single source.
At the standards level, the C2PA coalition (Coalition for Content Provenance and Authenticity) — led, among others, by Microsoft, Adobe, Intel, BBC, OpenAI, Google, Meta, and Amazon — is developing an open standard for cryptographically signed metadata that travels with digital content. Content Credentials function as a "nutrition label for digital content," based on SHA-256 cryptographic fingerprints and X.509 certificates. The standard is at version 2.3, with post-quantum cryptography planned.
At the regulatory level, the NIST AI Risk Management Framework 1.0 (January 2023) emphasizes data quality and provenance as key risk indicators. In February 2026, NIST launched the AI Agent Standards initiative for industry standards for AI agents. The OWASP Top 10 for language models identifies data and model poisoning and prompt injection as critical vulnerabilities.
In 2023, Schwartz made a mistake that seemed isolated — he blindly trusted a chatbot's output for legal citations. Three years later, this is no longer the problem of an individual. It is a systemic property of an information environment in which economic incentives — paid access, freemium models, search-engine optimization of content — systematically degrade freely available content, while artificial-intelligence systems draw on this degraded content and distribute its errors to billions of users.
The data shows a clear trend: the share of machine-generated content in search results has roughly doubled over two years. The feedback loop is accelerating.
For architects of search engines and knowledge systems, three conclusions follow. First: single-source retrieval is unreliable — every system must have cross-validation from multiple independent sources, with thoughtful reliability modeling. Second: detection of differential content serving must be part of the search processing chain — the agent-aware cloaking technique is trivially simple, and there is no technical barrier to its mass deployment. Third: content-provenance metadata (C2PA) represents the only scalable path to verifiable content, but its adoption is in an early stage.
The deepest challenge is fundamentally economic. The highest-quality human-created content is being pulled behind paid access. The open web is filling up with synthetic and degraded content. And artificial-intelligence systems — dependent on the open web — are becoming a means of spreading this degraded content. Without a systemic solution at the level of data licensing, regulatory requirements, and technical standards, the quality of AI outputs will continue to diverge from the quality of real knowledge.
Schwartz paid $5,000. Feldman's client lost the entire dispute. The question is not whether the problem will worsen, but how fast — and who will pay for it next time.
This article is based on research conducted in March 2026 using web search and the analysis of academic sources, court decisions, and industry reports. Primary sources include: peer-reviewed studies (arXiv, Nature, Nature Medicine, IEEE S&P, J. Empirical Legal Studies), court decisions (S.D.N.Y., 2nd Circuit, 5th Circuit, California Court of Appeal, the Czech Constitutional Court, the Supreme Administrative Court of the Czech Republic), industry reports (Stanford HAI AI Index, the Cloudflare blog, the BuzzStream study), regulatory texts (the EU Regulation on Artificial Intelligence, the NIST AI RMF), and investigative journalism (404 Media, NPR, Gizmodo, Česká justice).
Limitations: The Mart study (2018) is the most recent thorough comparison of legal databases, but it dates from before the massive integration of artificial intelligence into legal tools. The figures on the blocking of artificial-intelligence crawlers are changing rapidly. The platform-degradation statistics (Stack Overflow, Wikipedia, Reddit) come from various methodologies and periods. The report on the fabrication rates of AI legal tools (Stanford RegLab) tested the versions available at the time of the study; current versions may differ.
Transparency of creation
The conception, structure, and editorial line of the article are the work of the author, who prepared the content outline, established the key theses, and directed the entire creation process. Generative AI (Claude Opus 4.6, Anthropic) was used as a tool for research, fact-checking, and fleshing out the author's draft.
The author verified the key findings and approved the final wording. No part of the text was published without conscious authorial oversight. The factual data were verified against the publicly available sources cited in the text.
The procedure complies with the transparency principles of EU Regulation 2024/1689 (the AI Act). #poweredByAI
Read the Czech original on Médium.cz.
AI · Claude — machine translation, may contain inaccuracies.