The Hancock Paradox: Why We Write More Irony Where No One Will Recognize It

The article unpacks what it calls the Hancock paradox: although written text is demonstrably the weakest channel for conveying irony (people recognize it less well there, and even AI models like GPT-4 reach no more than 39.8% F1), we produce more irony in writing than in speech. Drawing on a game-theoretic model of indirect speech (Pinker), the concept of "strategic ambiguity," and the figure of Švejk, the author argues that textual ambiguity is not a flaw but a desirable feature — it enables deniability, where the "right" audience decodes the intent while the "wrong" one hears only the literal content. At the same time, it shows the dark sides of this mechanism (Poe's law, political "dog whistles," irony poisoning) and the limits of compensations like the winking emoji, which works only for some people and contexts.
In 2004, Jeffrey Hancock of Cornell ran an experiment that hinted at something strange about how we communicate online. He had participants converse in pairs — once face to face, once over text chat — and counted how often they used irony. The result ran against intuition. In text chat, which lacks tone of voice, facial expression, gesture — that is, everything on which the recognition of irony depends — participants did not use irony less. They used it more.
Hancock's finding was confirmed by Aguert and colleagues (2016) in adolescents. And Raymond W. Gibbs (2000), in an analysis of sixty-two conversations between friends, found that irony makes up roughly 8 % of all conversational turns — in ordinary face-to-face talk.
This gives rise to a paradox worth examining. Of all modalities, text communication is demonstrably the worst at conveying ironic intent. And yet we produce irony in text more abundantly than anywhere else. Either we are irrational — or the ambiguity of text is not a flaw we are trying to work around, but a property because of which we choose irony in text.
What text lacks
That text is a poor carrier of irony is not impressionism. It is an experimentally measurable deficit.
Bromberek-Dyzman, Jankowiak and Chełminiak (2021) tested Polish-English bilingual participants on the recognition of irony in three modalities: reading, listening, and audiovisual presentation. The text modality showed the lowest accuracy of the three. Even more interesting was an incidental finding: in text, participants responded to ironic and non-ironic statements at comparable speed. No hesitation, no cognitive signal that something was off. In the spoken modalities, by contrast, processing irony took longer — the brain detected the mismatch and invested cognitive resources in reinterpretation. In text, this process often did not occur at all. The reader simply missed the irony.
Why? Recognition of irony depends on detecting a mismatch between a statement and its context. In spoken language this mismatch is signaled by the voice (ironic tone, exaggerated intonation), facial expression (a raised eyebrow, a smile), and gestures. Cheang and Pell (2011) quantified the strength of the vocal signal alone in a cross-linguistic experiment with English and Cantonese: native English speakers recognized sarcasm from the voice in 85 % of cases. But when they heard a sarcastic tone in Cantonese, which they did not understand, accuracy dropped to near chance level. The voice carries the ironic signal reliably, but it needs linguistic experience as an anchor. Text has neither.
A large 2024 study in the Journal of Psycholinguistic Research (574 participants, six languages) confirmed the hierarchy: facial expression dominates over voice, voice dominates over text. In the purely textual modality the only remaining cue is context — and in online communication that is often insufficient, ambiguous, or entirely absent.
But perhaps the deficit is the point
This is where the story breaks. Why would we voluntarily choose a communication channel that systematically devalues our irony? The answer toward which several independent research lines converge is provocative: precisely for that reason.
In 2008, Steven Pinker, Martin A. Nowak and James J. Lee published in the Proceedings of the National Academy of Sciences a game-theoretic model of indirect speech — the category into which irony falls. Their key insight formalizes an intuition shared by everyone who has ever said something they did not mean and hoped the right person would understand it and the wrong person would not.
Consider a scenario the authors analyze in the article: a driver pulled over by a police officer wants to offer a bribe but does not know whether the officer is corrupt. Direct speech ("I'll give you money if you let me go") is catastrophic if the officer is honest. Silence is safe but brings no benefit. Indirect speech ("Couldn't we settle this here on the spot?") is the optimal strategy: a cooperative officer grasps the implication, an honest officer cannot react, because the literal content is innocent.
Irony works identically. The ironist says X and means not-X, relying on the "right" audience to decode the intent while the "wrong" audience hears only the literal content. The text medium amplifies this strategy, because the absence of prosodic cues makes deniability even more robust. No one can prove you meant it ironically — but the right people will know.
Eric Eisenberg (1984) named this dynamic "strategic ambiguity" and attributed four functions to it: unifying diverse interpretations, preserving privileged positions, enabling deniability, and facilitating change. In doing so he drew on Labov and Fanshel (1977), who put the basic thesis aptly: people need a form of discourse that is deniable in order to communicate — and if one did not exist, they would create it.
Švejk as an algorithm of deniability
The Czech reader does not need game theory to understand this mechanism. Švejk is enough.
Hašek's protagonist is a perfect implementation of strategic ambiguity in text. His ironic modus operandi consists in an idiotic, seemingly sincere agreement with authority that simultaneously subverts it — but whose subversive layer is unprovable. Švejk never admits the irony. Every utterance of his is compatible both with the literal interpretation (he really is stupid) and with the ironic one (he is a brilliant subversive). The reader oscillates between the two readings, and Hašek deliberately never resolves this oscillation.
This is not merely a literary discovery. It is a cultural code that Czechs have internalized to such a degree that Monty Python's Michael Palin declared that Czechs have "the best sense of humor in the world." The researchers Dewaele and colleagues (2021) quantified the cultural variability of irony: 68 % of Britons interpreted the courtesy phrase "with the greatest respect" as "I think you're an idiot," whereas 49 % of Americans took it literally. The British and Czech traditions share roots in dry, deniable humor — and both have run into the same problem: their ironic code is untransferable without a shared cultural context.
And that context is precisely what text systematically strips away.
When the strategy fails: Poe's law
Deniability, as described by both Pinker and Eisenberg, is useful as long as it works asymmetrically — the right audience decodes, the wrong one does not. But what if no one decodes? Or, conversely — what if everyone decodes, including those who were not supposed to?
In August 2005, a user named Nathan Poe wrote, on the Christian discussion forum christianforums.com, a sentence that became one of the few internet memes with epistemological status: without an overt signal of humor, it is impossible to parody an extreme position in such a way that someone will not take it as seriously meant. Poe's law was originally limited to creationism, but it soon turned out to apply universally.
Aikin (2009) subjected it to the first academic analysis and demonstrated that it represents a legitimate problem in the epistemology of online discourse — not a joke about a joke, but a structural property of the text medium. Historically, however, Poe was not the first to point out the problem. As early as 1983, Jerry Schwarz warned on Usenet: avoid sarcasm and facetious remarks, because without vocal intonation they are easily misunderstood.
The consequences reach further than harmless misunderstandings. A study in Communication Theory (2024) documents how politicians exploit polysemous messages with multiple plausible interpretations for different audience segments — so-called "dog whistles" with one meaning for the mainstream and another for radicals. Irony, which was meant to be a tool of subversion from below (as in Švejk), has become a tool of manipulation from above. And the phenomenon of "irony poisoning" — a state in which layered irony erodes the ability to distinguish sincerity from parody — is a direct consequence of the dynamic that Poe's law names.
39.8 %: AI as the perfect literal reader
If text makes irony hard for humans to recognize, how do machines fare?
The answer is offered by the study SarcasmBench (Zhang et al., 2024), to date the most comprehensive comparison of eleven large language models and eight pre-trained models across six benchmarks. And also by the iSarcasm dataset (Oprea & Magdy, 2020), unique in that the labels come from the tweet authors themselves — that is, from people who know whether they meant it ironically.
GPT-4 in zero-shot mode on the iSarcasm dataset achieved an F1 score of 39.8 %. With few-shot examples it improved to 52.3 %, but still lagged far behind human raters. For comparison: BERT fine-tuned on the large SARC corpus (1.3 million Reddit comments self-tagged with /s) reached an F1 of 91.7 % — but with massive training context and metadata, not from plain text without a hint.
The most remarkable finding, however, concerns chain-of-thought prompting. The technique that improves LLM performance on most cognitive tasks (breaking a problem down into steps) worsens performance on sarcasm detection. Data from SarcasmBench show that the Qwen 2 model recorded a drop in F1 of 25.1 %. This finding corresponds with the interpretation of the SarcasmCue study (Yao et al., AAAI 2025): irony detection is a holistic and non-analytic cognitive process that resists analytical decomposition. Irony cannot be "explained step by step," because its essence lies in the simultaneous grasp of context, tone, and intent — that is, exactly what text lacks and what analytical decomposition cannot replicate.
The winking smiley: a patch that half works
Compensatory mechanisms exist. But they work only partially.
Filik and colleagues (2016), in an experiment with 192 participants, found that in an unambiguous context emoticons have no significant effect on the recognition of irony — context alone is enough. In an ambiguous context, however, a winking emoticon (;-)) significantly increased the likelihood of a sarcastic interpretation. Thompson and Filik (2016) confirmed that the tongue-out (:p) and winking emoticons are the principal textual markers of sarcastic intent.
A breakthrough finding came from Weissman and Tanner (2018): an ironic emoji elicits in the brain the same electrophysiological responses (the P200 and P600 components) as verbal irony. The brain processes an ironic emoji by the same neural mechanism as ironic words.
But: Howman and Filik (2020), in an eye-tracking study, showed that a winking emoticon increases sarcastic interpretation only in younger adults (18–30 years), not in older ones (65+). On Reddit a system of tone indicators arose (/s for sarcasm), but it is used chiefly for socio-moral topics, where a misunderstanding would have serious consequences — not universally. The emoji is a patch that works for some people, in some contexts, for some of the time. It is not a universal solution.
Where the truth is more complicated
It would be convenient to tell a story about a medium that cannot do irony. But the data are more nuanced.
First, context helps more than isolated experiments would suggest. In real conversations between friends, colleagues, or members of a community, the shared context is rich — participants know the other party's humor, know what is "typical" ironic behavior for their counterpart, and recognize allusions to shared experiences. Experiments typically test the recognition of irony between strangers in artificial conditions, which systematically underestimates baseline accuracy.
Second, Hancock's data show not only more irony in online communication, but also that participants recognized most of the irony in chat — just not reliably. The problem is not binary (text cannot do irony) but gradual (text conveys irony less reliably than voice and facial expression). That is an important distinction.
Third, Poe's law itself has limits. It holds most strongly for extreme positions and anonymous communication. In communities with a shared context — from corporate Slacks to family WhatsApp groups — the ironic code is functional, because its decoding depends on the relationship, not on the medium.
Conclusion: the paradox that remains
Hancock's paradox does not dissolve even when we add nuance. Text communication is demonstrably the weakest channel for conveying ironic intent — and yet in it we produce irony more abundantly than anywhere else.
The explanation is less absurd than it seems. We produce irony in text not in spite of its ambiguity, but thanks to it. Deniability, strategic ambiguity, the ability to say and at the same time not say — these are properties that spoken language, with its prosodic signals, offers far less. Švejk would not work in a podcast. He needs text, where his idiotic agreement remains undecipherable.
The winking smiley is so far the closest humanity has come to a textual tone of voice. It works — but only for some people, in some contexts, and not for pensioners. And perhaps that is as it should be. Perhaps irony that can be signaled by a smiley ceases to be irony. Perhaps its power lies precisely in the fact that no one can identify it with certainty — including the billion-dollar machines that toss a coin over it.
The conception, structure, and editorial line of the article are the work of the author, who prepared the content sketch, set the key theses, and directed the entire creative process. Generative AI (Claude, Anthropic) was used as a tool for research, fact-checking, and elaborating the author's outline.
The author edited the outputs continuously, verified the key findings, and approved the final wording. No part of the text was published without human review. All factual data were verified against the publicly available sources cited in the text.
The procedure complies with the transparency requirements for AI-generated content under Art. 50 of EU Regulation 2024/1689 (AI Act). #poweredByAI
Read the Czech original on Médium.cz.
AI · Claude — machine translation, may contain inaccuracies.