When a Second Look Helps and When It's Just the Same Mistake Twice

Drawing on data from mammography, software inspections, and clinical trials, the article shows that what determines the usefulness of review (an appellate court, peer review, code review) is not the number of reviewers but the mutual independence of their errors. Parallel independent assessment adds up findings, whereas sequential chaining inherits the anchor of the previous verdict and is the weakest; moreover, reconciliation and arbitration mainly serve to remove false alarms rather than to uncover new mistakes. AI can be a valuable second reader only if its blind spots do not overlap with human ones.
Appellate courts, scientific peer review, code review, and the auditing of financial statements all rest on the same assumption: that an independent reviewer will reliably find the error. Data from mammography, software inspections, and clinical trials show that the outcome is not decided by the number of reviewers, but by one hard-to-secure property — their mutual independence.
The cohort of the British CO-OPS study comprised 805,206 women, and each of their mammograms was read independently by two radiologists. Where they disagreed, a third physician decided in arbitration. The second reader added 627 tumors that the first had missed — just under 9% of all those found. And arbitration simultaneously cut the rate of unnecessary recalls of women for further examination from 6.19% to 4.08%, that is, below the level achieved by reading by a single radiologist (4.76%). More cancer found and fewer false alarms at the same time; the data were published by Taylor-Phillips' team in Radiology in 2018.
Mammography is exceptional in that the usefulness of review is measured in lives saved and in the rate of unnecessary recalls. But we use the same architecture — someone does something, someone else reviews it — everywhere. An appellate court reviews a first-instance judgment. A scientific reviewer assesses someone else's study. A programmer reads a colleague's code. And in the last two years an exceptionally prolific source of mistakes that need checking has been added to this: text and code generated by artificial intelligence. The question of whether a second look even helps, and how to organize it, has thereby ceased to be academic.
The answer that science gives is uncomfortably precise. A second look does help, but far less and under far narrower conditions than we believe. What decides is not the number of reviewers; it is their mutual independence. And the topology we intuitively consider the most thorough — chaining, where a review is reviewed by a further review — is in fact the weakest of all.
One has to begin with why review is more fragile than it seems. Verifying someone else's work is not a weaker version of that work; it is a different and, in an important respect, harder task. To genuinely verify the result, the reviewer would have to reconstruct the entire problem the author worked through — but with less context, less time, and without the tacit knowledge of why each choice was made. In the extreme, the only way to truly verify something is to do it again. The reviewer therefore does the only thing within their power: they look for what looks wrong, not for what is wrong. They settle for plausibility. When the British Medical Journal inserted nine deliberate major errors into test manuscripts and sent them to 607 reviewers, they found on average 2.58 of them; training improved this only marginally and its effect soon faded (Schroter et al., 2008). The journal's former editor-in-chief, Richard Smith, summed it up by saying that peer review rests not on verification but on trust.
This is where the choice between topologies is decided. The sequential chain — author, first reviewer, second reviewer who sees the first one's verdict — inherits an anchor. The second reviewer does not read a clean original but the opinion of their predecessor, and their task imperceptibly shifts from the question "what is the correct answer?" to the question "is the person before me wrong?". Confirmation bias is, moreover, demonstrably stronger in a sequential arrangement than when a person receives the information all at once (Jonas et al., 2001). A randomized experiment with reviewers showed that a revised assessment stays closer to the original than an unbiased fresh look would be (Liu et al., 2024). The chain thus tends toward the intersection of errors that slip through both sieves — not toward their sum.
The parallel arrangement reverses that ratio. Two reviewers who assess the same original independently and blind to each other capture the sum of their findings. In the cleanest experiment on the screening of studies, dual independent assessment achieved a sensitivity of 97.5%, whereas a single assessor only 86.6%; alone they missed 13% of relevant works, the pair 3% (Gartlehner et al., 2020). In data extraction from studies, a single extraction made 21.7% more errors than dual independent extraction (p = 0.019), but was 36.1% faster (p = 0.003) — a trade-off between accuracy and time, not between accuracy and nothing (Buscemi et al., 2006). The logic is simple and merciless: if each of the pair finds 85% of detectable errors and their mistakes are independent, together they find almost 98%. If their mistakes overlap, they find barely more than one alone.
This brings us to the question with which this text began: is it worthwhile to add to a dual assessment a further phase in which the reviewers exchange their outputs and assess each other, with a third resolving any disagreement? The answer is yes, but not for the reason we would expect. Reconciliation and arbitration serve primarily not to find new errors — they serve to remove false alarms. In the Florence program, 1,217 discordant double readings went through arbitration. The arbiter sent 476 of them (39.2%) for further examination and thereby uncovered 30 tumors; the remaining 741 cases (60.8%) were dismissed, sparing those women an unnecessary examination — and of the 311 dismissed cases followed up so far, only two tumors appeared (0.64%). Thirty tumors caught against two missed: arbitration cut off most of the unnecessary recalls and paid for it with a minimum of escaped findings (Ciatto et al., 2005). And in software inspections, where this question has been measured the longest, an even harsher finding emerges. When Lawrence Votta in 1993 measured what an inspection meeting adds to the errors found during individual preparation, he came up with very little — and at the same time, in his view, meetings cost far more development time and developer time than anyone realizes. Replications confirmed it: a nominal team, whose members work separately and whose findings are simply summed, beats a real team with a meeting across all defect classes, because "meeting losses" — errors that one inspector found but that were not raised at the meeting — outweigh the "meeting gains" (Porter et al., 1995; Bianchi et al., 2001).
The strongest objection to this runs: surely discussion finds what an individual missed. And sometimes it indeed does. Multidisciplinary oncology boards change the original diagnosis or treatment plan in a considerable proportion of cases — a prospective study in gynecologic oncology recorded a change of plan in 27.1% and a change of diagnosis in 9.4% of patients (B. Lee et al., 2017), and an international ASCO survey reports changes to the treatment plan in 44 to 50% of breast and colon cancer cases (El Saghir et al., 2015). In forecasting research, in turn, teams in the Good Judgment Project beat both estimate-sharing and fully independent forecasters (Mellers et al., 2014). But these very cases reveal the condition under which discussion helps. An oncology board is not the same view twice; it is a radiologist, a pathologist, a surgeon, and an oncologist, each of whom brings information the others do not have. And the superforecaster teams were trained to share arguments, not conclusions, and to hold a high standard of evidence. Where reviewers share training and look at the same artifact with the same eyes — like software inspectors — a meeting consumes a coordination cost and brings no proportionate gain. A structured exchange of positions by the Delphi method proved, in a direct comparison, to be exactly as good as a face-to-face meeting, and on two questions out of ten even better (Graefe and Armstrong, 2011).
Through everything said so far, one and the same factor shines through. The usefulness of any review topology is bounded from above by how independent the reviewers' errors are. When their mistakes overlap, they do not cancel out. What destroys independence is well mapped: the visibility of the previous verdict, shared training and doctrine, and a hierarchy that forces the junior to defer to the senior. Moreover, groups do not restore independence on their own — in the classic hidden-profile experiment, only 18% of groups chose the best option when each member had different information, but 83% when everyone had the same, because groups preferentially discuss what they already share (Stasser and Titus, 1985). For the same reason, adding further reviewers yields steeply diminishing returns: in a software inspection two reviewers find almost as much as four, and the best configuration is two sessions of two (Porter et al., 1997). The outcome is decided more by the expertise of the individual than by the number of reviewers (Sauer et al., 2000).
It is worth adding a sober note that the literature search brought to light: more procedural rigor is not the same as more truth. When an independent committee in the ADVANCE clinical trial re-adjudicated 2,443 events reported by the investigators, it confirmed 2,077 of them (85%) and added a few dozen — but the estimate of the treatment effect barely moved (Hata et al., 2013). Adjudication here served defensibility and the audit trail, not a more accurate result.
And then there is a new player. When researchers at METR had sixteen experienced programmers work on their own projects with and without advanced AI tools, they were 19% slower with AI — although they believed they were 20% faster (METR, 2025). It is a snapshot from early 2025 on a small sample, and the team itself cautiously qualified it in February 2026, when newer measurements gave no clean signal; what remains, however, is that gap between perception and measurement — reviewing the output exacted a toll that the brain systematically underestimated. Microsoft Research, meanwhile, found among 319 knowledge workers that the higher the trust in AI, the less critical thinking (Lee et al., 2025). Yet the same technology can also be that rare second reader whose blind spots do not overlap with the human ones. In the evaluation of study screening, an automated workflow with a large language model beat a pair of human assessors (sensitivity 96.7% versus 81.7%; Bobrovitz et al., 2025), and in mammography, replacing the second human reader with artificial intelligence would reduce the radiologists' workload by 30 to 44.8% (Sharma et al., 2023). The condition, however, remains the same as with two humans: the benefit arises only if the machine's errors are genuinely different from the human's.
The cheapest second opinion is the one you force inside your own head — even a mere "consider the opposite" restores roughly half the gain that a real second person would bring (Herzog and Hertwig, 2009). For every editorial board, appellate panel, and review committee, then, there remains an uncomfortable test that the British radiologists passed: are your two readers really two, or is it the same trained eye looking a second time?
Citations are differentiated by strength. Primary = the original peer-reviewed study or official report; secondary = a summary or secondary reporting.
Note on method: the claim about software inspection meetings (Votta and subsequent work) is strong evidence against the meeting as a tool for error detection, not against the meeting as such — some authors credit it with value for knowledge transfer and training. The cost-effectiveness of double reading of mammograms is contested (Posso et al., 2016 versus Taylor-Phillips et al., 2018), and recommendations differ between European and American programs.
Transparency of creation:
The conception, structure, and editorial line of the article are the work of the author, who developed the content sketch, established the key theses, and directed the entire creative process. Generative AI (Claude, Anthropic) was used as a tool for research, locating primary sources, and the verbal elaboration of the author's content sketch.
The author edited the outputs continuously, verified the key findings, and approved the final wording. No part of the text was published without human review. All factual data were verified against the publicly available sources cited in the text.
The procedure complies with the requirements of Article 50 of EU Regulation 2024/1689 (the AI Act) on the transparency of AI-generated content. #poweredByAI
Read the Czech original on Médium.cz.
AI · Claude — machine translation, may contain inaccuracies.