Mind · the wisdom and madness of crowds

The Crowd That Watched Itself

A crowd can out-guess its own smartest member: and there is an exact equation for why. The unsettling part is that the same equation, read backwards, tells you precisely how a wise crowd becomes a confident, wrong one. What it spends is the one thing that made it wise.

In 1906, at a country fair in Plymouth, the statistician Francis Galton watched a weight-judging competition: pay sixpence, guess the dressed weight of a fat ox after slaughter, win a prize if you're close. Eight hundred or so people entered - butchers and farmers who knew cattle, but also clerks and shopkeepers who didn't. Galton collected the tickets afterward. He was 84, a believer in the rule of experts, and he expected the data to expose the foolishness of the common voter. He found the opposite.

Of the tickets, 787 were legible. Their middlemost guess: the median, which Galton chose deliberately as the fairest "one vote, one value" verdict of the crowd: was 1207 lb. The ox's dressed weight, as Galton printed it, was 1198 lb. The crowd, as a body, was wrong by nine pounds: three-quarters of one percent.

"This result is, I think, more creditable to the trustworthiness of a democratic judgment than might have been expected." - Francis Galton, "Vox Populi", Nature 75, 7 March 1907
Companion film: 2:35 The argument told as motion. Galton's 787 guesses, rebuilt live from his 1907 centiles (printed 1198 · median 1207 · mean ≈1197); then the Diversity Prediction Theorem balancing on screen to the decimal - crowd error = average individual error − diversity: as the members' disagreement is swept wide (the crowd stays nailed to the truth) and then given a shared bias (it slides off, and no diversity rescues it). The social-influence collapse runs the same model the page does: 260 estimates under DeGroot consensus averaging: until diversity → 0, confidence soars ×20,000+, and the truth falls outside the crowd. Condorcet closes it: 501 independent voters each 60% right are correct >99.99% of the time: flip them to 40% and the same crowd is wrong essentially always. Every figure is read live from the model the verifier runs, so the picture and the numbers cannot drift; the music is a minor-pentatonic bed whose percussion density is the collapse model's own diversity, emptying out as the crowd converges. Companion film for a Wasteland layer.

IThe ox

Galton published only the shape of the crowd: a table of its percentiles - not the raw tickets. Below is that distribution, rebuilt from his printed centiles (the full reconstruction is K. Wallis, Statistical Science 2014). Drop in your own guess and see where you land against the body of the crowd.

1000 lb1100120013001400 lb
the 787 guesses printed weight (1198) crowd median (1207) crowd mean (≈1197)

On Galton's printed figures the crowd's median missed by 9 lb; its mean by about 1. Almost no single person did that well. Place a guess and find out where you'd have stood.

Here is the first quiet correction. The number everyone repeats: that the crowd's guess was "right to within a pound": is the mean, about 1197 lb. But the mean is not the statistic Galton reported. He championed the median, 1207, as the true voice of the people, and he supplied the famous near-perfect average only three weeks later, in a reply letter to Nature ("The Ballot-Box", 28 March 1907), after a correspondent wrote in asking for the mean. The crowd was astonishing either way. It just wasn't astonishing in quite the way the story tells it.

The second correction took a century and an archive. Kenneth Wallis re-checked Galton's worksheets (Statistical Science, 2014) and found transcription slips in the printed article: the ranked middlemost estimate was actually 1208 lb, and the organiser's letter put the dressed weight at 1197 lb. So the median really missed by eleven pounds, not nine, and the mean missed by nothing at all. The numbers on this page keep to Galton's printed 1907 table, because that table is what the histogram above rebuilds; the archival figures are the ones to quote as history.

IIWhy it works: and it's not magic

A crowd's collective error is not usually smaller than its members'. It is always smaller: provably, by an exact identity. For any set of guesses and any true value:

The diversity equation

(crowd's error) = (average individual error) − (diversity of the guesses)

This is not a statistical tendency. It is an algebraic identity: the same parallel-axis decomposition that splits any variance: so it holds exactly, every time, to the last decimal. Scott Page named it the Diversity Prediction Theorem. Because diversity (the spread of the guesses) can never be negative, the crowd is never worse than its average member, and is strictly better the moment people disagree. Disagreement is not noise to be averaged away. Disagreement is the accuracy.

Build a crowd. Two dials: how much its members disagree (diversity), and how much they share a bias (everyone wrong in the same direction). Watch the three quantities in the equation, and watch them balance: always.

−100truth+100
0 = 0 + 0
avg indiv. error diversity crowd error
✓ balances exactly
the crowd beats
-
crowd is off by
-

Slide diversity up with bias at zero and the lesson is almost shocking: the individuals get worse and worse, yet the crowd stays nailed to the truth, and the fraction of members it beats climbs toward everyone. Diversity is free accuracy: but only the idiosyncratic kind, the errors that point in every direction and cancel. Now push shared bias up. The crowd's error climbs with it, and no amount of diversity rescues it. This is the honest boundary the cheerful version omits: a crowd cancels the part of the error its members don't share. The part they share: a rumour, a framing, a number everyone half-remembers: it averages straight into the answer and calls it consensus.

IIIThe madness: when the crowd watches itself

Galton judged his entries "unbiassed by passion and uninfluenced by oratory and the like", though he allowed that some competitors were "probably guided by such information as they might pick up". Independence was an assumption about that room, never a measured fact; what the result actually needed was errors that did not all lean the same way. So what happens when a crowd can see itself: when each person, sensibly, nudges their estimate toward the room?

In 2011, Jan Lorenz and colleagues sat 144 people down to estimate real quantities about their own country: things like the number of new immigrants in a year, or the length of the Swiss border: over several rounds, paying them for accuracy. Some saw nothing of others' answers; some saw the group's estimates and could revise. The result was precise and bleak: social influence made the estimates converge : the crowd's diversity collapsed: and made people far more confident. It did not make them more accurate. Worse: as the range of guesses shrank, the true value often ended up outside it. The crowd talked itself into a tight, certain, wrong consensus.

The diversity equation tells you exactly why, with no extra psychology needed. Pulling everyone toward the average preserves that average: it adds no information: but it destroys diversity. And diversity was the crowd's entire edge. Spend it, and the crowd's edge is gone: the collective error cannot improve, because the average never moved, while the average member's error falls to meet it. Everyone ends exactly as wrong as the crowd, and far more sure. The same identity that grants the wisdom revokes it.

Below: a crowd that starts independent and slightly biased: like real estimators of a large unknown number. Set how strongly each member copies the room, then drag time forward, round by round of "talking," and watch.

▲ truth
diversity (the wisdom)
-
crowd error
-
confidence
×1.0
truth inside the crowd?
yes

Turn influence to zero and the rounds do nothing: an independent crowd stays wise forever. Turn it up and drag time forward: the cloud contracts to a confident sliver, the confidence multiplier soars, the crowd error sits bolted to where it started (this averaging conserves the mean, so it cannot move), and at some round the green truth-marker slips outside the band entirely. Everyone now agrees. Everyone is now more sure than they have any right to be. And the thing they agreed on is no closer to true than where they started: they simply burned the disagreement that was doing the work.

IVThe binary cousin: voting, juries, and Condorcet

The same shape governs yes/no decisions, where it's older still. In 1785 the Marquis de Condorcet proved his jury theorem: if each voter is independent and right more often than not, a majority vote is more reliable than any single voter, and approaches certainty as the crowd grows. The numbers are stark: with each person right just 60% of the time:

voters (independent, each 60% right)majority is correct
160.0%
973.3%
3187.2%
10197.9%
501>99.99%

But flip each voter to below half: say a shared misconception makes everyone right only 40% of the time: and the theorem runs in reverse with equal force: a crowd of 501 such voters reaches the wrong verdict essentially every time (~0.0003% chance of being right). The lever is always independence. Correlate the voters: a pundit they all watch, a poll they all saw, a panic they all feel: and "more people" stops adding wisdom and starts amplifying whatever they share. Information cascades (Bikhchandani, Hirshleifer & Welch, 1992; confirmed in the lab by Anderson & Holt, 1997) are this failure at its sharpest: rational people, each watching the choices before them, can all rationally ignore their own private evidence and stampede: together, confidently: off a cliff.

VAsk yourself again

Living arm Extended 2026-08-21: the crowd page gains a crowd.

You will see one field whose number of dots is exact by construction. Estimate it, lock that estimate, then make a genuinely fresh second estimate without seeing the first one scored. Only after you approve the exact payload and send it does the page reveal the count.

Dot field, preparing

Checking the sealed analysis before opening the instrument.

The living crowd, Arm B

The arm opened 2026-08-21. Its state is being checked.

Each target is centered on its generated truth before the crowd median is taken. Different-reader pairs are irrevocable non-overlapping adjacent arrivals within the same target: first with second, third with fourth, and so on. The MSE values are descriptions of the rows so far, not fixed-sample estimates with ordinary intervals.

Anytime-valid means: the 95% confidence sequence covers its bounded mean at every viewing time simultaneously, under its assumptions. An ordinary fixed-n interval repeatedly refreshed under this live counter would not make that guarantee. The displayed win outcomes are bounded to 0 or 1: a strict improvement is 1, and a tie or loss is 0. The bankroll tests mean win rate 0.5 and preserves its running maximum. The evidence is therefore about the strict-win rate, not the MSE gap: a run of tied guesses alone can bankrupt the 50% null while every displayed MSE stays equal, so read the bankroll beside the equality rate and the MSE columns.

The published comparison, not a reproduction

Vul and Pashler's 2008 cohort had 428 internet-subject-pool participants answer eight percentage questions. The immediate group had 255 analyzed participants; the group that returned three weeks later had 173. Their interpolation valued a second immediate guess at 0.11 of a different person's opinion, about one tenth, and a second delayed guess at 0.32, about one third. Those are properties of those cohorts, questions, lags, and that interpolation. This one-item visual web arm does not estimate the same fraction.

2008 comparisonprinted t2014 reanalysis t
immediate: guess 1 vs average2.254.41
immediate: guess 2 vs average6.089.90
three weeks: guess 1 vs average3.946.22
three weeks: guess 2 vs average6.599.85
gain contrast, df 4262.122.68, p=.008

Steegen and colleagues recomputed the right column after receiving the original records privately. As of 2026-08-21, the scout could not find a public copy of the original 428 by 16 response matrix or verified redistribution permission. The public OSF files it found were the 2014 replication, not the 2008 records. This page therefore does not label the printed aggregates reproduced, does not silently substitute the replication, and does not compute the original interpolation.

Checks shown: the 2008 primary paper, pp. 645-647; Steegen et al.'s 2014 tables and correction note; the OSF data-component listing checked 2026-08-21; and Howard et al. 2021 for time-uniform confidence sequences.

The check for the new material

Why the truth is knowable. The sealed module deterministically shuffles a 20 by 12 grid for each of 16 target identifiers, chooses 96 to 220 cells, and places one non-overlapping dot in each chosen cell. The answer is the generated array length, not a hand-entered label. It remains hidden until a contribution succeeds.

Free choices, named. Sixteen targets, their seed, the 8-row display threshold, squared error, the median of signed first-guess errors, strict wins with ties counted as zero, same-target arrival-order pairing, a 401-point confidence-sequence grid, and a 0.05 error rate were fixed in analysis.mjs.

Selection and identity. These are self-selected visitors, not Galton's paying fairgoers and not Vul and Pashler's recruited pool. A browser mark and server rate limits discourage repeats but cannot establish one person per row. Coordinated arrivals, bots, shared devices, cleared storage, and population drift remain possible. Every valid stored row is included because the frozen payload contains no honest screening variables.

What small n means. Below 8 rows the page shows only the count. After that, the confidence sequences remain valid under continuous viewing, but they do not repair dependence, biased sampling, automation, a changing population, or target-specific difficulty. One immediate visual task cannot test a three-week delay, sleeping on a decision, a general inner probability distribution, or a universal benefit from asking twice.

No machine arm. The task is visual by construction. A text model would not receive the same evidence, so no machine population is forced into this comparison.

The seal. Expected SHA-256: 84886b881be92cdc241c7e77595f30201e8c688697f3ff7f25809b9e4407c9eb. checking the shipped bytes

Off-page check appears after the file is read.

VIWhat to actually take from this

The wisdom of crowds is real, and it is not a mystery or a miracle. It is one equation: a crowd's accuracy is its members' accuracy plus their diversity. That gives a short, usable recipe, and it is mostly about protecting the second term: