Philip E. Tetlock on Forecasting and Foraging as a Fox
Philip Tetlock — political psychologist at the University of Pennsylvania, creator of the Good Judgment Project and author of Superforecasting — joins Tyler Cowen in Ep. 93 to argue that accuracy is rarely the first thing we want from forecasters, that ordinary people can be trained to beat experts, and that the fox who knows many small things tends to predict better than the hedgehog who knows one big thing. Recorded on 26 March 2020, the conversation runs through prediction markets, the reform of the CIA, the trade between democracy and technocracy, and why counterfactual reasoning has long been the last refuge of ideological scoundrels.
Key ideas
- Foxes forecast better than hedgehogs — but hedgehogs win the prizes. Tetlock’s signature distinction, borrowed from Isaiah Berlin: the hedgehog views the world through one big organising idea; the fox holds many small, loosely connected models and tolerates contradiction. Foxes are the more accurate forecasters because they update, hedge, and synthesise opposing arguments. Yet Tetlock concedes that the great achievements in science usually come from hedgehogs — single-minded persistence, not many-sided caution, breaks new ground. He calls himself a fox who had to ‘commit career suicide’ to become one.
- Superforecasters are ordinary people who, trained and aggregated, beat the experts. Through the Good Judgment Project — run for IARPA, the intelligence community’s research arm — Tetlock found that a self-selecting pool of amateurs, scored on real geopolitical questions, outperformed credentialled analysts, sometimes including those with classified access. Their edge is temperament and method: relentless updating, breaking a hard question into tractable parts, and active perspective-taking. A statistical composite of their forecasts, extremised, beat 99.8% of the individuals it was built from.
- Scoring makes forecasting falsifiable — and almost no one submits to it. A forecast counts only when it carries a date and a probability, so it can be checked. The Brier score (a running tally of how far a forecaster’s stated probabilities sat from what actually happened — lower is better) turns punditry into something testable. There is no Tetlock score scrolling under the pundit on the evening news, so most public forecasting stays unaccountable; the high-status experts Tetlock invited to compete repeatedly declined, because for them a tournament offers, at best, a tie.
- Calibration and base rates are the craft, and machines do not yet own it. Calibration means your confidence matches reality — of the things you call 90% likely, about nine in ten should happen. A base rate is the background frequency of an event before you know anything specific (how often a struggling state actually collapses, say). For the intelligence community’s questions — how long the Syrian civil war will run, what Russia does in Ukraine — base rates are elusive and machine intelligence does not dominate as it does in chess or Go. Simple extrapolation is hard to beat; the skill is knowing when to depart from the trend.
- Cognitive diversity, value pluralism, and a leader who shuts up. The best composite forecast comes from people who normally disagree converging — a signal to extremise. Forecasters who hold conflicting values (‘value pluralism’) and tolerate the resulting dissonance reason in more integratively complex ways, and predict slightly better for it. To run such a team Tetlock would pick, after his teacher Irving Janis on groupthink, a leader ‘who knows how to shut up’ — eliciting independent judgements before conformity pressure sets in, then opening them to critique.
Content
Accuracy is only one of the things we want from a forecaster
Cowen opens by asking whether we want our forecasters accurate, or want them as a portfolio that keeps us alert to extreme events. Tetlock’s answer reframes the whole field: accuracy is ‘often not the first thing’. We look to forecasters for ideological reassurance, for entertainment, and for the minimising of regret — we pump up the probabilities of the disasters we would most hate to have ignored. On the coronavirus, he resists the idea that more vivid doom-mongers would have served us. The warning was already in the portfolio: epidemiologists had flagged SARS-like viruses in horseshoe bats and the exotic-meat trade as a time bomb for two decades. The problem was never that the forecast was missing; it was that scientists lose the competition for attention, and that crying wolf at a 30 or 40% annual chance of a viral leap, year after year it fails to land, eventually drains the warning of force.
This thread runs through the episode. When Cowen later asks what mistake the smart people are making about COVID-19, Tetlock declines to play amateur epidemiologist and instead points to a procedural confusion: a public-health expert who believes overestimating the threat is far safer than underestimating it is not forecasting at all. He is engaged in manipulation and social influence — a legitimate activity, but a different one, and it should not be scored as though it were a prediction.
The Good Judgment Project and the superforecasters
The substantive engine of Tetlock’s reputation is the forecasting tournament. IARPA — the intelligence community’s equivalent of a venture R&D lab, born of the reforms that followed the WMD-in-Iraq fiasco — ran tournaments in which teams put calibrated probabilities on geopolitical questions and were scored against what actually happened. Tetlock’s Good Judgment team won decisively, and within it a subset of amateurs, the superforecasters, beat the rest by a wide margin.
What distinguishes them is not raw brilliance but a skill set and a temperament. They will ‘delve into the details of really pretty obscure problems for very minimal compensation’ — intrinsically, cognitively motivated to an extraordinary degree. Tetlock is candid that he is not one of them. In the 2012 tournament he tried second-guessing the aggregation algorithm — which was already outperforming 99.8% of the forecasters feeding it — and by overriding the composite where he thought it too aggressive, he degraded his own accuracy and finished mid-pack. The lesson is double-edged: the statistical composite is formidable, and even the field’s founder cannot reliably beat it by hand.
Markets, machines, and the limits of automation
Cowen presses repeatedly on whether financial and betting markets are simply superforecasters already. Tetlock grants that arbitrage is superforecasting ‘in a form’ and that powerful actors like Goldman Sachs are continually doing exactly this. But he resists ‘best’ and ‘optimal’: prediction markets, in his own experiments with the intelligence community, tend to underperform tournaments — though that is no decisive rebuke, since the intelligence community will not let those markets run deep and liquid the way Wall Street’s do. On talent-spotting by venture capitalists, he reframes their success as a base-rate advantage: Silicon Valley’s talent pool is rich, so the needle is easier to find, and a few hits pay for many false positives.
On machine learning he is pointed. For credit-card screening, machine intelligence will dominate humans totally. For the questions the intelligence community actually cares about — the Syrian civil war, Russia and Ukraine, settlement on Mars — base rates are elusive and there is little evidence the machines are there yet. Hybrid human-machine forecasting ‘sounds like a great idea’ and ‘who could be against it?’, but the devil is in the details and it does not deliver automatically: ‘It’s trench warfare here.’ Against simple statistical algorithms, meanwhile, humans have been ‘repeatedly humbled’, a finding he traces back to Paul Meehl’s work on clinical versus actuarial prediction. His distilled advice: ‘Be humble.‘
Accountability that favours accuracy
Cowen sets a trap from Tetlock’s own early work, which found that making people accountable can breed evasion and self-deception — so might making pundits accountable just push them further from objective standards? Tetlock distinguishes types of accountability. A tournament is a ‘stark monistic’ regime where one thing and only one thing matters: accuracy. You earn no points for ideological cheerleading; pump up your team’s preferred probabilities and you take a reputational hit. That alignment is ‘extremely unusual in the social world’, where most accountability sits inside organisations full of distortions, and where the rational response is to drift towards the views of important others.
This explains who shows up. Tetlock recounts a 1980s correspondence with William Safire: the upwardly mobile young see tournaments as a chance to rise, while senior figures — the 65-year-old China analyst asked to compete on a level field against 25-year-olds — see only a way to lose, and move to nix the whole thing. Economists and sociologists react to tournaments differently and revealingly. The economist asks, ‘if these things are so great, how come they’re not everywhere?’ The sociologist asks why anyone would be naive enough to think an organisation would want one — they are status-disruptive.
Granularity, value pluralism, and integrative complexity
How many decimal places should a forecaster use? It depends on the game. For geopolitical questions, three or four decimals would be absurd; the National Intelligence Council moved from five degrees of uncertainty to seven, and Tetlock’s rounding experiments suggest the best forecasters can genuinely distinguish 10 to 15 — more than they think. Pseudo-precision is when moving a forecast from 0.6 to 0.65 does not, on average, improve accuracy; then the extra digit is noise.
Asking who gave the most cognitively complex speeches in the British House of Commons lets Tetlock unfold integrative complexity — the capacity to hold opposing considerations and synthesise them (‘on the one hand, on the other hand, and then synthesis’). In Bob Putnam’s interview data the centrists scored highest, slightly left of centre. The deeper driver, he argues, is value pluralism: the more your own values genuinely conflict, the more cognitive dissonance you must metabolise, and the more you are pushed into complex synthetic thinking rather than the easy outs of denial and bolstering. People raised across two cultures gain a related edge — a richer internal dialogue, better perspective-taking, which is itself central to superforecasting.
Reforming the CIA, and the trade between democracy and technocracy
Asked how he would reform the CIA, Tetlock locates the opening in history: the WMD fiasco forced the intelligence community to take score-keeping, training, and the monitoring of accuracy more seriously, and IARPA was the institutional result. He defends the aspiration to value-neutral analysis — ‘just the facts, ma’am’ accuracy, neither liberally nor conservatively skewed — while conceding Cowen’s cynical point that there is ‘no view from nowhere’: even a perfectly objective scoring system can be skewed by who chooses the questions. He folds in the politics: Michael Gove and Dominic Cummings both invoked his Expert Political Judgment during Brexit, weaponising its finding that subject-matter experts often fail to beat simple extrapolation, and he reads in this an ‘intellectual fissure’ between a largely liberal social science and a conservatism wary of its advice — a fissure he thinks helped slow the response to COVID-19.
Counterfactuals: the next project
The work Tetlock wants to dedicate years to is linking counterfactual reasoning about the past to conditional forecasting about the future. Counterfactuals matter enormously for drawing lessons from history and for policy argument, yet they are unresolvable — you cannot rerun history, so they become ‘a place where ideologues can retreat’ and make up whatever facts justify their priors (no matter how badly Iraq went, one can always claim Saddam would have been worse). His FOCUS programme uses simulations like Civilization V, where you can rerun history, to find people and methods that generate superior counterfactual forecasts against a known ground truth — discovering, for instance, that you still get something like World War I 37% of the time even after undoing the assassination of the archduke. The longer-term hope is to validate counterfactual beliefs indirectly, through their logical links to conditional forecasts in a Bayesian network, and to show that better counterfactual reasoners are less ideologically polarised. Closing on his own influence, Tetlock counsels caution: he is running against the psychological grain, because people treat their beliefs as ‘ego-defining, quasi-sacred possessions’ rather than falsifiable, probabilistic propositions.
Related
- Philip Tetlock — speaker; political psychologist, creator of the Good Judgment Project
- Tyler Cowen — host
- Annie Duke — speaker; decision strategist who applies probabilistic thinking to choices under uncertainty
- Annie Duke on Poker, Probabilities, and How We Make Decisions — kindred Conversations with Tyler treatment of calibrated, probabilistic decision-making
- Daniel Kahneman — coined ‘adversarial collaboration’; his camp and Gigerenzer’s collaborated, via Barb Mellers, on the conjunction fallacy
- Rationality — concept; falsifiable, probabilistic belief and the discipline of updating