Survivorship bias
This article is also available on Spotify as an extended audio version. Spotify · Apple Podcasts
Survivorship bias is one of the most deceptive statistical distortions. It occurs when we analyze only the cases that survived, succeeded, or reached us, while ignoring those that disappeared from the data entirely. The result is conclusions that may sound logical, but are actually wrong. The following examples show how easily this bias can arise in war, science, animals, and also in sport and education.
When more injuries mean fewer deaths
One of the less known but very revealing examples of survivorship bias is the introduction of more modern military helmets during World War II. At a certain stage of the conflict, military commanders noticed an apparently alarming trend. After new, higher quality helmets were introduced, the number of injured soldiers transported to field and rear hospitals began to rise sharply. This effect raised doubts among commanders about the effectiveness of the new protective equipment. From a purely intuitive perspective, it looked as if the new helmets were failing because they were producing more injured soldiers. There were even discussions about returning to older helmet designs or limiting the use of the new ones.
The core problem, however, was the way the data were evaluated. The analysis relied almost exclusively on data from military hospitals and medical stations. These data included only soldiers who survived long enough to be evacuated and treated. In contrast, data on those killed in action were not systematically included, even though other parts of the military recorded them, such as burial units. Once these datasets were later connected, a completely different picture began to emerge. It turned out that the new helmets significantly reduced mortality caused by head injuries. According to military medical statistics from World War II, the share of fatal head injuries dropped by tens of percent after the introduction of modern steel helmets, with some units reporting a decline of approximately 40 to 50 percent compared with earlier equipment.
At the same time, the number of soldiers with non fatal head injuries, concussions, or shrapnel wounds increased, but they survived. These soldiers would have been far more likely to die on the battlefield in earlier stages of the war, and they would never have appeared in hospital statistics at all. The rise in the number of injured soldiers was therefore not evidence that the protective gear had failed. It was evidence that it worked. Helmets were turning immediately fatal hits into survivable injuries. Statistics based only on hospital data, however, completely distorted this positive effect.
In this case, survivorship bias did not only lead to a wrong statistical conclusion. It almost led to a wrong strategic decision. Only by including complete data on both the injured and the dead was it possible to interpret the situation correctly and confirm that the new helmets were saving lives.
Aircraft in World War II: where to add armor
The most classic and now textbook example of survivorship bias is the analysis of damage to military aircraft during World War II. This problem was addressed by the British and American air forces as they tried to reduce the heavy losses of bombers during raids over occupied Europe. Every lost aircraft meant not only a destroyed machine, but also the loss of a multi person crew, which created a serious strategic problem.
The military therefore began to systematically collect data on aircraft that returned from combat missions. For these planes, bullet and shrapnel hits were recorded in detail. The statistics showed that the most frequent damage appeared on the wings, the rear sections of the fuselage, and around the tail surfaces. At first glance, it seemed logical to add armor to these parts of the aircraft, because that was where the largest number of hits occurred.
This intuitive conclusion was challenged by the mathematician Abraham Wald. Wald pointed out a fundamental error in reasoning. Only the aircraft that returned from missions were being analyzed. Data on the aircraft that were shot down and never returned to base were completely missing. These invisible aircraft were the key to understanding the problem. Wald noticed that some parts of the aircraft, especially the engines, the cockpit, and the fuel systems, showed surprisingly few hits on the planes that came back. This did not mean those areas were hit less often in combat. It meant that hits in those areas were highly likely to be fatal. If the engine or cockpit was struck, the aircraft simply did not return, and therefore never entered the dataset.
Based on this reasoning, Wald recommended adding armor precisely to the places where hits on surviving aircraft were almost absent. Specifically, this included the engines, the pilot compartment, and certain fuselage sections near fuel tanks. At first, this approach seemed counterintuitive to military leadership because it went directly against the visible data. After implementation, however, there was a measurable reduction in bomber losses. According to later analyses, optimized armor placement led to a reduction in aircraft losses by single digit to lower double digit percentages depending on mission type and deployment.
This case became an iconic example showing that missing data can carry critical information. Wald’s analysis demonstrated that focusing only on surviving objects leads to systematically incorrect conclusions. Here, survivorship bias was not just a theoretical statistical issue. It was a matter of life and death. This example still serves as a warning that when analyzing data, we must ask not only what we see, but also whether something is missing from the data, and why. We should not look only at results, but also view the case from other angles that may reveal what the numbers alone do not.
Why athletes are winter born and academics are autumn born
Survivorship bias also appears very strongly in birth date data among successful people. Among elite athletes, this phenomenon is one of the best documented. Research consistently shows that athletes born in the first part of the year are significantly overrepresented compared with those born later in the year. For example, in the Canadian NHL, analyses by Rodger Barnsley found that approximately 40 percent of players were born in January through March, while only 8 to 10 percent were born in October through December. In the general population, birth distribution is close to even. This difference therefore cannot be random.
A similar pattern appears in European football. UEFA studies show that in youth academies of elite clubs, up to 60 percent of players are born in the first half of the year. Among famous athletes born very early in the year are, for example, Cristiano Ronaldo (5 February), Neymar (5 February), Buffon (28 January), all football, Wayne Gretzky (26 January), Jaromír Jágr (15 February), or Phil Esposito (20 February). These visible examples stand out, while thousands of less fortunate talents born later in the year disappear from the data entirely.
The mechanism is simple. Children born in January are almost a full year older than children born in December within the same youth category. In youth sport, this difference creates a major physical and psychological advantage. Coaches therefore more often select the older children because they show better immediate performance. These children receive more training, better conditions, and greater trust. Professional sport statistics then show only those who survived this selection. The others, often equally talented, are missing from the data. Here again, survivorship bias is at work.
Interestingly, the opposite pattern appears in academic success, especially among Nobel Prize winners. Analyses suggest that Nobel laureates have an above average share of births in autumn, especially in September and October. For example, studies published in the Journal of Biosocial Science found that roughly 30 to 35 percent of Nobel Prize winners born in the Northern Hemisphere fall in the period September through November, while spring months are represented much less. Well known Nobel Prize winners born in autumn include, for example, Niels Bohr (7 October), Richard Thaler (12 September), or Marie Curie (7 November). Autumn births repeatedly appear among scientists who followed long academic careers.
The explanation again does not lie in biology, but in the system. The school year begins in autumn. Children born shortly after the start of the school year are the oldest in their class. They often have greater cognitive maturity, better early educational results, and more frequent positive feedback from teachers. This increases their confidence and willingness to continue studying. These small advantages accumulate over time and increase the probability that a child will pass through the entire education system to its top levels. Nobel Prize statistics then again show only the survivors, not everyone who could have succeeded but dropped out earlier.
Survivorship bias in this case creates the impression that success is linked to the birth date itself. In reality, it reflects the structure of the systems people move through. The data are correct, but without understanding who is missing from them, they lead to incorrect conclusions. Exceptions can of course be found in both sport and science, but those are precisely the cases that had lower odds of reaching the top.