How statistics lie
Even accurate statistics can lead to wrong conclusions if they are presented without proper care and context. The way numbers are visualized, summarized, and framed strongly influences how they are understood by the reader. Charts, averages, or percentages may appear objective, yet they always highlight some aspects of reality while downplaying others. Without clear explanation and transparent presentation, correct results can easily be misunderstood or become misleading. That is why responsible statistical communication is not only about getting the numbers right, but about presenting them in a way that supports fair and accurate interpretation.
US Presidential Election 2016: when the map lies but the data do not
The 2016 US presidential election is one of the most well known examples of how data visualization can fundamentally distort reality. Donald Trump won the election even though he received fewer votes than his opponent. This fact alone is surprising to many people because we intuitively expect the winner of an election to be the candidate who receives the highest number of votes. In the American electoral system, however, the outcome is decided by electors, not by the direct popular vote.
Shortly after the election, a map of the United States spread widely across media and social networks. The map was almost entirely colored red. At first glance, it appeared to be clear evidence of a sweeping Republican victory across the entire country. For many people, it became a visual confirmation that Trump had the support of the majority of America. This impression, however, was largely misleading.
The problem was not the election data themselves, but the way they were displayed. A traditional state map primarily shows geographic area, not the number of inhabitants or voters. Large and sparsely populated states therefore occupy vast areas on the map, even though relatively few people live there. These states also have a low number of electors, which limits their real influence on the election result. In contrast, densely populated states such as California or New York appear relatively small on the map, despite being home to tens of millions of people. Their actual weight is therefore visually strongly underestimated. The map highlights space rather than population and creates an image that does not correspond to the distribution of votes. The result is a visualization that supports a very strong but misleading narrative.
Once a cartogram or so called tile map is used instead, where one unit corresponds to a specific number of inhabitants or electors, the picture of the election changes dramatically. It suddenly becomes clear where support for individual candidates was truly concentrated. Small but populous states are no longer visually marginalized and their importance becomes much more apparent. At the same time, it becomes evident that vast areas with few inhabitants do not carry the decisive weight suggested by the traditional map. The same data begin to tell a completely different story.
In neither case are the numbers incorrect. The difference lies solely in what the visualization emphasizes and what it hides. This example clearly shows that a map is not a neutral tool, but a powerful interpretive framework. It can very easily shape public perception of reality and political sentiment. That is why it is always important to ask what exactly a given map shows and especially what it does not show. The data did not lie. The visual interpretation was misleading.
Manipulating scale: how to turn a small difference into a crisis
Another very common way statistics can create a misleading impression is through manipulation of axis scales in graphs. This issue appears mainly in media, political communication, and various presentations where a chart is meant to support a particular narrative. At first glance, such a graph may appear completely correct and professional.
A typical example is a comparison of two values, such as inflation at 4.0 percent and 5.5 percent. On paper, this represents a difference of 1.5 percentage points, which is a change with some significance but not an extreme one. However, if the graph is designed so that the axis does not start at zero but instead at 3.8 percent, the difference begins to look dramatic. Bars or curves visually diverge much more, and the change appears far larger than it actually is. The viewer may then gain the impression that inflation has surged or that the economic situation has deteriorated dramatically.
In reality, the absolute change has not changed at all. Only the method of visualization has. The same data plotted on an axis starting at zero would appear much calmer and more moderate. The difference would still be visible, but it would not feel alarming. This contrast shows that the numbers themselves are not the problem. The problem lies in the context into which they are placed and the visual framing chosen by the author of the chart. The axis scale determines whether a change is perceived as a minor deviation or as a crisis.
The same principle can also be used in the opposite direction. If we want to obscure or downplay a difference, we choose a very wide scale. In that case, graphs look almost identical even when the values differ significantly. Such a graph can create the impression that nothing has really changed, even in situations where the change has real consequences. This approach often appears in presentations intended to reassure the public, investors, or voters. Negative developments are visually diluted within a wide scale and lose urgency. Readers or viewers may not notice the difference at all or may underestimate it.
Axis scaling thus becomes a very powerful tool for interpreting data. In such cases, the graph does not convey information but works with emotions. Sometimes it triggers panic, other times false calm. Yet the underlying data remain the same in both cases. That is why it is essential to always check where an axis starts and where it ends. Without this context, a graph may appear convincing while being misleading at the same time. Proper visualization should reflect the real significance of a change, not serve political, media, or marketing intentions. Once scale becomes a tool of manipulation, the graph stops informing and starts influencing, even when the values themselves are correct.
Is the average or the median better?
Misleading interpretation of statistics does not concern only elections or sports. It appears very often in everyday economic and social data as well. A typical example is the use of averages, which may seem intuitive and objective at first glance.
Average wages are frequently used when comparing cities, regions, or entire countries. When we look more closely, however, we find that averages can significantly distort reality. Imagine a city where most people earn around 3,500 dollars per month, but where several large companies also operate with extremely high executive salaries. These extreme incomes pull the average sharply upward. The result may be an average wage of 5,000 dollars per month, even though most residents are nowhere near that income. The average then does not describe the experience of the typical person, but rather reflects the existence of a small group of very highly paid individuals.
In such situations, the median wage is a much more appropriate indicator. The median shows how much the person exactly in the middle of the income distribution earns. Half of the people earn less and half earn more. If the median in a given city is, for example, 3,800 dollars, it corresponds much better to the reality of most residents. The difference between the average and the median also reveals income inequality. The larger the gap between these two values, the more unevenly incomes are distributed.
A similar issue appears when comparing regions, where one large city can significantly improve the average for an entire region. People in smaller towns or rural areas then feel that the presented figures do not match their lived reality at all. This is largely because people generally understand averages more easily than medians. When we tell them that the average wage in a region is 3,500 dollars, they grasp it faster than if we say that the median income in the region is 3,800 dollars.
Distortion also often arises from poor choice of comparison periods. If we compare an extreme year with a normal period, results may appear either excessively positive or catastrophically negative. Crisis years often serve as a very low baseline. Any return toward normal then looks like rapid growth, even though it is only partial recovery. This was clearly visible during the COVID 19 pandemic, when the economy contracted sharply in 2020 due to lockdowns. If GDP then grew year on year by, for example, 5 percent in 2021, it could appear as an extraordinary success, even though the economy still had not reached pre pandemic levels.
In such cases, the data are technically correct but taken out of context. A simple year on year comparison does not indicate whether we are seeing real growth or merely recovery from an unusually weak period. For correct interpretation of statistics, it is therefore crucial to ask not only how much, but also who it affects, which period the data come from, and how the values are distributed. Only the combination of these questions allows us to understand what the numbers truly say and what they only seem to suggest.
Sports statistics: when percentages mislead goaltenders
Sports statistics can appear very convincing, but without context they often lead to incorrect conclusions. A typical example is evaluating ice hockey goaltenders solely by save percentage.
Imagine two goaltenders. The first has a 100 percent save rate because he stopped all 10 shots he faced. The second has a save percentage of 88 percent because he faced 50 shots and allowed 6 goals. At first glance, it appears that the first goaltender delivered a much better performance. This conclusion, however, is misleading.
The goaltender with a perfect save rate may have played behind a very strong team and faced only a few non dangerous situations. In contrast, the second goaltender may have been under extreme pressure throughout the game. Fifty shots represent a very high workload, and even very strong performances usually result in goals against. In this context, an 88 percent save rate may be above average, even though the percentage itself looks negative. Both values are calculated correctly, but without understanding the circumstances, they tell a distorted story.
The problem therefore lies not in mathematics, but in interpretation. Save percentage does not account for shot quality, team defense, or the overall workload of the goaltender. That is why it is necessary to consider additional metrics such as shot volume, expected goals, or game context. This example shows that a single number is never sufficient on its own and that even correct statistics can be misleading without context.
Why correct data are not enough
All the previous examples show one fundamental truth. Statistics themselves do not lie. Misleading conclusions arise when we interpret them without understanding context. Whether we are dealing with election maps, axis scales in charts, sports percentages, or average wages, the problem is not the number but how it is interpreted.
Visualization can strongly influence how we perceive data. It can emphasize one aspect of reality while hiding another. An average can conceal inequality, a map can highlight space instead of people, and axis scaling can turn a small difference into a crisis. That is why it is essential to approach data critically and ask what the numbers truly tell us and what they do not. Proper work with data is not only about calculation, but primarily about fair interpretation. Only then can statistics genuinely help us understand reality instead of distorting it.