관리
← All articles

Mean Versus Median: Which Number to Read in Skewed Data

This article was translated from its source language with AI assistance. Please check technical terms and equations against the original.

Mean and median calculator

Inputs are processed on this screen only and are never sent to the server or saved.

Calculate the mean and median yourself

Enter numbers separated by commas or spaces. Up to 500 values are allowed, each with an absolute value no greater than one trillion. Compare values with the same units. Inputs are calculated only on this screen.

Mean Versus Median: Which Number to Read in Skewed Data — Original concept illustration
Original concept illustration
Numbers to calculate

The illustrative example has a mean of 100 and a median of 30. This calculation cannot establish the sample's representativeness or the data's accuracy.

JavaScript is disabled. Formula: mean=sum÷count; median=the middle value after sorting by size (for an even count, the mean of the two middle values).

The mean and median answer different questions

An announcement that “the average waiting time is 10 minutes” can make it seem that most people waited about 10 minutes. Yet just a few long waits can raise the mean. The median is read from the middle of data sorted by size. Rather than automatically calling one more accurate, first distinguish whether you want to understand the total burden or a typical experience.

This article uses hypothetical waiting-time data. These numbers are illustrative examples, not any institution's actual statistics or measurements. We follow calculations using five simple values and examine errors from combining means alone and the effects of handling missing data. Official explanations from NIST and the Australian Bureau of Statistics inform the article, but the situations and calculations below were newly prepared to help readers make judgments.

Before examining a mean or median, the data must have matching units and populations. Numbers differ depending on whether they are customer waiting times or daily averages by service counter, seconds or minutes, and whether people who did not wait are included. Even a correctly calculated representative value can mislead if the question and data scope differ.

1. Calculate directly with the same five people's data

Assume illustrative waiting times of 2, 3, 4, 5, and 36 minutes. Their sum is 50 minutes, and dividing by five people gives a mean of 10 minutes. Since they are already sorted, the third value, 4 minutes, is the median. The mean of 10 and median of 4 come from the same data; neither is a calculation error.

For comparison, replace only the last value with 6 minutes, giving 2, 3, 4, 5, and 6 minutes. The sum is 20 minutes, the mean is 4 minutes, and the median is also 4 minutes. Both datasets have the same middle value but different total waiting times. This demonstrates that the mean and median extract different information. You can directly observe how greatly changing one value affects the mean.

Hypothetical data / minutesMeanMedian
2, 3, 4, 5, 620 ÷ 5 = 4 minutes4 minutes
2, 3, 4, 5, 3650 ÷ 5 = 10 minutes4 minutes
2, 3, 5, 616 ÷ 4 = 4 minutes(3 + 5) ÷ 2 = 4 minutes

2. The median need not be an observed value

With an odd number of values, use the single middle value; with an even number, use the mean of the two middle values. For the four values 2, 3, 5, and 6 minutes above, the middle values are 3 and 5, so the median is 4 minutes. The median can be 4 even when nobody actually waited 4 minutes. The interpretation that “there must be someone exactly at the median” is incorrect.

Forgetting to sort produces the wrong middle position. If the input order is 36, 2, 5, 3, and 4 minutes, do not choose the third input, 5 minutes, as the median. Find the middle after sorting by size. If changing the original table's order is difficult, retain original row IDs and sort a calculation copy. Recording whether the original order was preserved in the report also helps maintain links to other columns.

Distinguish the ability to sort values from the meaningfulness of arithmetic. Categories without a size ordering, such as colors, do not readily have means or medians. For satisfaction rankings, the meaning of order may differ from the meaning of intervals between grades. Numeric coding alone does not justify adding and dividing values as though they were times or lengths.

3. Questions requiring a mean and questions requiring a median

The relationship between count and mean is useful for understanding total waiting time. In the hypothetical five-person example, multiplying the mean of 10 minutes by 5 gives a total of 50 minutes. Such totals can help compare accumulated delays at service counters. However, the sum of overlapping waits may not equal the actual elapsed time during which a counter stopped, so state what the measure means.

When describing a representative individual's experience, presenting the median as well is helpful. The median of 4 minutes in the example shows the waiting time at the middle position. But it alone does not show how many people waited a long time. Replacing the mean with the median is insufficient to explain the waiting problem fully.

A practical report can include count and range: “For five people, the mean was 10 minutes, the median 4, the minimum 2, and the maximum 36.” The fact that there are only five people also matters. One unusual circumstance can greatly change the result, so do not claim these figures predict every visitor's wait the next day.

4. Do not immediately delete a large value

Automatically deleting 36 minutes because it looks larger than the others can conceal a genuinely long wait. First determine whether it is an input error or a real event. Investigate whether a seconds value entered a minutes column, whether the waiting start time was entered incorrectly, or whether a procedure took unusually long. If an error is confirmed, record the correction reason; if it is a real observation, decide inclusion according to the analytical purpose.

Mean Versus Median: Which Number to Read in Skewed Data — Original illustration of the key points
Original illustration of the key points

If the data include a legitimate long wait, present the mean and median together and explain their difference. If a value is confirmed as a typo, preserve the link between the original and corrected values. Do not exclude large values merely because they prevent the desired result. Where both before- and after-exclusion figures are needed, present both alongside the exclusion criteria.

NIST's explanation of measures of location discusses how means and medians can respond differently to skewness or long tails. Do not turn this into a rule that “the mean must exceed the median in every small right-skewed dataset.” Examine the actual values and distribution together. The direction of the difference alone should not determine the entire shape of the data.

5. Check group sizes when combining means

Assume hypothetical counter A has two users with a mean wait of 10 minutes, while counter B has eight users with a mean of 20 minutes. Adding the two counter means and dividing by two gives 15 minutes. But the mean across all ten users must account for different group sizes. Adding A's total of 20 minutes to B's 160 minutes and dividing by ten gives 18 minutes.

Here, 15 minutes averages the two counters' means with equal weight, while 18 minutes is the overall mean giving equal weight to each user. They answer different questions. Without stating the purpose, either can be mistaken for a calculation error. A report about overall user experience should verify calculations weighted by each group's user count.

Nor can you average two group medians to find the overall median. The median requires knowing the middle position in the complete data. Group summaries alone may provide insufficient information. In that case, explain why original data or distribution information is needed rather than arbitrarily estimating an unknown figure.

6. Blanks and zero change the denominator

Consider illustrative values of 2, 3, and a blank. Treating the blank as unobserved gives a mean of 2.5 for the two known values. Replacing it with an actual zero gives approximately 1.67 across three values. The table looks similar, but the calculations have different meanings. Distinguish zero for someone who did not wait from a blank for someone whose wait was not measured.

State the denominator and missing count, such as “two valid observations and one missing observation.” An original population of three does not always make the mean's denominator 3. Conversely, excluding a valid zero wait as a blank changes the representative value. Check each code's meaning before calculating, and record how missing values were handled during preparation.

Mixed units cause similar problems. If one row contains 120 seconds and another 3 minutes, but only the numbers 120 and 3 are stored, do not average them directly. Convert to consistent units first and preserve the original units. Decide decimal places at the final display stage, checking whether intermediate rounding changes the result.

7. Read graphs and numbers together

The five illustrative values can be displayed as a dot plot with minutes on the horizontal axis. Four dots lie near 2, 3, 4, and 5, with one distant dot at 36. Showing the mean of 10 and median of 4 together makes their difference easier to understand. Actual statistical graphs should also state axis units, sample size, and data source.

When comparing large datasets, distributions and group composition matter alongside representative values. If mornings contain many short tasks and afternoons many long ones, differences between time-of-day means may not indicate counter-performance differences alone. First check who was included and whether measurement criteria matched. Aggregating more numbers requires revisiting the original data conditions.

The maximum and minimum of a small table are easy to calculate but limited in describing common experience. A single dot stretching the graph can obscure differences between short waits. If trimming the displayed axis, clearly state its range and display method, and do not hide the large value. Visualization should reveal the distribution, not make a representative value look favorable.

8. Final questions for checking statistical statements

First ask whose data these are, what they measure, their units, how many valid values exist, and how missing observations and exclusions were handled. Next identify which questions the mean and median answer, and check weights when groups are combined. Finally look for distribution or range information that reveals differences behind the representative value.

When writing a report yourself, include conditions: “In five illustrative observations, the mean is 10 minutes and the median 4, including one long wait.” This states the calculation more clearly than “people normally wait 10 minutes.” Providing both the table and explanation lets readers use both summaries to make the judgments they need.

The difference between mean and median is more than statistical vocabulary: it helps check the viewpoint from which data were summarized. Understanding how totals and middle experiences differ in the same data supports more accurate reading of figures in news, work reports, and service information. For actual decisions, examine the original scope and conditions before applying illustrative calculations.

Official sources and writing standards

References checked: 2026-10-03. This explanation was written with AI assistance based on official sources actually opened. Separately marked calculations, code, and checking examples are illustrative, not results from directly testing or measuring the user environment. Recheck changes to features and references on the publication date.

Original illustrations created to help explain this article.

Original on Tistory ↗