What this calculator does
When data arrives as a frequency table rather than a list, the mean and standard deviation can only be estimated, by treating every value in a class as though it sat at the class midpoint. A table with midpoints 5 to 45 and frequencies 4, 9, 15, 8, 4 gives an estimated mean of 24.75.
The estimate is genuinely an estimate. The midpoint assumption is exactly right only if values are evenly spread within each class, which they rarely are. In practice the mean comes out close and the standard deviation is slightly understated, because within-class variation is discarded entirely.
The formula
Each class midpoint is weighted by its frequency to give the mean. The standard deviation uses the same weighting on squared deviations from that mean, with the total frequency minus one in the denominator for the sample version. Both sample and population forms are shown; the sample form is the usual choice unless the table covers an entire population.
| Term | Meaning |
|---|---|
| Class midpoint | The centre of a class interval, used to stand in for every value in it. |
| Frequency | How many observations fall in each class. |
| Grouping error | The inaccuracy introduced by the midpoint assumption. |
| Total frequency | The sum of all frequencies, which is the effective sample size. |
The inputs explained
| Field | What to enter |
|---|---|
| Class midpoints (comma separated) | Class midpoints, comma separated. For a class covering 0 to 10, the midpoint is 5. |
| Frequencies (comma separated) | Frequencies, comma separated, in the same order as the midpoints. The two lists must be the same length. |
When to use it
Working from a published table
Official statistics are often released as frequency tables with the raw data unavailable, and this is the only route to summary figures.
Summarising survey bands
Surveys that collect age or income in bands produce exactly this kind of table.
Checking a histogram
Deriving the mean and standard deviation from the bars gives a numerical summary to go with the picture.
Worked examples
Every figure in the tables below is produced by this page’s own calculator at build time, so the numbers and the tool always agree. Select any row to load that scenario.
How does the distribution shape affect the estimates?
The same classes with three different frequency patterns.
| Frequencies | Estimated mean | Estimated sample standard deviation | Total frequency |
|---|---|---|---|
| 20, 10, 5, 3, 2 | 14.250 | 11.851 | 40 |
| 4, 9, 15, 8, 4 | 24.750 | 11.206 | 40 |
| 2, 3, 5, 10, 20 | 35.750 | 11.851 | 40 |
Questions
How accurate is the grouped estimate?
The mean is usually close, often within a percent or so. The standard deviation tends to be understated because all variation inside each class is thrown away. Narrower classes reduce both errors, which is one reason published tables with many bands are more useful than those with few.
What do I use for an open-ended class?
You have to assume a boundary, such as treating "65 and over" as 65 to 85 and using 75 as the midpoint. This is the weakest point of the whole method, since the assumption is arbitrary and open-ended classes usually sit in the tail where they most affect the result.
How do I find a class midpoint?
Average the lower and upper boundaries of the class. For a class covering 20 to 30, the midpoint is 25. For discrete data recorded as 20 to 29, the true boundaries are 19.5 and 29.5, giving a midpoint of 24.5, which is a distinction worth keeping straight.
Should I use the sample or population standard deviation?
The sample version unless your frequency table genuinely covers the entire population of interest. For a census the population form is correct; for a survey the sample form is. The difference shrinks quickly as the total frequency grows.
For choosing the classes in the first place, see the class width calculator. For summary statistics from raw data, see the descriptive statistics calculator.