What this calculator does
The sum of squares, written SS, adds up how far every value in a data set sits from the mean, with each of those distances squared before it is added. Squaring does two jobs at once: it stops positive and negative deviations cancelling each other out, and it makes a few large deviations count for far more than many small ones.
The name is ambiguous, and that is the usual source of confusion. A sum of squares statistics calculator returns the total squared deviation, Σ(yᵢ − ȳ)², the quantity variance and standard deviation are built from. It is not the same as adding up the squares of the raw values, Σyᵢ², which is a different number entirely. This page computes the first of the two.
The formula
The mean of the data set is found first. Every value then has the mean subtracted from it, that difference is squared, and the squares are added together to give SS. Dividing SS by n − 1 gives the sample variance and dividing it by n gives the population variance, which is why SS is usually described as the building block the rest of the spread measures rest on.
| Term | Meaning |
|---|---|
| SS | Sum of squares: the total of every squared deviation from the mean. |
| ȳ | The mean of the data set, the point all the deviations are measured from. |
| Deviation | The signed distance of one value from the mean, before it is squared. |
| n − 1 | The divisor used for sample variance, one less than the count of values, correcting for the mean having been estimated from the same data. |
The inputs explained
| Field | What to enter |
|---|---|
| Data set (comma separated) | The data set, as numbers separated by commas. At least two values are needed, since a single value has no spread to measure. |
When to use it
Working through a variance or standard deviation by hand
SS is the step in the middle of both calculations. Getting it separately makes the working easier to check, since an error in the mean or in one deviation shows up here rather than being buried inside a final figure.
Comparing the spread of two data sets of the same size
SS is directly comparable between sets that have the same number of values, because it has not yet been divided by anything. Across sets of different sizes, the variance is the fairer comparison, since a larger set accumulates more squared deviations simply by having more of them.
Checking a figure quoted in regression or ANOVA output
Both methods split variation into several sums of squares. The figure this calculator produces for the response values is the total sum of squares, the one the explained and residual parts add back up to.
Worked examples
Every figure in the tables below is produced by this page’s own calculator at build time, so the numbers and the tool always agree. Select any row to load that scenario.
How does the sum of squares change as a data set spreads out?
Five data sets, all with five values and all with a mean of exactly 10, so the only thing changing between rows is how widely the values are spread.
| Data set | Sum of squares (SS) | Mean | Sample variance |
|---|---|---|---|
| 10, 10, 10, 10, 10 | 0 | 10.000 | 0 |
| 9, 10, 10, 10, 11 | 2.000 | 10.000 | 0.5000 |
| 8, 9, 10, 11, 12 | 10.000 | 10.000 | 2.500 |
| 6, 8, 10, 12, 14 | 40.000 | 10.000 | 10.000 |
| 2, 6, 10, 14, 18 | 160.000 | 10.000 | 40.000 |
Does shifting every value change the sum of squares?
The same five consecutive numbers, shifted up by 10, then 100, then 1,000.
| Data set | Sum of squares (SS) | Mean |
|---|---|---|
| 1, 2, 3, 4, 5 | 10.000 | 3.000 |
| 11, 12, 13, 14, 15 | 10.000 | 13.000 |
| 101, 102, 103, 104, 105 | 10.000 | 103.000 |
| 1001, 1002, 1003, 1004, 1005 | 10.000 | 1,003.00 |
Questions
What is the difference between the sum of squares and the sum of squared values?
The sum of squares in statistics subtracts the mean first: it is Σ(yᵢ − ȳ)². The sum of squared values simply squares each number as it stands and adds those, Σyᵢ². For the set 1, 2, 3, 4, 5 the first is 10 and the second is 55. Only the first says anything about spread.
Why square the deviations instead of just adding them up?
Because the plain deviations always add to zero, by definition of the mean, so they would measure nothing at all. Squaring removes the signs, and it also weights the result toward the values furthest from the mean, which is usually what a spread measure is wanted for.
Should I divide the sum of squares by n or by n − 1?
Divide by n − 1 when the data is a sample and you want an estimate of the variance of the wider population it came from. Divide by n when the data is the whole population. This calculator shows both, so the right one can be read off directly.
Can the sum of squares be negative?
No. Every term in it is a square, so every term is zero or positive and the total cannot fall below zero. SS is exactly zero only when every value in the set is identical to the mean, which also means identical to each other.
For the mean, median, quartiles and standard deviation of the same data set in one pass, see the descriptive statistics calculator. For a spread measure built on absolute distances rather than squared ones, see the mean absolute deviation calculator.