StatGardenREF. DESK
Calculators/Statistics/Confusion matrix
Statistics

Confusion matrix calculator

Turns the four counts of a binary confusion matrix into accuracy, precision, recall and F1.

Published 9 August 2026 · Updated 14 August 2026

What this calculator does

Accuracy is the share of predictions a classifier, test or model got right overall: correct predictions divided by total predictions. This accuracy calculator works from the four counts of a binary confusion matrix, true positives, false positives, false negatives and true negatives, and reports accuracy alongside precision, recall and F1 score, since accuracy alone often hides a misleading picture.

Accuracy can look excellent while still being a poor measure of performance, particularly when one outcome is much rarer than the other. A model that always predicts "no" for a condition that only affects 1% of cases scores 99% accuracy while catching zero real cases, which is exactly the kind of situation precision and recall are needed to reveal.

The formula

FormulaAccuracy=(TP+TN)/total, Precision=TP/(TP+FP), Recall=TP/(TP+FN), F1=2·Precision·Recall/(Precision+Recall), Specificity=TN/(TN+FP)

Accuracy is (true positives + true negatives) divided by the total of all four counts. Precision is true positives divided by everything predicted positive (true positives + false positives). Recall is true positives divided by everything that was actually positive (true positives + false negatives). F1 score combines precision and recall into a single balanced figure.

TermMeaning
AccuracyThe share of all predictions that were correct: (TP + TN) / total.
PrecisionOf everything predicted positive, the share that was actually positive: TP / (TP + FP).
Recall (sensitivity)Of everything actually positive, the share correctly identified: TP / (TP + FN).

The inputs explained

FieldWhat to enter
True positivesTrue positives: cases correctly predicted positive.
False positivesFalse positives: cases wrongly predicted positive when they were actually negative.
False negativesFalse negatives: cases wrongly predicted negative when they were actually positive.
True negativesTrue negatives: cases correctly predicted negative.

When to use it

Evaluating a machine learning classifier

Reporting accuracy alone on a test set can hide serious weaknesses; checking precision and recall alongside it shows whether the model is actually catching the cases that matter, not just getting most predictions right by favouring the more common class.

Assessing a screening test or fraud detector

In screening for a rare condition or rare event, both accuracy and the trade-off between false positives (unnecessary alarms) and false negatives (missed cases) matter, and this calculator surfaces both from the same four counts.

Comparing two models or test versions

Two versions of the same classifier tested on the same data can be compared directly on accuracy, precision, recall and F1 score together, rather than relying on a single figure that might favour one version for the wrong reasons.

Worked examples

Every figure in the tables below is produced by this page’s own calculator at build time, so the numbers and the tool always agree. Select any row to load that scenario.

How does accuracy and precision change as false positives increase?

Fixed true positives, false negatives and true negatives, across a range of false positive counts.

50 true positives, 5 false negatives, 85 true negatives
False positivesAccuracyPrecision
096.4%100.0%
593.1%90.9%
1090.0%83.3%
2084.4%71.4%
4075.0%55.6%
8061.4%38.5%
With zero false positives, accuracy reaches 96.4% and precision is a perfect 100.0%. Raising false positives to 80 drags accuracy down to 61.4% and precision down to 38.5%, since more and more predicted positives turn out to be wrong.

How does accuracy change with the number of true negatives?

Fixed true positives, false positives and false negatives, across a growing pool of true negatives.

50 true positives, 10 false positives, 5 false negatives
True negativesAccuracySpecificity
5087.0%83.3%
8590.0%89.5%
15093.0%93.8%
30095.9%96.8%
50097.3%98.0%
100098.6%99.0%
Accuracy climbs from 90.0% at 85 true negatives to 98.6% at 1,000, even though the false positive and false negative counts never change, which is exactly why accuracy alone can look better simply because correct negatives dominate a large, imbalanced dataset.

Questions

Why can accuracy be misleading?

When one class is much more common than the other, a classifier can score high accuracy just by favouring the common class, while barely ever correctly identifying the rare one. Precision and recall, focused specifically on the positive class, reveal that weakness where accuracy alone would not.

What is a good accuracy score?

It depends entirely on the problem and the balance between the classes being predicted. There is no universal threshold; a score should be judged against a baseline (such as always predicting the more common class) and alongside precision and recall, not on its own.

How is this different from the diagnostic test accuracy metrics calculator?

This calculator uses general true positive, false positive, false negative and true negative labels, suited to classifiers and predictive models generally. The diagnostic test accuracy metrics calculator uses the same underlying maths but is framed specifically for medical and diagnostic screening tests, including sensitivity, specificity and predictive values by name.

What does an F1 score add beyond accuracy?

F1 score balances precision and recall into one number, which is useful when both false positives and false negatives matter and neither should be ignored. Unlike accuracy, F1 is not affected by how many true negatives are in the dataset, making it a steadier measure when the classes are imbalanced.

For the same maths framed specifically around medical or diagnostic screening tests, see the diagnostic test accuracy metrics calculator.