What this calculator does
Accuracy is the share of predictions a classifier, test or model got right overall: correct predictions divided by total predictions. This accuracy calculator works from the four counts of a binary confusion matrix, true positives, false positives, false negatives and true negatives, and reports accuracy alongside precision, recall and F1 score, since accuracy alone often hides a misleading picture.
Accuracy can look excellent while still being a poor measure of performance, particularly when one outcome is much rarer than the other. A model that always predicts "no" for a condition that only affects 1% of cases scores 99% accuracy while catching zero real cases, which is exactly the kind of situation precision and recall are needed to reveal.
The formula
Accuracy is (true positives + true negatives) divided by the total of all four counts. Precision is true positives divided by everything predicted positive (true positives + false positives). Recall is true positives divided by everything that was actually positive (true positives + false negatives). F1 score combines precision and recall into a single balanced figure.
| Term | Meaning |
|---|---|
| Accuracy | The share of all predictions that were correct: (TP + TN) / total. |
| Precision | Of everything predicted positive, the share that was actually positive: TP / (TP + FP). |
| Recall (sensitivity) | Of everything actually positive, the share correctly identified: TP / (TP + FN). |
The inputs explained
| Field | What to enter |
|---|---|
| True positives | True positives: cases correctly predicted positive. |
| False positives | False positives: cases wrongly predicted positive when they were actually negative. |
| False negatives | False negatives: cases wrongly predicted negative when they were actually positive. |
| True negatives | True negatives: cases correctly predicted negative. |
When to use it
Evaluating a machine learning classifier
Reporting accuracy alone on a test set can hide serious weaknesses; checking precision and recall alongside it shows whether the model is actually catching the cases that matter, not just getting most predictions right by favouring the more common class.
Assessing a screening test or fraud detector
In screening for a rare condition or rare event, both accuracy and the trade-off between false positives (unnecessary alarms) and false negatives (missed cases) matter, and this calculator surfaces both from the same four counts.
Comparing two models or test versions
Two versions of the same classifier tested on the same data can be compared directly on accuracy, precision, recall and F1 score together, rather than relying on a single figure that might favour one version for the wrong reasons.
Worked examples
Every figure in the tables below is produced by this page’s own calculator at build time, so the numbers and the tool always agree. Select any row to load that scenario.
How does accuracy and precision change as false positives increase?
Fixed true positives, false negatives and true negatives, across a range of false positive counts.
| False positives | Accuracy | Precision |
|---|---|---|
| 0 | 96.4% | 100.0% |
| 5 | 93.1% | 90.9% |
| 10 | 90.0% | 83.3% |
| 20 | 84.4% | 71.4% |
| 40 | 75.0% | 55.6% |
| 80 | 61.4% | 38.5% |
How does accuracy change with the number of true negatives?
Fixed true positives, false positives and false negatives, across a growing pool of true negatives.
| True negatives | Accuracy | Specificity |
|---|---|---|
| 50 | 87.0% | 83.3% |
| 85 | 90.0% | 89.5% |
| 150 | 93.0% | 93.8% |
| 300 | 95.9% | 96.8% |
| 500 | 97.3% | 98.0% |
| 1000 | 98.6% | 99.0% |
Questions
Why can accuracy be misleading?
When one class is much more common than the other, a classifier can score high accuracy just by favouring the common class, while barely ever correctly identifying the rare one. Precision and recall, focused specifically on the positive class, reveal that weakness where accuracy alone would not.
What is a good accuracy score?
It depends entirely on the problem and the balance between the classes being predicted. There is no universal threshold; a score should be judged against a baseline (such as always predicting the more common class) and alongside precision and recall, not on its own.
How is this different from the diagnostic test accuracy metrics calculator?
This calculator uses general true positive, false positive, false negative and true negative labels, suited to classifiers and predictive models generally. The diagnostic test accuracy metrics calculator uses the same underlying maths but is framed specifically for medical and diagnostic screening tests, including sensitivity, specificity and predictive values by name.
What does an F1 score add beyond accuracy?
F1 score balances precision and recall into one number, which is useful when both false positives and false negatives matter and neither should be ignored. Unlike accuracy, F1 is not affected by how many true negatives are in the dataset, making it a steadier measure when the classes are imbalanced.
For the same maths framed specifically around medical or diagnostic screening tests, see the diagnostic test accuracy metrics calculator.