StatGardenREF. DESK
Calculators/Statistics/Diagnostic test accuracy metrics
Statistics

Diagnostic test accuracy metrics calculator

Sensitivity, specificity, predictive values and accuracy from a 2×2 test result table.

Published 8 August 2026 · Updated 25 September 2026

What this calculator does

From four counts, true positives, false positives, false negatives and true negatives, a set of standard measures follows. A test with 85 true positives, 15 false positives, 10 false negatives and 90 true negatives is 89.5% sensitive and 85.7% specific.

Sensitivity and specificity are properties of the test. The predictive values are not: they depend on how common the condition is in the group being tested. Move the same test to a population where the condition is rarer and the positive predictive value falls, even though nothing about the test changed. That distinction is the source of most confusion in this area.

The formula

FormulaSensitivity=TP/(TP+FN); Specificity=TN/(TN+FP); PPV=TP/(TP+FP); NPV=TN/(TN+FN); Youden J = Sens+Spec−1

Sensitivity is true positives over all who have the condition, and specificity is true negatives over all who do not. Positive predictive value is true positives over all positive results, and negative predictive value is the mirror. Accuracy pools everything, F1 is the harmonic mean of sensitivity and precision, and Youden J is sensitivity plus specificity minus one.

TermMeaning
SensitivityThe share of true cases the test catches. Also called recall or the true positive rate.
SpecificityThe share of non-cases the test correctly clears.
PPVOf those testing positive, the share who actually have the condition. Depends on prevalence.
F1 scoreThe harmonic mean of precision and recall, useful when one class is much rarer than the other.

The inputs explained

FieldWhat to enter
True positivesTrue positives: have the condition and tested positive.
False positivesFalse positives: do not have it but tested positive.
False negativesFalse negatives: have it but tested negative.
True negativesTrue negatives: do not have it and tested negative.

When to use it

Evaluating a diagnostic test

The full set of measures describes different failure modes, and which one matters depends on the consequences of each error.

Assessing a classifier

The same 2x2 table underlies machine learning evaluation, where precision, recall and F1 are the usual names.

Choosing a threshold

Moving a cutoff trades sensitivity against specificity, and the metrics show the shape of that trade.

Worked examples

Every figure in the tables below is produced by this page’s own calculator at build time, so the numbers and the tool always agree. Select any row to load that scenario.

How do the metrics respond to different errors?

The same rough sample size with errors distributed differently.

Around 200 tested
False negatives (cases missed)Sensitivity (recall)SpecificityPrecision (PPV)
5 missed94.4%85.7%85.0%
10 missed89.5%85.7%85.0%
25 missed77.3%85.7%85.0%
50 missed63.0%85.7%85.0%
Missing more cases drives sensitivity down from 94.4% to 63.0%, yet specificity holds at 85.7% and precision at 85.0% throughout. Neither of those depends on the false negative count, which is why a test can look precise while missing a third of the cases it was meant to find.

Questions

What is the difference between sensitivity and specificity?

Sensitivity is about catching real cases: of everyone who has the condition, what share does the test find. Specificity is about not raising false alarms: of everyone who does not, what share does the test correctly clear. A test can be excellent at one and poor at the other.

Why does the positive predictive value change between populations?

Because it depends on how common the condition is. In a low-prevalence group, most positives come from the large healthy majority, so PPV falls. Sensitivity and specificity stay fixed, which is why they are the properties usually quoted for a test.

Which metric should I optimise?

It depends on the cost of each error. Screening for a serious treatable disease favours sensitivity, since missing a case is worse than a false alarm. A test triggering an invasive follow-up favours specificity. Accuracy alone is misleading whenever the classes are unbalanced.

Why is accuracy misleading?

Because a rare condition can be predicted with high accuracy by never predicting it at all. A condition affecting 1% of people gives 99% accuracy from a test that always says negative, while having zero sensitivity. Report sensitivity and specificity alongside it.

For the combined summary score, see the Youden’s index calculator. For the underlying table, see the confusion matrix calculator.