What this calculator does
Allele frequencies come from counting alleles rather than individuals. Each individual carries two, so a sample of N individuals holds 2N alleles. A homozygote contributes two of the same, a heterozygote one of each.
That gives p = (2 × AA + Aa) ÷ 2N and q = (2 × aa + Aa) ÷ 2N. This is a direct count from observed data, with no model assumed, which makes it the figure to calculate first and then compare against Hardy-Weinberg expectations rather than the other way round.
The formula
The total allele count is twice the number of individuals. Dominant alleles are counted as two from each homozygous dominant individual plus one from each heterozygote, and the same logic applies to the recessive count. Each is divided by the total to give a frequency, and the two frequencies sum to one.
| Term | Meaning |
|---|---|
| p | Dominant allele frequency, counted directly from the genotypes observed. |
| q | Recessive allele frequency. p + q = 1 for a locus with two alleles. |
| 2N | Total alleles in the sample, twice the number of individuals since each is diploid. |
| Genotype count | The number of individuals of each of the three genotypes, which is the raw observation. |
The inputs explained
| Field | What to enter |
|---|---|
| Homozygous dominant count (AA) | Number of homozygous dominant individuals (AA). |
| Heterozygous count (Aa) | Number of heterozygous individuals (Aa). |
| Homozygous recessive count (aa) | Number of homozygous recessive individuals (aa). |
When to use it
Analysing genotyping results
Genotype counts are what a sequencing or typing run produces, and allele frequencies are what population genetics works in, so this conversion is the first step in almost any analysis.
Testing for Hardy-Weinberg equilibrium
Calculate the frequencies from the observed counts here, then feed q into the Hardy-Weinberg calculator to get expected genotype proportions and compare the two.
Tracking frequencies over time
Repeating the calculation on samples from successive generations shows whether an allele is rising or falling, which is the direct evidence of selection or drift.
Worked examples
Every figure in the tables below is produced by this page’s own calculator at build time, so the numbers and the tool always agree. Select any row to load that scenario.
What allele frequencies do these genotype counts give?
Three genotype distributions, with the heterozygote and recessive counts held fixed while the homozygous dominant count changes.
| AA count (Aa held at 40, aa at 10) | Dominant allele frequency (p) | Recessive allele frequency (q) | Individuals sampled | Alleles sampled (2N) |
|---|---|---|---|---|
| 50 | 0.7000 | 0.3000 | 100 | 200 |
| 25 | 0.6000 | 0.4000 | 75 | 150 |
| 90 | 0.7857 | 0.2143 | 140 | 280 |
Questions
Why multiply the homozygote count by two?
Because each homozygous individual carries two copies of the same allele. A heterozygote carries one of each, so it contributes one to each count. Counting individuals rather than alleles is the most common error in these calculations.
Do p and q always add to one?
For a locus with exactly two alleles, yes, by construction. Loci with three or more alleles need a frequency for each and they sum to one collectively, which this two-allele calculator does not handle.
Is this the same as Hardy-Weinberg?
No, and the distinction matters. This counts what is actually there, assuming nothing. Hardy-Weinberg predicts what genotype frequencies should be given those allele frequencies if the population is in equilibrium. Calculating the first and comparing against the second is how equilibrium is tested.
How large a sample do I need?
Larger than is usually convenient, particularly for rare alleles. Estimating a frequency of 0.01 reliably needs hundreds of individuals simply to observe the allele more than once or twice. Small samples give frequencies with very wide confidence intervals.
For expected genotype frequencies from an allele frequency, see the Hardy-Weinberg calculator. For multi-gene cross probabilities, see the multi-gene cross calculator.