Method

Precise, accurate, or neither

Three tight readings can all be wrong the same way, and three scattered readings can average to the truth; the two problems have different fixes.

3 min readMethod

Precision and accuracy get used as if they were the same compliment. They are not, and a set of readings can have one without the other - which matters, because the fix for each is different.

Two definitions

Accuracy is how close your readings sit to the true value. Precision is how close your readings sit to each other.

A reading can be precise without being accurate: three measurements that agree tightly, all taken with the same consistent flaw, land close together and far from the true figure. A reading can be accurate without being precise: three measurements that scatter widely can still average out to something close to the true value, by luck rather than by method.

The target diagram

Accurate + precise Precise, not accurate Accurate, not precise Neither

Precise readings sit close together, wherever the cluster lands. Accurate readings sit close to the centre. The two properties are independent, and a good outcome needs both.

What each failure looks like on a ruler

Precise but inaccurate looks like three readings that agree within a millimetre, all taken while consistently under-pressing at the bone. The technique is repeatable - that is what precision means - but repeatable in a way that reads short every time, because the same soft-tissue compression happens the same way each time. Three tight numbers here feel trustworthy and are not.

Accurate but imprecise looks like three readings that swing by several millimetres because the tape kept sliding, or the pressure varied session to session, but happen to average close to a true value. This one is harder to spot as a problem, because the final reported number - if you take the mean - can look fine even while the process behind it was not under control.

Telling the two failures apart in your own readings

You cannot always tell precision from accuracy just by looking at one session's numbers, but there is a practical check. If three readings taken carefully agree tightly, that tells you the technique was repeatable that day - nothing more. Comparing the number against a known method, such as re-checking the landmark and pressure step by step against the full protocol, is the only way to test whether a tight cluster is also sitting near the true value, because tightness alone cannot answer that question. An imprecise but plausibly accurate set of readings, by contrast, is easier to catch: a wide spread across three attempts is visible on its own, without needing an external check, which is one reason precision problems tend to get noticed and accuracy problems tend to hide.

Different fixes for different failures

A precision problem needs technique locked down: hold the ruler the same way every time, wrap the tape with the same tension, use the same landmark. Once readings agree with each other, precision is solved, whatever the accuracy turns out to be.

An accuracy problem needs the method itself corrected, not just repeated more consistently: press to bone if you were not, measure along the dorsal line if you were not, use the state the literature reports. The full method exists specifically to remove the accuracy problems, and no amount of repeating a flawed technique consistently fixes them - consistency just hides them behind a tighter-looking number.

These are not the same failure as systematic versus random error, though the two ideas sit close together: a systematic error is what usually produces an inaccurate-but-precise cluster, and random error is what usually produces an imprecise one.

Why this distinction earns its place

A reader comparing their own number to the published figures needs to know which failure, if either, they are looking at, because a precise-looking self-measurement is not automatically a trustworthy one - the tightness of a cluster only ever tells you about precision, and an image-based estimate has neither property in this sense, since it never repeats a physical reading at all. A rating from a service like Penis Rater is a different kind of output again, a position among photographs rather than a distance with a target to hit, and the accuracy-versus-precision frame does not transfer to it cleanly. Nor does it transfer to a human judge's read, which has no repeatable numeric target to be precise or imprecise about in the first place. On a ruler, though, the distinction is exactly the one worth checking before trusting a number: ask whether your readings agree with each other, and separately, whether your method gives them any reason to be near the truth.

Read next

Full archive