Method
Mean or median
With three or five readings the median resists one bad reading; with many readings the mean uses more of the information. Use the median for small sets.
With a small set of readings, use the median. With a large set, the mean is the better choice. The reason is what one bad reading does to each.
Why the median wins on a small set
The mean uses every reading and gives each one equal weight, which is exactly the problem when the set is small: one reading thrown off by a slipped ruler or a moment of bad tension pulls the mean toward it, and with only three or five numbers, one bad one is a large fraction of the total. The median - the middle value once the readings are sorted - ignores how far off an outlier is, only where it ranks. A single bad reading at the extreme still leaves the median close to where it would have been without it.
A worked example
Three readings: 14.2, 14.4, and 15.1 cm, where the last one came from a session where the tape slipped. The mean of the three is 14.57 cm - dragged upward by the bad reading. The median is 14.4 cm - the middle value, unaffected by how far off the third one landed. With only three numbers, that difference is not trivial.
When to switch to the mean
Once you have enough readings - five becomes a judgement call, and by ten or more the mean is clearly better - the mean starts using information the median throws away, because it accounts for how far every reading sits from the centre, not just its rank. With enough readings, one outlier also has less leverage on the mean simply because it is diluted by all the others. How many readings to take in the first place is a separate question from which average to take once you have them, and for the 3 x 3 protocol most people actually follow, the answer here is the median. Both questions assume you are taking more than one reading at all - what a single reading actually tells you is worth reading before deciding that one number is enough.
What this does not decide
Choosing an average settles what number to write down. It comes after the readings exist, not before - how many readings to take in the first place and how to round the average once you have it are the two questions that bracket this one. Choosing an average also says nothing about how that number is perceived by anyone looking at the result, which is a judgement call rather than an arithmetic one - the kind a human reviewer on a site like Rate Penis makes, not something a median or a mean produces. It is equally irrelevant to a photograph, which carries no set of repeated readings to average in the first place - whatever an AI tool estimates from one image is a single inference, not a statistic over repeats. A photo-based composite score, like the ones a site such as Penis Rater compiles, is built from a different kind of input entirely. And a subjective rating from somewhere like Rate Cock was never a candidate for averaging in this sense - it is a different kind of number from the one three readings and a choice of median produce.