Data
One method change, one shifted mean
Switching a protocol from non-bone-pressed to bone-pressed, or from stretched to erect, moves an entire study's mean by an amount that dwarfs its confidence interval.
Two studies can measure the same kind of population, with the same rigor, and still report means that differ by a centimetre or more, for no reason more exotic than one of them changed the protocol. That shift is systematic, not random, and systematic shifts do not shrink with sample size the way sampling noise does.
Random error shrinks with more people, systematic error does not
A confidence interval narrows as a study enrols more participants, because it is bounding how far the sample mean is likely to sit from the true population mean given random variation between individuals. A protocol difference is not random variation between individuals. It applies to every participant in the same direction, every time, so no amount of additional enrolment corrects for it - a larger sample measured the wrong way just gives you a more confident estimate of the wrong quantity.
A worked, hypothetical illustration
Say a hypothetical study of 200 men measures non-bone-pressed length and reports a mean with a narrow confidence interval, tight because the sample is reasonably large. Now suppose a second hypothetical study measures a similar population, this time bone-pressed, and its mean comes in a centimetre lower. Both intervals might be a few millimetres wide. The gap between the two means is several times wider than either interval, and none of that gap is explained by sampling error - it is explained entirely by the fact that one protocol includes the pre-pubic fat pad in the reading and the other does not. These figures are illustrative arithmetic, not drawn from any specific paper, but the shape of the effect is the real one the bone-pressed method is built to control for.
Why this is why methods sections matter more than sample size
A reader comparing two published means without reading the methods section is comparing two numbers that may not be measuring the same thing. Sample size tells you how precisely a study estimated its own mean. It tells you nothing about whether that mean is comparable to a different study's mean, which is a question the protocol answers and the confidence interval cannot. Not every study in the pooled literature specifies bone-pressed measurement consistently, which is exactly this problem showing up inside a single meta-analysis rather than between two separate papers.
The same logic applies to stretched versus erect
A study reporting stretched flaccid length and one reporting erect length are also not directly comparable means, even though stretched length correlates reasonably well with erect at the individual level. The average gap between the two states, applied across an entire sample, produces the same kind of coherent, whole-sample shift as the bone-pressed change does - it just runs along a different methodological axis. Why studies disagree with each other more broadly covers population, instrument and state alongside method; this post is about the mechanism by which one changed variable moves an entire mean rather than just adding noise around it.
Why this fools careful readers, not just careless ones
The trap here is not a reader skipping the methods section out of laziness - it is that two well-designed, well-powered studies can each be internally excellent and still produce means that should not be placed side by side. Nothing about statistical rigor within a study protects against a difference introduced by the protocol itself. A reader who checks sample size, confidence interval and p-values, and stops there, has checked everything except the one thing that actually explains the gap.
What to check before comparing two means
Read the methods section, or at minimum the state and landmark description, before treating two published averages as comparable. A confidence interval answers "how precisely did this study estimate its own number." It does not answer "is this study's number the same kind of number as that other study's number," and conflating the two is the single most common way a size figure gets misread.
This is a distinctly different problem from the ones a rating platform runs into. Rate Cock produces a photo-based rating rather than a measured mean, so a shift in scoring criteria there changes what a score represents in a more direct, immediate way than a measurement protocol does. Penis Rater is upfront that its tools are scoring instruments rather than measuring instruments, and the same distinction applies to a human's opinion on Rate Penis or an AI system's read of a photo on AI Penis - none of the three are quoting a mean length at all, so this specific failure mode does not touch them, though each has its own version of the same underlying lesson: know what produced the number before you compare it to another one.