Method
How the erect figures were taken
Erect measurements in the literature come from pharmacologically induced erections, self-stimulation in private, or self-report of an at-home reading; each is a different dataset.
An erect measurement in a published study did not come from a man arriving already erect and being measured on the spot. Producing a comparable erect reading requires reaching that state under some kind of controlled circumstance, and the literature has used three different ways of doing that, each of which hands back a slightly different kind of number.
Pharmacologically induced
The most controlled route is an erection induced by injection, administered by a clinician as part of a diagnostic procedure, with the measurement taken as a secondary data point once full rigidity is reached. This is the version closest to a lab measurement: a trained observer, a fixed protocol, a state that does not depend on the participant reaching or holding arousal on demand. It is also the smallest and most selected dataset of the three, because injected erections are typically only obtained from men already undergoing that procedure for a clinical reason unrelated to a size study in the first place.
Self-stimulation in private
A more common route in dedicated size studies is self-stimulation to full erection in a private room at the clinic, with a health professional taking the measurement immediately afterward. This keeps the measurer - and so the technique, landmark and reading - clinical and controlled, while leaving the erection itself to the participant, which is closer to how the state is reached outside a study. Firmness at the point of measurement still has to be full for the reading to be usable, and a study using this route depends on the participant reaching and holding that state for long enough to be measured properly.
Self-report of an at-home reading
The third route removes the clinician from the room entirely. A participant measures himself at home, following written instructions, and reports the figure back to the study. This is the least controlled of the three - nobody is checking the landmark, the pressure, or whether full rigidity was actually reached - and it is also the route most studies without funding for an in-clinic protocol have had to rely on. Self-reported figures run high for reasons that are well documented, and a self-reported erect figure carries all of that bias on top of whatever the participant's actual state was at the time.
What a reader can and cannot tell from a mean
A published table rarely spells out which route produced its erect figure in the sentence that quotes the mean - that detail usually sits in the methods section, if it is stated at all. This is worth checking before comparing a figure to your own reading, because the three routes do not simply add noise around a shared true value; they can shift a mean in a consistent direction. A self-reported dataset, for instance, is not just a noisier version of a clinician-measured one - the same self-report pressures that inflate other figures apply here too, so a self-reported erect mean tends to sit above a clinician-measured one for reasons that have nothing to do with the underlying population being different.
Sample size and route are also linked in a way that matters. The clinician-measured routes are more trustworthy per participant but recruit far fewer men, because they require either a clinical justification for the induced-erection route or a private room and staff time for the self-stimulation route. Self-report scales more easily and so tends to dominate the largest surveys, which means the biggest sample is not automatically the most trustworthy one - a small, tightly controlled study and a large, loosely controlled one are answering the same question with different reliability, and size alone does not settle which to weight more.
Why the route matters
These are not three measurements of the same underlying number taken with more or less care. They are three different kinds of dataset, shaped by who is holding the ruler, how the erection was reached, and how much of the process anyone other than the participant can verify. Reading how a given study was designed means knowing which of these three routes it used, because the route sets a ceiling on how much a reported figure can be trusted before a single number is even looked at. Veale et al. (2015), the meta-analysis this site returns to most often, pools studies that used more than one of these routes, and it says so.
None of this is a question a photograph or a rating can answer either way. Rate Cock and services like it score an image against other images, not reconstruct which of these three routes produced a state - that is a different exercise from a measurement entirely, as an image-based accuracy discussion elsewhere makes clear. A subjective score, whether it comes from a photo-rating service or a human judge's opinion, does not carry a route at all. It is answering what something looks like, not how a number was obtained under a protocol, and the two questions do not substitute for each other.