Data
Fewer erect measurements
Far fewer men in the pooled data were measured erect than flaccid or stretched, which is why the erect figures carry wider uncertainty.
Within the studies pooled by Veale et al. (2015), BJU International, not every man contributed a reading in every state. Flaccid and stretched flaccid measurements are far more common across the pooled sample than erect ones, and that imbalance carries directly into how confident the reported erect figures can be.
Why the erect count is smaller
An erect measurement requires a full erection obtained and held in a clinical setting, which is a harder thing to arrange as a matter of course than a flaccid or stretched reading. Some studies obtained it pharmacologically, some relied on self-stimulation in private, and some used self-report of an at-home reading - three genuinely different approaches, each with its own practical limits on how many participants a study could put through it.
Flaccid and stretched flaccid readings need none of that. They can be taken quickly, in the same visit as almost any other clinical measurement, which is why studies that were not even primarily about size sometimes still recorded a flaccid or stretched figure in passing, while an erect reading required a study built specifically to obtain one.
What a smaller pool does to the estimate
A mean calculated from fewer data points has a wider confidence interval around it than a mean calculated from more, all else equal. This is a general statistical fact, not something specific to size data, but it applies here directly: the erect mean in Veale 2015 rests on a smaller slice of the total 15,521 men than the flaccid or stretched means do, so it is the figure with the least statistical cushion of the three.
The mean itself is not necessarily biased by this - a smaller but still substantial sample can still centre on the right value. What changes is how tightly that value is pinned down, and how much a single additional study could still move it.
What this means for reading the erect figure
Treat the erect mean as the best available estimate, not as a number with the same statistical weight as the flaccid or stretched figures sitting next to it in the same table. It is still the most useful erect figure available, drawn from more erect measurements than any single study on its own has managed, but it is the one figure in that table with the most room left for a future large study to shift.
This is a data point about sample composition, not about method. How each of the three states is actually obtained in a clinical setting is a separate question from how many men contributed to each pool, and answers it in more detail than belongs here.
It is also a different kind of uncertainty from the one attached to a subjective rating or a human judge's opinion - those carry no sample size at all in the same sense, because they are not estimating a population mean. And it has nothing to do with what an image model can infer from a photograph, which faces a different limitation regardless of how many images it has seen. A public distribution of ratings, like the kind a scoring hub displays, answers yet another kind of question again.