Data
When the reward depends on the number
Some self-measured studies gave participants a product sized to their reported figure; that design changes who takes part and how carefully they measure.
Most self-report studies just ask for a number and record it. A smaller line of research does something different: it sends the participant a product sized to whatever they report, most often in fit research for products like condoms. That single design choice changes the incentives around the measurement itself, and it is worth separating from the ordinary self-report problem.
Two pulls in opposite directions
When a reported figure determines what arrives in the mail, a participant has a reason to get the number right that a plain survey respondent does not. An anonymous online survey rewards nothing for accuracy - reporting a larger number costs nothing and may flatter the respondent, which is a large part of why self-reported figures run high. A fit-driven study changes that calculation, because a wrong number produces a product that does not fit.
But the pull does not run only toward accuracy. Ego still has a vote even when a product is at stake, and a participant choosing between "the honest measurement" and "the flattering measurement" may still round toward the second, especially if the instructions allow some latitude in how the measurement is taken or if the fit tolerance is wide enough to absorb a small exaggeration. The honest answer is that the net effect is not established in a way this site can quote as a magnitude. What is defensible is the direction of two competing forces, not a single corrected number.
Who volunteers changes too
A study offering a sized product also selects its participants differently than a study offering nothing. People willing to measure themselves and report the result for a free or discounted item are not drawn evenly from the population, and that selection effect sits on top of whatever the incentive does to individual honesty. It is the same participation problem that runs through volunteer-based size research generally, sharpened by the fact that the "prize" here is specifically tied to the trait being studied.
How to read a study built this way
Three questions are worth asking of any self-measured, product-fitted dataset before treating its mean as informative: was the measurement instruction detailed enough to be repeatable, was there any independent check of a subsample against a clinician measurement, and does the paper report on who declined to participate as well as who did. Few papers in this space answer all three. That does not make the data worthless - it makes it a different category from the clinician-measured pool that Veale et al. (2015) assembled, and the two should not be averaged together or cited as if they carry the same weight. The condom-fit literature specifically is its own dataset with its own conventions, and it is worth reading on those terms rather than folded into the general size picture.
Not the same question as sizing
None of this is about how a product like a condom gets sized once the data exists - that is a separate, narrower question about converting a girth figure into a nominal width, and it belongs elsewhere. It is also not a verdict on self-report as a category; the broader mechanism of why unincentivized self-report runs high is covered on its own. This post is narrower: what happens to a self-measured figure specifically when the number determines a reward.
A method note that gets skipped here as everywhere: a self-measured figure, incentivized or not, is still self-measured, and the instrument in someone's own hands at home is not the instrument a clinic uses. That is a design fact about the study, not a judgement about the people in it. If what you actually want is a subjective read rather than a study design to interrogate, that is a different kind of question - a photo-based score from a service like Rate Cock is answering something else entirely, not this one, and neither an AI model's read on a photo nor a rating drawn from uploaded photos is a self-measured figure in a study at all. Nor is a human judge's verdict - all three are different data sources built for different purposes, and none of them is what this post is describing.
Take the incentive structure as one more variable to log when you see a self-measured mean, the way you would log method or state on your own reading. A number without its collection conditions is not comparable to anything, whichever way the incentive happened to push it.