Gear
Two instruments, one measurement
Taking girth with a tape and again with a paper strip separates instrument error from technique error, because the two instruments do not share failure modes.
A single reading tells you a number. It does not tell you whether that number is close to true, because you have no second reading to compare it against.
What "a second instrument" actually buys you
Take girth with a cloth tape. Then take it again, same session, same site, with a strip of paper wrapped and laid against a ruler. If the two agree to within a couple of millimetres, you have real evidence the reading is sound. If they diverge by more, one of the instruments is doing something the other is not.
The reason this works is that a tape and a paper strip fail in different ways. A tape can stretch with age, and stretch reads long. A tape has a metal end tab that shifts the printed zero if you use the tab edge instead of the printed mark. A tape's own thickness sits one layer out from the skin, which adds a small amount to every wrap. A paper strip has none of those failure modes. It cannot stretch under tension the way fabric or vinyl can, and there is no tab to misread. Its own error is different again: the width of the pen mark where the strip overlaps itself, and creasing if the strip is handled roughly. Tape or ruler covers the instruments individually; this is about running two of them side by side.
Two instruments with unrelated error sources that agree are unlikely to be wrong in the same direction by coincidence. That is the whole argument, and it does not require any statistics beyond noticing whether two independent numbers land close together.
What it does not catch
This method only separates instrument error from everything else. It does nothing for technique error, because technique is not the instrument - it is you, and you bring the same habits to both tools.
If you pull the tape too tight, you will likely pull the paper strip too tight as well, for the same reason: your sense of "snug" is the thing that is off, not the tape. If you measure at the wrong site on the shaft, both readings will be taken at that same wrong site, because it is the site you chose, not the tool you chose it with. If your erection state has drifted between readings, the second instrument sees the drifted state just as clearly as the first one did. Where measurement error comes from breaks down how much of a typical spread is technique and state rather than the instrument, and it is most of it - a second instrument closes a smaller gap than people expect, which is worth knowing before you buy a second instrument expecting it to fix a number that technique produced.
So a second instrument is a check on the tool, not a check on the person holding it. Agreement between two tools rules out one category of error and says nothing about the others. Disagreement is more useful than agreement, in a way: it tells you specifically that something instrument-side is wrong, and which of the two readings is more likely to be the artefact, based on which instrument has the failure mode that would explain the gap.
When it is worth doing
Once, when you are establishing a baseline figure you intend to keep and refer back to - the kind of number that goes in a measurement log rather than a passing curiosity. Not every session. Running two instruments every time doubles the effort for a check that, once passed, does not need repeating unless you change instruments or suspect one has degraded. A tape left in a drawer for a year is worth re-checking against a second instrument before you trust it again, because tapes stretch and rulers wear over time in ways you cannot see happening.
The same logic extends past girth. Length taken with a rigid ruler has no equally good second instrument - a tape is the wrong tool for length because it bends rather than transmitting bone pressure, so cross-checking length this way is not available to you the way it is for girth. For length, the closer equivalent is a second session on a separate day, which catches state and technique drift rather than instrument drift, and how to measure properly covers what to hold constant when you do that.
Not the same question as a rating
None of this produces a judgement, only a number with more confidence attached to it. A subjective score, like the ones an image-based rating service such as Rate Cock produces, is not something a second ruler improves - it is a different kind of output built to answer a different question, and cross-checking instruments has nothing to say about it, the same way a second photo judged against a scoring rubric does not converge on a length the way two rulers converge on one. The same separation holds for a human opinion: a person weighing in, the way sites like Rate Penis are set up for, is offering an impression rather than a measurement, and no amount of instrument agreement changes what that impression is worth. An AI system estimating from a photograph is answering yet another question again - AI Penis is honest about what an image model can and cannot recover, and it is not a length in centimetres either way. A tape and a paper strip agreeing with each other only ever tells you about the tape and the paper strip.