Evidence Note
Evidence summary
- Question
- Puberty blockers: what does the evidence actually show — and what is it that people are really arguing about?
- Overall certainty
- 🟡 Mixed — what the drugs do is firm; whether they deliver the benefit they are given for is low-certainty
- Clinical relevance
- High
- Reading time
- 5 minutes
- Companion review
- Available →
Almost every public argument about puberty blockers is really two arguments wearing one coat. One is about evidence: what do the studies show, and how much can we trust them. The other is about policy: what should be done about it. They get welded together constantly, and the weld is where most of the shouting happens.
Pulled apart, the picture is less dramatic than either side would like — and more useful. On the evidence there is more agreement than the volume suggests. On what should follow from it, there is a real disagreement that no study can settle.
What is the question?
Gonadotrophin-releasing hormone agonists — “puberty blockers” — pause puberty in adolescents experiencing gender dysphoria or incongruence. This note asks what the evidence supports, outcome by outcome, and where the certainty runs out. It is the short version of the full review, which grades each claim rather than restating it. It does not recommend for or against the treatment, and it is not a policy position.
What does the evidence say?
They do what they are designed to do — at the hormonal level. Once suppression is established, the gonadally-driven progression of puberty largely halts. Two honest qualifiers: changes already made do not reverse — a voice that has broken stays broken — so this is prevention of further change, not undoing; and because the drug acts on the gonads, not the adrenal glands, some pubic hair, acne and body odour can still advance.
The benefit they are principally given for is among the outcomes the evidence supports least well. Blockers are given, above all, to relieve psychological distress. The foundational Dutch study found general functioning improved — but gender dysphoria itself did not resolve during suppression. The most direct attempt to test the claim, a UK study, found no significant group-level change. The studies reporting clear benefit almost all measure blockers and subsequent hormones together, so they cannot tell you what the blocker did on its own. That is a statement about the evidence, not a verdict on the treatment.
Bone: a real effect, routinely overstated in both directions. Bone-density accrual slows during treatment. It largely recovers once gender-affirming hormones follow — though a shortfall at the lumbar spine can persist in those taking oestradiol. Whether any of this leads to more fractures later in life has never been adequately studied. Adult height, by contrast, appears broadly unaffected.
Fertility is more layered than “sterilising” or “fully reversible”. The blocker alone does not destroy the gonad. But starting early and going straight on to gender-affirming hormones may leave no window for conventional fertility preservation, because the conventional methods need a maturity that early suppression prevents. And the decisive study does not exist: no one has followed adolescents who began suppression in early puberty through to attempting conception in adulthood.
The famous “98%” does not mean what it is used to mean. The number comes from a Dutch cohort of people who started blockers and went on to hormones — 704 of 720 were still taking those hormones at follow-up. It measures how many stay on hormones once they start them. It is not a measure of how many blocker-starters go on to hormones in the first place. Used as though it were, it is a denominator error — and it is quoted that way constantly, by people arguing in both directions.
Neurocognitive effects: genuinely unknown. Two small human studies point in opposite directions and neither can settle it; the clearest adverse signal is from sheep. “No cognitive impact” overstates the reassurance; “irreversible brain damage” overstates the alarm. The honest word is unknown, and the reason is specific — nobody has tracked cognition from before treatment through into adulthood.
How strong is the evidence?
Split it, because the halves are not equal. What the drugs do — suppress the hormonal axis, pause pubertal progression, slow bone accrual — is on firm ground. What they achieve — psychological benefit, long-term safety, what happens to fertility and cognition — is where the certainty runs out. The systematic review commissioned for the Cass Review and the NICE appraisal both found the evidence low quality across the outcomes that matter most. A more recent review read the mental-health signal more favourably; that disagreement turns on which studies you include and whether evidence combining blockers with hormones can answer a question about blockers alone. It does not settle the point.
Where are the uncertainties?
Two of them are worth naming precisely, because they are where both sides overreach. The first is the difference between absence of evidence and evidence of absence: weak evidence of benefit is not proof the treatment fails, and weak evidence of harm is not proof it is safe. The second is the false absolute — “reversible” and “irreversible,” “safe” and “harmful” are wielded as though the evidence licensed certainty at either pole, when for most outcomes it licenses neither.
And the deepest uncertainty is the one the argument turns on. Very few who reach hormones later stop. Does that show the right young people were carefully selected — or that starting suppression makes continuing more likely? Both readings are coherent. The same number fits each. To tell them apart you would need to know what would have happened to those same adolescents had they not been treated, and nobody has observed that.
“The evidence is weak” is a claim about evidence. “Therefore it should be banned” is a claim about values. The first does not carry the second — in either direction.
What does this mean for practice?
It means holding two things at once, which is harder than picking a side. The evidence question has an answer: for the outcomes that weigh most on the decision, certainty is low. The policy question — what should follow from evidence of that quality — is about how much precaution a treatment for minors warrants, weighed against the risk that withholding it also carries. That is a judgement about values, which is why well-resourced health systems, reading substantially the same evidence, have reached opposite conclusions. It is not that one of them is following the evidence and the other isn’t.
For anyone actually weighing this — a clinician, a family, a young person — the honest posture is to hold uncertain benefit against uncertain harm openly, to counsel fertility and bone explicitly and early, and to resist the false comfort of a confident answer in either direction. None of this is a reason to start, stop or withhold treatment on your own: these are decisions to make with a clinician who knows the young person, and this note is not a substitute for that conversation.
This is the short version. The full evidence — every outcome graded, the 98% denominator error set out in detail, and the policy divergence characterised without being adjudicated — is in the companion Evidence Review →