Eden Openly

Masculinising and Feminising Hormones in Adolescents

Evidence Note

Evidence summary

Question
What is actually known about masculinising and feminising hormones in adolescents — and what has been assumed?
Overall certainty
🔴 Very low on every psychological outcome. 🟢 Firm on the physical changes. And on the treatment’s central intended outcome, weaker than anyone has been told.
Clinical relevance
High — a national policy decision is pending
Reading time
8 minutes
Before you quote it
What this review does not conclude →
Companion review
Available →

Three formal appraisals — NICE in 2020, Taylor for the Cass Review in 2024, and NHS England’s ten evidence reviews in 2026 — have reached the same verdict about hormones for adolescents: the evidence is of very low certainty. That finding is correct, and this note does not dispute it.

What this note is about is what happened to the evidence on the way to that verdict. Because when the primary studies are examined alongside the evidence summaries, several things turn out not to be what the chain has been passing along.

What is the question?

The clinical question is narrow: what do testosterone and oestradiol do to an adolescent, how certain are we, and how much of it lasts if treatment stops?

The full review grades every outcome on four axes — did the treatment cause it, how big is the effect, would a patient notice, and was it measured in adolescents at all — and keeps them apart, because those questions are often conflated.

What does the evidence say?

The physical changes are firm. Almost nothing else is. Testosterone reliably produces voice deepening, facial and body hair, clitoral growth and menstrual suppression. Oestradiol reliably produces breast development. These are not in doubt, in any population, and nobody seriously disputes them.

The psychological evidence is observational, uncontrolled and of very low certainty. The largest prospective study — 315 adolescents and young adults, two years, no comparison group — found depression falling by about 1.27 points a year on a 63-point scale. Whether that is a benefit, a trivial change, or regression to the mean cannot be determined from the design.

And nobody has agreed what “better” means. NHS England’s own PICO states, in terms: “There are no known minimal clinically important differences and there are no preferred timepoints for the outcome measures selected.” A minimal clinically important difference is the smallest change on a scale that a patient would actually notice or a clinician would act on — the thing that converts a number into a meaning. NHS England is saying that no such threshold has been agreed for these outcomes. It then commissioned ten reviews to measure them. So the field is measuring change against no threshold of meaning — and has been saying so for six years.

Three things this review found that were not in the summaries

The errors are almost never in the primary studies. They are in what happens to them afterwards.

1. The one study measuring the treatment’s central outcome changed its measuring instrument halfway through.

Gender dysphoria — the thing the drugs are actually for — has been measured in adolescents on hormones exactly once, in a study of twenty-three people. Its headline figure is a fall from 57.1 to 14.7 on the Utrecht Gender Dysphoria Scale. That 42-point drop is quoted everywhere, by everyone, on both sides.

The scale has two forms, chosen by sex assigned at birth. The paper’s Methods state that it was “filled out for the sex assigned at birth at T0 and the self-identified gender at T1.” Participants completed one form before treatment and the other one afterwards. The two numbers are not on the same instrument, so the difference between them is not a change score.

The authors disclosed this plainly. NICE — whose only positive conclusion is that hormones are “likely to improve symptoms of gender dysphoria” — did not report it. Taylor did not report it. NHS England inherited it. The limitation was not carried forward into any of the subsequent evidence summaries — including, until 12 July 2026, this one.

2. An adverse effect affecting roughly one in four adolescents was never asked about.

In a cohort of 158 transmasculine adolescents on testosterone, 37 (23.4%) had documented pelvic pain, beginning a median of six weeks after starting. Where intensity was scored — in only eleven of them — ten were severe.

NHS England’s safety outcome list is long and specific. It names pulmonary oil microembolism. It names seizures. It names sleep apnoea. It does not name pelvic pain. The evidence review picked it up anyway, because the study happened to report it — but an outcome that is never asked for gets no certainty rating, no synthesis, and no place in the argument. It remains outside the structured synthesis.

3. The evidence base has no outcome for what happens when treatment stops.

All ten of NHS England’s PICO documents state that these medicines “may mitigate the unwanted endocrine and metabolic effects of hypogonadism.” The document covering blockers-without-replacement lists what hypogonadism does, by name: hot flushes, night sweats, headaches, muscle pain, reduced libido, reduced bone density.

And not one of the ten reviews asks what happens when those medicines are withdrawn from a young person taking them — which produces exactly that state. The only outcome about stopping is “detransition,” defined across all ten as something a patient “may choose”, with nine listed synonyms, every one of them a word about identity and not one about physiology.

The largest study of adolescents who actually stop — 1,050 of them — was excluded from the review as out of scope. It found that 93% were still taking hormones at last contact, and that 0.5% stopped because they no longer identified as transgender. The commonest reasons were having achieved their goals, and difficulty getting or taking the medication. The nine synonyms cannot record either of those.

And it happens to the harm findings too. All three examples above weaken a claim of benefit, and a reader is entitled to ask whether this review only found the errors that suited it. So here is one that runs the other way.

The study behind NHS England’s cardiometabolic evidence reports that transgender young people had their cholesterol checked roughly three times as often as the people they were compared with — 39% against 12%. Its outcomes are defined as two abnormal test results. Test one group three times as often and you will find more abnormal results in them, whether or not anything is wrong. The authors say so, plainly, in the paper, and they use it to explain their own findings.

NHS England’s evidence review copies the definition and leaves out the warning. The same shape as the dysphoria scale: the researchers disclosed the problem, and the document that summarised them did not carry it. Correcting this one does not strengthen the case against these medicines. It weakens it.

That is the point. The errors are not on one side. They are wherever a summary was read in place of a source — and they were found by opening the papers, not by deciding in advance which way they ought to fall.

How strong is the evidence?

Very low, and this review agrees with the appraisals. No randomised trial of hormones in adolescents exists. No study has an untreated comparison group capable of establishing cause. The largest studies are uncontrolled, single-centre, and follow people for one to two years.

Two things worth knowing about how that verdict is usually described:

NICE’s “very low certainty” and Taylor’s “moderate-quality evidence” are not answering the same question. NICE is rating confidence in an effect. Taylor is scoring the studies on a checklist — moderate means anything above 50% and up to 75%. A study can score 60% and still be a before-and-after series with no comparator, incapable of attributing anything to anything. These two ratings are not in conflict, and they are not in agreement.

And the figure of “40,906 participants” is misleading. Taylor’s own next sentence says: 8,164 received hormones. The rest are untreated adolescents and comparison groups. The evidence base for the treatment is five times smaller than the figure that is often quoted.

Where are the uncertainties?

Many of the outcomes most important to a long-term decision remain uncertain, and several have never been studied at all. Lifetime regret: no study of adequate duration exists. Fertility in transmasculine adolescents: no study at all. Fractures — the outcome that bone-density scores are a surrogate for: never measured. Cancer: the latency is decades; the cohorts are years old.

And two things this review corrected in its own drafts. The finding that early puberty suppression preserves fertility better than late is confounded — the adolescent and adult groups in that study received entirely different drugs, and the authors say so. And the finding that early suppression adds height rests on a crude difference of 1.9 cm which was not statistically significant, expanded fourfold by a statistical model across two groups that were never comparable.

An earlier version of this note said those two corrections “cut against the case for treatment.” That was asserted rather than worked out, and it was wrong about the first one. The fertility correction cuts both ways: it removes a claim that early suppression forecloses fertility, and it removes the credit given to early timing for preserving it. Both are true, and they point in opposite directions. The review’s correction record now states the direction of a correction only where that direction can actually be derived, says so where a correction cuts both ways, and gives no tally. A correction is kept because it is faithful to the source — not because of which way it cuts.

What does this mean for practice?

For clinicians. The physical effects are reliable and can be discussed with confidence. The psychological effects cannot. Pelvic pain in transmasculine adolescents on testosterone is common, often severe, poorly documented and badly treated — and it is not on anyone’s outcome list, so ask about it. And nobody has measured what stopping does.

For patients and families. “Very low certainty” does not mean the treatment doesn’t work. It means the studies cannot tell you how much, for whom, or for how long. Anyone who tells you the evidence proves benefit is overstating it. Anyone who tells you the evidence proves harm is overstating it in the other direction. Both are reading a verdict about the studies as though it were a verdict about the drugs.

For anyone reading the policy documents. Open the primary papers. This review identified sixteen apparent errors in NHS England’s evidence reviews — transposed drug names, a final adult height lower than the baseline height, a z-score of 21.0, and comparisons reported as statistically significant where the printed figures appear inconsistent with that interpretation. All sixteen were reported to NHS England on 12 July 2026. At the time of writing they have neither been accepted nor disputed. Correcting them does not change the certainty ratings, and several of them, corrected, would strengthen the case for caution.

In short. The physical changes are firm. The psychological evidence is very low certainty, and that is a fact about the studies, not about the drugs. The only published adolescent measurement of the treatment’s central intended outcome cannot be read as a conventional before-and-after change, because different versions of the instrument were used before and after treatment. An adverse effect affecting one in four was never asked about. And a policy proposing that adolescents stop treatment rests on an evidence base that never asked what stopping does.

The research is uncertain. That uncertainty is real, and correcting these errors does not remove it. But uncertainty should not be compounded by avoidable errors in how the evidence is summarised. Correcting them does not make the evidence stronger. It makes the uncertainty itself more accurately described.

This is the short version. The full evidence — all twelve outcomes graded on four axes, the persistence ledger, the instrument switch set out in detail, and the sixteen errors reported to NHS England — is in the companion Evidence Review →