Eden Openly

Masculinising and Feminising Hormones in Adolescents

Evidence Review · ER-015

Identifier: ER-015 · Version: 1.4 · Status: Current · Published: July 2026 · DOI (all versions): 10.5281/zenodo.21324879 · DOI (this version): 10.5281/zenodo.21364084

This is the full review — every outcome graded, every limitation stated, and the complete evidence behind them. A short companion Evidence Note distils it — see the Evidence Note → (8 min read). It is the companion to Puberty Suppression in Adolescents (ER-014), which grades the blocker.

Before you quote this review — or anyone quoting it — read what it does not conclude → It does not conclude that these medicines do not work. It does not conclude that they are safe. It does not recommend withdrawal, and it does not recommend unrestricted access. One page, and any quotation can be checked against it in thirty seconds.

This review grades the evidence on masculinising and feminising hormones — testosterone and oestradiol — in adolescents with gender dysphoria or incongruence. It is the companion to ER-014, which graded the blocker, and it reaches a conclusion of the same shape. The drugs reliably induce many of the physical changes for which they are prescribed. The evidence that those changes translate into reduced gender dysphoria or improved psychological functioning in adolescents is observational, uncontrolled, heterogeneous and of low certainty — and on the central outcome of gender dysphoria it consists of a single study of twenty-three people which changed the measuring instrument between the two assessments — insufficient to produce a reliable estimate of benefit, and equally insufficient to exclude one. Where this review differs from its predecessor is in what is at stake. The blocker was defended as a pause. These drugs are not a pause: several of the anatomical changes they produce substantially persist after treatment stops, and being exact about which ones is the central task of the review. It is one of the central factual questions in the policy argument, and it is the one both sides blur.

Purpose

To appraise the certainty of the evidence on masculinising and feminising hormones in adolescents with gender dysphoria or incongruence — what the drugs do, which of their effects are permanent, and how strong the evidence for each outcome really is — not to recommend for or against the treatment, and not to determine whether it ought to be provided. This is not a prescribing protocol and not a policy position. It should not be used to start, stop or withhold any treatment, and no clinical or policy decision should rest on it alone.

A word on what this piece is, and is not. This is a single-author narrative review. It is not a systematic review, it has not been peer reviewed, and it does not use formal GRADE. The certainty ratings are editorial judgements, made against a stated scheme, and they are open to challenge. The commitment is to grade the evidence symmetrically: to treat “the evidence is too weak to show benefit” and “the evidence is too weak to show harm” as the same kind of claim, and to hold association apart from causation, and absence of evidence apart from evidence of absence, in both directions. A reader who supports access and a reader who supports restriction should each be able to check their own position against the same graded findings, and each should find something they did not want to read.

Scope. Oestradiol and testosterone in adolescents — in the United Kingdom, predominantly at sixteen and seventeen; internationally, sometimes younger. The sequential pathway (blocker, then hormone) is in scope because it generated much of the long-term evidence and because it is where several of ER-014’s open questions resolve. Partial and non-binary regimens are in scope and are not collapsed into the binary case — though, as §7 sets out, the appraisals identified no eligible outcome evidence reported separately for partial or non-binary regimens, so nothing can be graded for them here. Adult feminising therapy (ER-005), adult masculinising therapy (ER-007), puberty suppression alone (ER-014), anti-androgens as adult monotherapy (ER-009) and surgery are all out of scope.

If you are affected by this. This is a subject that touches many people closely — young people, families and clinicians alike. If you or someone you care for is struggling, support is available, and it is worth reaching out to a qualified professional or a trusted person.

Part I — The frame

1. Two questions, kept apart

Almost every public exchange about these drugs runs two different questions together, and the confusion is not accidental — collapsing them is how each side converts a modest evidential claim into a sweeping conclusion. Pulling them apart is the whole method of this review, as it was of its predecessor.

The evidence question. What do the available studies show about the benefits and harms of giving testosterone or oestradiol to adolescents with gender dysphoria, and how much confidence can be placed in those findings? This is an empirical question, and it has an answer that neither side likes stated plainly. The drugs reliably produce many of the physical changes they are given to produce; several of those changes persist after treatment stops; and the evidence that those changes translate into reduced gender dysphoria or improved psychosocial functioning is observational, uncontrolled and of low certainty. Note that “benefit” is not a single endpoint: guidelines and consent documents variously describe the purpose as inducing desired secondary sex characteristics, relieving gender dysphoria, and improving wellbeing or quality of life. These are related but not interchangeable, they are measured with different instruments, and this review grades them separately in Part V rather than collapsing them.

The policy question. Given evidence of that quality, what should be done? Permit the treatment routinely, permit it only within research, restrict it, or withdraw it? This is not an empirical question. It is a question about how much weight to give uncertain benefits against uncertain and in some respects permanent harms; about how to treat an evidence gap when the alternative — untreated distress in a population with very high rates of self-harm — also carries risk; about who should decide; and about how much precaution a treatment for minors warrants. Reasonable, well-informed people reach opposite answers from the same evidence, because they weigh the uncertainties differently. That is a values disagreement, not a factual one. It does not follow that evidence is irrelevant to it. Values supply the weights; evidence supplies the probabilities and magnitudes those weights are applied to. Better data could rationally move a reader without changing anything they value. Evidence cannot settle the argument — but it can change the answer, which is a reason to gather it rather than to declare the question closed.

Why the seam matters. Watch how the two get welded. “The evidence is weak” is an evidence claim; “therefore the treatment should be withdrawn” is a policy claim, and the first does not entail the second — weak evidence can justify caution, more research, restriction or continued provision with consent, depending on values the evidence cannot supply. Run the weld the other way and it fails just as plainly: “young people report relief” is an evidence claim of limited certainty, and “therefore access should be unrestricted” is a policy claim it cannot carry. This review answers the evidence question and characterises the policy question without pretending the evidence settles it.

And here, unusually, the policy question is live. At the time of writing, NHS England has paused new prescriptions for sixteen- and seventeen-year-olds and consulted on withdrawing them as a routine commissioning option; the final policy has not been published. This review is written so that it stands whichever way that decision lands. If it reads, in retrospect, as having been written to anticipate one outcome, it has failed.

Confidence: not applicable — this section is the method, not a finding.

2. Three errors, named once so they can be watched

The evidence base is weak enough that both sides are tempted into the same mistakes, mirrored. Two were named in ER-014 and recur here. The third is specific to this review, and it is the one that catches careful people.

Absence of evidence as evidence of absence. Low-certainty evidence of benefit is not proof the treatment fails, and low-certainty evidence of harm is not proof it is safe. This error is about to be committed at scale, in both directions, and the wording that will do it is already in circulation: no evidence was identified. That is a statement about a literature, made within a defined methodological frame. It is not a statement about a drug. Section 4 sets out exactly what it does and does not mean.

The false absolute. “Reversible” and “irreversible”, “safe” and “harmful”, “settled” and “experimental” are wielded as though the evidence licensed certainty at either pole, when for most outcomes it licenses neither. This review reaches for a strong word only where a specific finding earns it — and, unlike ER-014, it does reach for one. Some of these changes are permanent, and saying so plainly is not a concession to one side. It is one of the central factual questions in the argument, and Part IV is devoted to it.

Mechanism as certainty — and the distinction that dissolves it. This is the new one, and it runs through every outcome that follows. Confidence that an effect exists, and in which direction, is a different object from certainty about the size of that effect. They come from different evidence and they can diverge wildly. That testosterone increases the risk of acne is strongly supported: the sebaceous gland is androgen-responsive, and androgen signalling is one well-established input into the sebum production, follicular keratinisation and inflammatory changes from which acne arises. How often, how badly, and in whom — for that, the adolescent evidence base is close to useless. High plausibility and low certainty, on the same claim, at the same time.

This is not a technicality. It is the reason a reader can be told, truthfully, both that a drug’s effects are well understood and that the evidence for them is very low quality, and conclude that someone is lying. Nobody is. The two statements are about different things. Where this review grades an outcome, it grades the estimate — not the plausibility — and where the two come apart, it says so.

Confidence: not applicable — this section is the method, not a finding.

3. How certainty is graded here

Four questions, kept apart. A single certainty rating cannot carry them, and collapsing them is how most public argument about this evidence goes wrong. Every confidence box in this review answers these four, in this order, and names any it cannot answer:

  1. Causal attribution. Confidence that the observed change was produced by the treatment rather than by something else — the surrounding care, the passage of time, regression to the mean, or who was selected into treatment. Answered by study design, not by the size of the change. Note the scope: the question is not “does testosterone cause virilisation” (it plainly does) but “did the treatment cause this observed improvement in this study.”
  2. Observed effect size. How large was the measured change? Answered by the data — and usually is estimable, even when the first question is not.
  3. Clinical importance. Would the observed change be noticeable, functionally meaningful, or decision-relevant to the patient? This can be assessed in several ways — validated severity thresholds, reliable-change indices, responder proportions, movement between clinical categories, patient-reported global change, functional or event outcomes, or an established minimal clinically important difference. An MCID is one method, not the definition. For most outcomes reviewed here, none of these has been adequately established in this population (section 9).
  4. Population directness. Was this measured in the population the claim is about? Answered by who was actually studied. A finding measured in adolescents is direct; one extrapolated from adults or from cisgender cohorts is indirect, and is weaker than it looks however large the study. This axis is graded in words rather than colours, because it is a property of the evidence’s provenance rather than of its strength.

These come apart constantly. A study can produce a precise effect estimate that establishes nothing causally, and whose clinical meaning is unknowable. That is not a hypothetical: it describes most of the evidence graded below.

Certainty ratings are per-claim, using the series scheme, as editorial judgements rather than formal GRADE:

Two clarifications, because both are routinely misread. Robust does not mean large. A 🟢 rating says the finding would be surprising to overturn — it says nothing about the size of the effect, and it is applied here to mechanism, to documentary fact, and to directly observed physical changes alike. And population directness is graded in words, not colours, on a fixed three-point scale: direct (measured in adolescents), mixed (the fact measured in adolescents, the detail borrowed), indirect (extrapolated from adults or cisgender cohorts). It is a statement about provenance, not about strength.

Population flags mark any finding extrapolated from a different population — most often adults, or cisgender cohorts. These flags are not decorative. Where a claim rests on adult data, the claim is weaker than it looks.

Part II — What the drugs are, and what they do

4. Mechanism, and why it is not the interesting part

Testosterone and oestradiol are the principal hormonal mediators of many pubertal changes, acting within a wider endocrine system — the hypothalamic–pituitary–gonadal axis, adrenal androgens, growth hormone and IGF-1, thyroid function and nutrition all contribute. Given exogenously, they act at the same receptors as the hormones the gonads make, across the larynx, the pilosebaceous unit, mammary tissue, adipose tissue, skeletal muscle, bone and the reproductive axis itself.

One refinement matters for the outcomes graded later. Several of the effects attributed to testosterone are not testosterone effects in any direct sense: in the pilosebaceous unit and the hair follicle, testosterone is converted by 5α-reductase to dihydrotestosterone, which binds the androgen receptor with substantially greater affinity.[33] Local conversion to dihydrotestosterone contributes importantly to several pilosebaceous and hair-follicle effects, although the relative contribution varies by tissue and by individual, and androgen action in these tissues also involves testosterone itself, receptor expression, co-regulators and other metabolites.[32,33] Voice, body hair, acne and male-pattern hair loss are therefore androgen effects rather than simple dose-effects of serum testosterone (ER-013), which is one reason serum testosterone correlates imperfectly with the outcomes patients notice.

Two caveats belong here rather than in a footnote. The preparations are not identical to endogenous secretion: most are esterified prodrugs — testosterone cypionate, enanthate, undecanoate; oestradiol valerate — and route and formulation produce peaks and troughs unlike continuous gonadal output. And the response depends on context: oestradiol induces breast development, but its degree varies with age, prior pubertal development, and how far endogenous testosterone is suppressed.

This matters for one reason, and it is the reason set out in section 2. The near-certainty of the mechanism does not transfer to the outcomes. That a feminising regimen induces breast development in the great majority of those treated is not in doubt; how much, how quickly, and in whom, is measured in adults and barely measured in adolescents. That testosterone causes voice deepening is not in serious scientific dispute; the adolescent-specific quantitative evidence on degree, timing and variability is limited, largely observational and inconsistently measured. This gap — firm mechanism, absent measurement — runs through Part IV and Part V and is the single most important thing to hold in mind while reading them.

Confidence: 🟢 Robust (that testosterone and oestradiol act through their cognate receptors and produce the secondary sex characteristics associated with them) / 🔴 Limited (the magnitude, timing and variability of any specific effect in this population — graded outcome by outcome in Part V).

5. The regimens, and the sequential pathway

The drugs, and how they are given. Common regimens use oral or transdermal oestradiol, and injectable or transdermal testosterone; in adolescents, subcutaneous and intramuscular testosterone have been compared directly and prospectively in a single small study.[3] Injectable oestradiol, oestradiol spray and testosterone pellets are used in some centres and jurisdictions but are not standard UK adolescent practice, and availability and licensing differ by country.[1,2]

Dosing is not uniform. For younger adolescents, and for those coming off puberty suppression, guidelines describe gradual escalation from sub-adult doses, intended to approximate pubertal induction.[1] That rationale does not extend automatically to a seventeen-year-old who has completed endogenous puberty, or to someone seeking a partial or non-binary regimen. Escalation strategies in practice are heterogeneous, and the evidence for non-standard and low-dose regimens is thinner still.

What adolescents are actually started on. A large published account of initiation practice is an Australian cohort of 158 transmasculine adolescents. Median age at testosterone initiation was 16.6 years. Just 15 (9.5%) had used a puberty blocker. Testosterone was started as a 3-monthly intramuscular undecanoate injection in 78 (49.4%), a 3-weekly intramuscular enanthate or mixed-ester injection in 63 (39.9%), and a daily topical preparation in 17 (10.8%). More than half — 84 of 158 (53.2%) — were started at what the authors classify as a high dose, defined as equivalent to standard adult male maintenance dosing (topical 32.5–50 mg; enanthate or mixed esters 250 mg; undecanoate 1000 mg). The low-dose arm was half of each.[5]

This is a single centre, retrospective, and outside the UK, and it describes 2007–2020. It is not a UK practice statistic and should not be read as one. It is included because it is the only cohort large enough to describe what initiation actually looks like, and because what it shows — a median age of sixteen and a half, almost no prior suppression, and adult maintenance doses from the outset in the majority — is not the graduated pubertal-induction protocol the guidelines describe.[1,5]

The dosing figures in URN 2417j do not match the primary

A worked example of adolescent testosterone titration was withdrawn from this section in an earlier version, because the dosing data transcribed from NHS England’s evidence review was pharmacologically impossible. The primary has now been read, and it resolves the question in one direction.

The evidence table in URN 2417j reports this cohort as using “3-weekly intramuscular testosterone undecanoate” (n=63) and “3-monthly intramuscular testosterone enanthate or mixed testosterone esters” (n=78).[47] Moussaoui et al. report undecanoate 3-monthly (n=78) and enanthate or mixed esters 3-weekly (n=63).[5] Undecanoate is the long-acting ester, given at intervals of roughly ten to fourteen weeks; enanthate is short-acting, given at intervals of weeks. The intervals and participant counts are paired correctly in both documents; the two drug names have been exchanged, so each ester carries the other’s interval and the other’s denominator.

Nothing in the review’s conclusions turns on it. It is recorded because this review reproduced the secondary version before checking. The discrepancy was reported to NHS England on 12 July 2026.[45]

The sequential pathway, and how common it actually is. The Dutch programme established a sequential model — a gonadotrophin-releasing hormone agonist from around Tanner stage 2, then the hormone — and several of the influential long-term studies, particularly on bone and adult height, follow it.[1,6,7]

But it is not the dominant route. In the largest prospective US hormone cohort, only 25 of 315 participants (7.9%) had received prior puberty suppression, and 262 (83.2%) were at Tanner stage 5 when hormones began.[13] In this large contemporary US cohort, most participants had completed endogenous puberty and had never received a blocker. That is enough to show that the Dutch sequential pathway is not representative of all current cohorts; it does not establish a continental prevalence, and this review does not claim one. The proportion in current UK practice is not established, and cannot be inferred from the ending of routine NHS blocker provision in 2024; that would require service data this review does not have.

Population flag. The long-run outcome data (bone, adult height, fertility histology) come disproportionately from Dutch sequentially treated cohorts. The mental-health and safety data come disproportionately from North American hormone-only, post-pubertal cohorts. These are different populations receiving different treatments, and findings do not transfer freely between them. Where a claim depends on prior suppression, it is stated.

Three consequences follow, and each recurs later in this review.

First, in sequentially treated cohorts the two exposures are entangled. An uncontrolled before-and-after analysis in a cohort that was suppressed and then given hormones cannot cleanly separate the contribution of each phase. That is not the same as saying they can never be separated — comparison by prior-suppression status, matched cohorts, segmented longitudinal modelling and dose-timing analyses could all do work here. What it means is that the common design cannot, and mostly has not. This is the mirror of ER-014’s central methodological finding about the blocker. Neither phase has been cleanly isolated from the other, and both reviews are obliged to say so.

Second, the pathway is where the blocker’s open questions resolve — or fail to. Bone density falls during suppression; what happens when the hormone is added is a question about the pathway, not about either drug alone, and it is answered in Part V. The same is true of adult height, of body composition, and of fertility.

Third — a fact about the evidence base rather than about the drugs. The sequential pathway was not evaluated as an integrated clinical pathway anywhere in the ten evidence reviews on which the current English policy proposal rests. Every one of them excludes, by design, studies in which GnRH analogues were used for puberty suppression: the monotherapy reviews exclude prior blocker exposure outright, and the combination reviews assess a hormone given with a GnRH analogue used as something other than puberty suppression.[8] The exclusion is methodologically defensible — it is the same objection ER-014 itself made about entangled exposures. Its consequence is that the classical sequential pathway falls outside every PICO. What that licenses, and what it does not, is set out in section 7.

Confidence: 🟢 Robust (that gradual pubertal induction from sub-adult doses is the recommended approach for younger or previously suppressed adolescents; that the ten evidence reviews exclude puberty-suppression exposure by design) / 🔴 Limited (what regimens are actually used, and with what escalation, in older post-pubertal adolescents and in partial or non-binary regimens).

Part III — The evidence base

6. Three appraisals, one conclusion

Three separately conducted appraisal exercises have now formally assessed this evidence, using different methods, different search dates and different inclusion criteria. Two were commissioned for NHS England (NICE in 2020; Solutions for Public Health in 2026); one was commissioned through the Cass Review (Taylor et al., 2024).

Appraisal Studies Quality Verdict on benefit
NICE, 2020[14]
for NHS England
10, all observational.
“No studies directly compared gender-affirming hormones to a control group (either placebo or active comparator).”
Very low certainty on every outcome (modified GRADE). Hormones are “likely to improve symptoms of gender dysphoria, and may also improve depression, anxiety, quality of life, suicidality, and psychosocial functioning.”
Taylor et al., 2024[15]
for the Cass Review
53. 12 cohort, 9 cross-sectional, 32 pre–post. One high-quality study — and it measured side effects only. 33 moderate; 19 low and excluded from synthesis. “Moderate-quality evidence suggests mental health may be improved during treatment, but robust study is still required. For other outcomes, no conclusions can be drawn.” The mental-health signal comes from mainly pre–post studies with 12-month follow-up.
Solutions for Public Health, 2026[8]
for NHS England
Ten separate reviews by PICO subcategory. Six returned no studies at all; four returned between one and eleven. In every review that returned studies, every graded outcome was rated very low certainty. None reached low. “Very low certainty evidence with inconsistent results” — the concluding phrase of the two reviews that returned substantive evidence (2417h and 2417j). The ten reviews issue no single overall verdict.

One thing to hold in mind before reading any count of studies. The adolescent literature is not fifty-three independent investigations. A substantial share of the long-term evidence — bone, adult height, fertility histology, the foundational psychological outcomes — comes from the same Amsterdam gender clinic and its overlapping cohorts, reported repeatedly across different papers with different endpoints and different follow-up windows. The bone papers cited in §18 explicitly acknowledge overlapping participants. Counting them as separate studies inflates the apparent breadth of the evidence base; a systematic review that includes eight Dutch papers has not sampled eight populations. This is a reason to be more cautious about the evidence, not less, and it cuts against both camps equally.

And one number that circulates and should not. Taylor reports that its 53 studies “included 40,906 participants.” Its own next clause reads: “of which 22,192 were adolescents experiencing gender dysphoria/incongruence (8,164 received hormone treatments and 14,028 did not), and 18,714 comparators.”[15] Eight thousand one hundred and sixty-four people in the entire evidence base actually received the treatment. The forty-thousand figure includes untreated adolescents and cisgender comparators. It is not a hormone-exposure denominator, and it should never be quoted as one.

Three appraisal exercises, conducted over six years, and a convergent answer: the evidence is of very low certainty or lacks high-quality studies, and none of the three concluded that the treatment does not work.

The three do not say the same thing about benefit, and the differences matter. NICE found a signal across several outcomes and called it “likely” for gender dysphoria. Taylor found a signal for mental health from mainly pre–post studies at twelve months, and explicitly said that for other outcomes no conclusions can be drawn. Solutions for Public Health, applying much narrower inclusion criteria, found what evidence there was to be inconsistent — improvement on some measures, no change or worse on others, within the same studies.[8]

What can be said across all three, and only this: each rated the certainty very low or the quality inadequate, and none concluded an absence of benefit. Both halves are load-bearing, and each side of the argument reliably drops one.

7. What “very low certainty” does not mean — and what “no evidence” does not mean either

It does not mean no benefit. NICE concluded that hormones are “likely to improve symptoms of gender dysphoria”, and may also improve depression, anxiety, quality of life, suicidality and psychosocial functioning — at very low certainty.[14] Taylor, commissioned by the Cass Review, concluded that there is “moderate-quality evidence” that mental health “may be improved during treatment”.[15]

Those two phrases are not answering the same question, and this review has until now treated them as if they were. NICE’s “very low certainty” is a GRADE-style rating of confidence in an effect estimate. Taylor’s “moderate quality” is a study-level appraisal score: the review used an adapted Newcastle–Ottawa Scale, converted the item scores to percentages, and classified studies as low (≤50%), moderate (>50–75%) or high (>75%), excluding the low-scoring ones from synthesis.[15] A study can score 60% on that checklist and still be a pre–post series with no comparator, incapable of attributing anything to anything. Taylor says so: the psychological conclusion rests “mainly” on pre–post studies.[15]

So “moderate-quality evidence” does not mean “moderate certainty that hormones improve mental health.” It means: the studies that survived a quality filter, most of which cannot establish causation, pointed in that direction. NICE and Taylor are not in conflict, and they are not in agreement. They are measuring different things. Any reading that sets one against the other — in either direction — is comparing a confidence rating with a checklist score. Neither concluded that the treatment does not work. Both concluded that we cannot be confident how much, for whom, or for how long. That is a statement about the studies, not about the drugs — the first of the two symmetric errors named in Part I.

Nor does it mean the treatment is vindicated. Very low certainty is not a technicality to be waved past. NICE’s own reasons are specific and they are not curable by reading the studies more sympathetically: no study had a control group; every observed improvement therefore remains compatible with regression to the mean, natural history, concurrent care, selection effects and residual confounding; comorbidities went unreported; concomitant treatments went unreported “in detail”, so it is not clear whether changes were due to the hormones or to something else the young person was also receiving; regimens were described so poorly that NICE could not tell whether they reflected UK practice at all.[14]

And “no evidence identified” is not the same as either. Six of the ten Solutions for Public Health reviews returned no studies. That phrase has already begun to do enormous work in public argument, and it is being misused in both directions. It is a statement about what a defined search, with a defined PICO, retrieved. In the non-binary categories, the exclusion tables show study after study screened out not because no such patients exist, but because no study reported its results separately for them.[8] That is a fact about how research is written up. It is not a fact about a drug — and the correct remedy for a reporting failure is research, not a conclusion.

8. A necessary caveat about randomisation

This review says repeatedly that no randomised or adequately controlled study of hormones in adolescents exists. That is true, and it is the principal reason all three appraisals rate the certainty very low.

It should not be read as an implicit demand for a trial that may not be possible. Whether a randomised trial of these drugs in adolescents is ethically or practically feasible is a separate question, and a hard one. Blinding is impossible, because the effects are visible. An untreated control arm raises the problem that eligible adolescents rarely wish to refrain from treatment, and that assignment to no treatment may itself drive withdrawal — a point made in the correspondence to the New England Journal of Medicine by the authors of the editorial accompanying Chen et al., who suggested a waiting-list control or a comparison between clinics using different approaches as more realistic alternatives.[21]

So: the absence of randomisation explains the certainty ratings. It is not itself evidence that the treatment fails, and it is not a demand for a design that may be unattainable. What it does mean is that better-designed observational studies — with comparators, with matched controls, with pre-specified outcomes reported in full — are both possible and overdue, and their absence is not explained by the ethics of randomisation.

9. The gap that runs underneath all three

There is one limitation that NICE names explicitly, that NHS England wrote into the specification of every one of its 2026 reviews, that is consistent with Taylor’s call for agreed core outcomes, and that almost nothing in the subsequent public argument has absorbed.

Nobody knows how much improvement counts as improvement.

NICE, in 2020:

“most outcomes reported across the included studies do not have an accepted minimal clinically important difference (MCID), making it difficult to determine whether any statistically significant changes seen are clinically meaningful.”[14]

And NHS England, six years later — not as a finding, but written into the specification of every one of the ten reviews it commissioned, in the outcomes section of each PICO:

“There are no known minimal clinically important differences and there are no preferred timepoints for the outcome measures selected.”[48]

Read those two together. NHS England stated, in the design document, that no threshold of clinical meaning existed for the outcomes it was about to measure — and then commissioned ten reviews to measure them. This is not a criticism of the reviewers, who did what the PICO asked. It is a description of the field.

The consequence is concrete. The largest prospective study in the literature reports depression falling by 1.27 points a year on the Beck Depression Inventory–II — a 63-point scale — over two years (95% CI −1.98 to −0.57), and the change is statistically significant.[13] The study does not establish whether that mean change meets a validated threshold of clinically important improvement for this population, and no such threshold has been agreed for these outcomes in this field.[8,14]

This is a real limitation and it should not be overstated either. A psychometric literature on BDI-II change, reliable-change indices and severity categories does exist; what is absent is an agreed threshold validated in adolescents with gender dysphoria, which is what the outcome measures here are being asked to carry. And a mean change can conceal substantial improvement in some participants and deterioration in others.

So a reader is offered “statistically significant improvement” with no established way to convert it into “better” — and, equally, none to convert it into “negligible.”

Stated once, plainly, because it is the sentence most likely to be dropped: the absence of a minimal clinically important difference is not the absence of clinically important improvement. It is the absence of an agreed instrument for deciding whether improvement occurred. The effect might be substantial. It might be trivial. The field has not built the ruler.

This is not a criticism of any individual study. It is a feature of a field that has not agreed its own outcome measures — which is consistent with Taylor’s recommendation, which is not simply for more research but for agreement on core outcomes first.[15] Taylor does not, so far as this review can establish, name the absence of minimal clinically important differences in those terms. The claim made here is that NICE names it, NHS England specifies it, and Taylor’s recommendation points the same way — not that all three used the same words.

Every psychological grading in Part V should be read with this in mind. Where this review says an effect is statistically significant but of uncertain clinical importance, it is not hedging. It is reporting the state of the field.

Documentary — verified. That NICE (2020), Taylor et al. (2024) and Solutions for Public Health (2026) each rated the evidence very low certainty or lacking in high-quality studies. NICE and Taylor reported possible favourable signals; the 2026 reviews reported inconsistent findings in the categories containing substantive evidence, and issue no overall verdict. These are quotations from published documents.

Causal attribution — not established. None of the three appraisals identified a randomised or otherwise adequately controlled study capable of establishing causality. Every observed improvement remains compatible with regression to the mean, natural history, concurrent care, selection effects and residual confounding.

Observed effect size — 🟡 estimable for most psychological outcomes. Clinical importance — 🔴 not estimable, because no minimal clinically important difference has been agreed for these outcomes in this population.

Part IV — The persistence ledger

10. What “permanent” means, and why the word needs replacing

Part I named the false absolute as one of the errors this review exists to avoid. A table that sorts changes into “reversible” and “irreversible” commits it. An earlier draft of this section contained exactly such a table, and it was wrong on several rows.

Persistence is not a property of an effect. It is an interaction between the tissue, the developmental stage at which the change occurred, how long exposure lasted, what hormonal environment follows, and what corrective interventions exist. This review therefore uses four categories and states which applies:

And one distinction that must not be collapsed. Anatomical persistence is not regret, and it is not harm. A permanent change that is wanted is not a harm; a reversible change that is unwanted still is one. The ledger below describes bodies. It says nothing about how anyone feels about them, and it must not be read as though it did.

11. The ledger, in both directions

Here is the correction that matters most, and the one an earlier draft of this review got badly wrong.

It is tempting to set the permanent effects of treatment against the reversible effects of waiting. That framing is false, and it is false for both pathways. Continued endogenous puberty also establishes characteristics that later hormone therapy does not reliably reverse. Deferral is not a return to a neutral baseline. It is a different exposure, with its own persistent consequences.

So the ledger has four cells, not two.

Pathway Persists after treatment stops Persists from continued endogenous puberty, if treatment is deferred
Masculinising
(testosterone)
Persists — laryngeal enlargement and lowered fundamental frequency (alterable by voice therapy or surgery; not by stopping); clitoral enlargement (any partial reduction unquantified); acne scarring.

Partly persists — facial terminal hair (calibre and rate may fall; established growth generally does not); male-pattern scalp loss once follicular miniaturisation is set.

Reverts variably — menstrual bleeding (usually resumes where the uterus and ovaries remain, but timing varies, and resumption is not equivalent to restored fertility); fat and muscle distribution (shift with whatever hormonal environment follows — not necessarily back).

Not quantified — vulvovaginal atrophic change; body-hair reversion.

Persists — breast development. Testosterone does not remove established glandular tissue; that is why chest surgery exists. Pelvic and skeletal development, and body proportions once the epiphyses have fused.

Possibly a small cost, not quantified as one — adult height. Early suppression followed by testosterone was associated with a modestly greater final height relative to target height than testosterone begun later (§20). How much of that difference is attributable to treatment rather than to pubertal stage at presentation and to selection is uncertain, and this review does not convert the association into a quantified cost of waiting.

Partly persists — fat distribution (partially androgen-responsive).

Feminising
(oestradiol)
Persists — established glandular breast development (volume may fall; removal requires surgery).

Not quantified — recovery of spermatogenesis (suppressed during treatment; may return in some, but the probability, timing and completeness are uncertain, particularly after treatment begun in adolescence); recovery of testicular volume (an imperfect proxy for fertility in any case).

Reverts variably — libido and erectile function; skin; fat and muscle distribution.

Persists — androgen-driven laryngeal enlargement and voice lowering; laryngeal prominence; established facial terminal hair (removal requires electrolysis or laser); major craniofacial skeletal dimensions, once established.

Not a cost of deferral — adult height. The full pathway, begun in early puberty, leaves adult height essentially unaltered and male-typical (§20). Deferral does not add a height consequence, because treatment does not avert one.

Population flag. Almost every entry above derives from adult cohorts, clinical consensus and developmental anatomy. The Endocrine Society timetables describe onset and maximum effect during treatment — they are not discontinuation studies, and citing them as a reversibility ledger is a category error.[1,2] The two partial exceptions are breast development, measured prospectively over three years in adults,[9,10] and clitoral growth, which has a dedicated review.[11] Direct discontinuation evidence in adolescents is, for practical purposes, absent for every row in this table.

11.2 Where each row of the ledger comes from

The ledger is the central claim of this review, and until now it has been the least sourced section in it. It assigns fairly specific persistence categories while stating that direct evidence on discontinuation in adolescents is effectively absent. Both of those can be true at once — but only if the provenance of every row is visible. This table makes it visible.

Read the fourth column first.

Claim What it rests on Population of that evidence Direct data on discontinuation in adolescents?
MASCULINISING — persists after stopping
Laryngeal enlargement and lowered fundamental frequency Mechanism (structural cartilage growth) + adult clinical observation + guideline consensus[1,2] Adults; cisgender pubertal physiology None. No cohort has measured voice after cessation in adolescents
Clitoral enlargement Adult clinical observation; guideline consensus[1,2] Adults None. Any partial reduction on cessation is unquantified in any population
Acne scarring Dermatological principle — established scar tissue does not remodel on hormone withdrawal. Stated as a principle, and cited to nothing. An earlier version cited [32] (acne pathogenesis) and [34] (post-inflammatory hyperpigmentation); neither supports it. A claim carrying two citations, neither of which bears on it, is worse than a claim carrying none. General dermatology None. No adolescent cohort reports sequelae at all (§22)
Facial terminal hair — partly persists Androgen-dependent vellus-to-terminal conversion; calibre and growth rate are androgen-responsive, established follicular conversion is not readily reversed[32,33] Adults; general dermatology None
Male-pattern scalp loss once follicular miniaturisation is set Androgenetic alopecia; DHT-mediated miniaturisation (ER-013)[32,33] Cisgender men; adults None
MASCULINISING — reverts, variably
Menstrual bleeding usually resumes Adult clinical observation; guideline consensus[1,2] Adults with intact uterus and ovaries None. Timing is not characterised, and resumption is not equivalent to restored fertility
Fat and muscle distribution shift with the hormonal environment that follows Adult body-composition data[1,2] Adults None. “Shift” does not mean “return”
Vulvovaginal atrophic change; body-hair reversion Not quantified anywhere. This review makes no claim None
FEMINISING — persists after stopping
Glandular breast tissue does not spontaneously regress Adult clinical observation; guideline consensus. The existence of chest surgery is the clinical evidence[1,2] Adults None. No adolescent cohort has measured breast tissue after cessation
Reduced testicular volume; germ-cell effects de Nie et al. — histology[23] Adult surgical specimens, from people who reached and chose gonadectomy (§19) A four-week pre-operative interruption only. Adherence unverified. Not a test of recovery (§19)
PERSISTS FROM CONTINUED PUBERTY, IF TREATMENT IS DEFERRED
Pelvic and skeletal development; body proportions once the epiphyses have fused Mechanism — epiphyseal fusion is terminal[9,10] Cisgender pubertal physiology Not applicable — this is the counterfactual, not the treatment
Voice deepening from endogenous male puberty is not reversed by oestradiol Mechanism — oestrogen does not shrink laryngeal cartilage; guideline consensus[1,2] Adults Not applicable
Adult height (§20) Three sequential-pathway cohorts[25,26,27] Adolescents — the one row with direct adolescent evidence — (not a discontinuation question)

The fourth column is the finding. Of the thirteen persistence claims this review makes, direct evidence on discontinuation in adolescents exists for none of them. One row rests on adult surgical histology with a four-week interruption that cannot test recovery. The rest rest on mechanism, on adult clinical observation, and on guideline consensus. Every one of them is an extrapolation, and the ledger should be read as one.

This does not make the ledger wrong. Laryngeal cartilage does not shrink; epiphyses do not unfuse; glandular breast tissue does not regress on hormone withdrawal, which is why chest surgery exists. These are not contested propositions, and treating them as unknown would be its own distortion. But they are not measurements, and this review will not present them as measurements. The distinction between “we know this because it was measured in this population” and “we expect this because of how bodies work” is the distinction the whole review is built on, and the ledger must be held to it too.

11.3 What the ledger does not measure: the act of stopping

The ledger above asks what persists and what reverts. It does not ask what the transition costs — and neither, it turns out, does anything else.

Withdrawing sex steroids from a young person whose hypothalamic–pituitary–gonadal axis has been suppressed for years does not return them to a neutral state. It produces hypogonadism: a defined clinical condition with a well-characterised symptom picture — vasomotor symptoms, fatigue, low mood, cognitive complaints, loss of libido, and accelerated bone loss. This is not contested, and it is not speculative. It is the reason hormone-replacement therapy exists, the reason androgen-deprivation therapy in prostate cancer carries the side-effect profile it does, and the reason surgical menopause is managed rather than simply observed.

And the evidence architecture governing this policy knows it. Every one of NHS England’s ten PICO documents states, in its background section, that these medicines “may mitigate the unwanted endocrine and metabolic effects of hypogonadism, which follow gonadectomy or the suppression of sex hormones produced by the body.”[48] The PICO for GnRH-analogue monotherapy — a treatment which is, by design, unreplaced hypogonadism — lists the consequences by name in its safety section: hot flushes, night sweats, headaches, muscle pain, reduced libido, reduced bone density.[48]

And not one of the ten reviews contains an outcome for it. This review has read all ten PICO documents. In every one, the only outcome addressing cessation is detransition, defined as something “important to patients because they may choose to discontinue treatment”, with nine listed synonyms — detransitioner, desistence, discontinuation, cessation, termination, reversion, reversal, disidentification, reidentification — every one of which is a word about identity or decision, and not one of which is a word about physiology. In the safety sections, withdrawal appears only as a consequence of a side-effect (“acute side effects that may lead to withdrawing the treatment”), never as an event with consequences of its own.[48]

So the position is this. The evidence base states that these medicines mitigate hypogonadism. It names what hypogonadism does. It then never asks what happens when they are withdrawn from a young person who is taking them — and the policy under consultation is one whose effect, for some, is exactly that.

What can and cannot be said

Mechanism — 🟢 robust, and not in dispute. Withdrawal of sex steroids from a suppressed axis produces a hypogonadal state. The clinical picture is well described.

Magnitude in this population — 🔴 not estimable. No study has measured it. There is no adolescent cohort reporting vasomotor symptoms, mood, cognition, bone or quality of life following cessation of gender-affirming hormones. Population flag. Everything known about hormone withdrawal in this context is extrapolated from adults — from surgical menopause, from androgen-deprivation therapy in prostate cancer, from GnRH-analogue treatment for other indications, and from cisgender hormone-replacement cohorts. None of those populations is an adolescent who has been on gender-affirming hormones. The extrapolation is mechanistically sound and empirically unmeasured, and this review will not pretend otherwise.

Frequency — unknown, and structurally unknowable from the current evidence. The adolescent cohorts record that some participants discontinued (§24), but none records what happened to them when they did.

This is not an argument that treatment must continue, and it is not offered as one. It is an observation about what can be weighed. A decision to withdraw a treatment is a decision that its costs exceed its benefits. The benefits here are of very low certainty — this review has said so at length. But the costs of withdrawal have not been measured at all, in any study, in this population. An evidence base that cannot describe the cost of stopping cannot weigh it against the cost of continuing. That is not a criticism of the conclusion. It is a description of a gap in what the conclusion was drawn from.

And the one study whose entire subject is stopping does not measure it either. Boskey et al. followed 1,050 transgender adolescents who began gender-affirming hormones between 2007 and 2022 — by some margin the largest discontinuation cohort in the literature. It reports the rate of discontinuation and the reasons for it. It reports nothing physiological: no vasomotor symptoms, no mood, no bone, no cognition, no quality of life off treatment.[43] Read in abstract; full text not obtained.

So the absence is total. Not one study in this literature — including the one designed to examine cessation, in a thousand adolescents — reports what happens to a young person’s body when the hormones stop.

12. What follows — and it turns on pubertal stage, not age

Read the ledger and the first thing that disappears is any clean asymmetry between the two pathways. Both carry persistent treatment effects. Both carry persistent effects of continued endogenous puberty. There is no arm of this decision without lasting consequences — only a choice about which ones, and for whom.

What the right-hand column costs depends almost entirely on pubertal stage, not chronological age — and that cuts hard against a rhetorical move both camps make, this review’s earlier draft included.

Where endogenous puberty is still progressing — an adolescent at Tanner stage 2 or 3 — deferral genuinely does permit further development that later hormones will not reverse. For a transfeminine adolescent, a larynx that keeps growing and facial hair that establishes. For a transmasculine adolescent, breast tissue that develops and pelvic skeletal development that progresses. Deferral there is a live exposure with a running cost, and the claim that waiting is the risk-free option is not sustainable.

Where endogenous puberty is largely complete — and in the largest contemporary cohort 83.2% were at Tanner stage 5 when hormones began[13] — most of the right-hand column has already happened. The voice has broken. The skeleton is set. The breast tissue is there. Further delay adds comparatively little to a bill already paid.

This is where an argument I found rhetorically attractive collapses, and it should be said plainly rather than quietly dropped. “Every month of delay imposes irreversible change” is materially true at Tanner 2. It is much weaker at the ages this policy debate is actually about — sixteen and seventeen, in young people who have overwhelmingly finished puberty. For adolescents who have already completed puberty, many of the persistent consequences of endogenous puberty have already occurred before presentation. That does not make deferral harmless, and it does not make the treatment unnecessary. It means the “delay is itself irreversible” argument does far less work at 16–17 than its proponents suppose — and this review supposed it too, until it checked.

13. What this does not settle

The ledger constrains the argument. It does not resolve it, and three things it cannot supply are worth naming.

It cannot tell you which persistent outcome is worse to risk. A reader who regards unwanted permanent breast tissue in a young person whose goals later change as the graver error, and a reader who regards leaving a young woman with a masculinised larynx she will spend years and considerable money addressing as the graver error, are disagreeing about weights, not facts.

But it does not follow that the disagreement is beyond the reach of evidence. As Part I set out, values supply the weights and evidence supplies the probabilities and magnitudes those weights are applied to. Better data on treatment-goal persistence, on discontinuation, on fertility recovery, on outcomes under deferred treatment, and on who benefits and who does not, could rationally move either reader without either changing anything they value. The gaps identified in this review are therefore not an argument for closing the question. They are an argument for filling them.

And the relevant uncertainty is not “will this young person turn out to be trans?” That framing is too binary, and this review used it. The real uncertainty concerns future embodiment goals and treatment preferences, which may persist, evolve, become non-binary, or settle into satisfaction with some effects and regret about others. Treatment-goal persistence is not identity persistence: a young person may continue to identify as trans and still wish they had made different decisions about their body. The evidence on all of it is graded in Part V, and it is thin.

Confidence — graded on four axes, because a single colour cannot carry documentary fact, causal direction, effect magnitude and population directness at once.

Causal direction — strongly supported. That testosterone produces laryngeal enlargement and voice lowering which do not spontaneously revert; that clitoral enlargement appears largely persistent after cessation, the extent of any spontaneous reduction being unquantified; that oestradiol induces glandular breast tissue which appears anatomically persistent after cessation; that continued endogenous puberty establishes characteristics later hormones do not reliably reverse. The breast and clitoral claims rest on adult evidence and clinical observation; adolescent discontinuation evidence is absent, and the wording here is held to what the ledger in §11 can support. Consistent, mechanistically coherent, universally observed in clinical practice.

Effect magnitude — 🔴 limited, and for several rows not estimable. The degree, timing and variability of every entry in the ledger, in adolescents specifically. Direct discontinuation evidence in this population is effectively absent.

Population directness — indirect. Almost every row is extrapolated from adult cohorts, adult consensus and developmental anatomy.

Not graded. Which persistent outcome is worse to risk. That is not an evidence question, and this review does not answer it.

14. The evidence map

Before any outcome is graded, it is worth seeing the shape of the evidence base in one place. The column that matters is not the certainty rating. It is the second one.

Outcome Adolescent studies Best available design Largest cohort
Physical virilisation / feminisation A handful; voice measured in one, breast volume in two Uncontrolled pre–post, several with unvalidated measures n = 232 (breakthrough bleeding)
Gender dysphoria One. Plus one cross-sectional body-satisfaction study Prospective, with a cisgender comparison group; different UGDS versions at baseline and follow-up n = 23 (16 transmasculine, 7 transfeminine)
Mental health Several; three carry the argument Prospective, 2 years, no comparison group n = 315
Bone Four, with overlapping Dutch participants; plus one untreated cross-sectional Prospective cohort, ~12-year follow-up; uncontrolled n = 889 (untreated, cross-sectional)
Adult height Three Retrospective cohort with non-random comparator n = 161
Spermatogenesis (transfeminine) One Histology at orchiectomy; no untreated comparator n = 214
Fertility (transmasculine) None. No study has assessed it.
Cardiometabolic Several; one with matched controls Cross-sectional, propensity-matched (4,172 vs 16,648); plus retrospective cohorts, 1–3 year follow-up n = 429 (thrombosis)
Skin / acne Three, plus one excluded from the NHS England review Retrospective; the only incidence study was excluded n = 119
Pelvic pain One. Not an outcome in any appraisal Retrospective chart review; no comparator n = 158
Discontinuation Two, incidental to their main purpose 2-year follow-up in a specialist clinic n = 315
Lifetime regret None of adequate duration. NHS England: no evidence identified, in every PICO.
Cognition Four, all small None designed to answer the clinical question Small
Cancer None. The latency is decades; the cohorts are years old.
Fractures None. The outcome bone density is a surrogate for has never been measured.

Read the second column, not the last one. The certainty ratings in the sections that follow are downstream of this table. Three of the outcomes people argue about most — lifetime regret, transmasculine fertility, fractures — have no adolescent studies at all. The central intended outcome of the treatment, relief of gender dysphoria, has one, in twenty-three people. And the counts above overstate the breadth of the evidence rather than understate it, for the reason given in section 6: the same Dutch cohorts recur across papers, and a review that includes eight of them has not sampled eight populations.

Part V — The outcomes, graded

Each outcome is graded on the four axes set out in §3: causal attribution (was the observed change produced by the treatment), observed effect size (how large, and is that knowable), clinical importance (would a patient notice), and population directness (was this measured in adolescents, or borrowed). Where an effect is statistically significant but its clinical importance cannot be established, that is said — because, as section 9 sets out, the field has not agreed what a meaningful change would be.

15. Intended physical effects — the best-evidenced thing here, and still not well evidenced

This is what the drugs are for, and it is the one domain where the appraisals agree without hedging. Taylor found “consistent evidence demonstrating induction of puberty, although with varying feminising/masculinising effects.”[15] Hormone levels rise as expected; secondary sex characteristics develop.[15]

Then the measurement thins out immediately.

Masculinising. One study reports increased facial, abdominal, chest and limb hair, and voice deepening in all participants at follow-up.[15,16] Menstrual suppression is reported in most: 85% cessation at six months in one cohort; 80% with no breakthrough bleeding at twelve months in another; and in a third, suppression in nearly all on 200 mg subcutaneous testosterone but only just over half on 140 mg.[15] Against that, one retrospective series found breakthrough bleeding in 58 of 232 (25%) adolescents on long-term testosterone, beginning at a mean of 24.3 ± 17.2 months.[17] And the same study reports two things that belong here. In 46 of those 58 (79.3%) no cause was identified. And no management method was found to be superior to any other. So the amenorrhoea that adolescents are counselled to expect is not universally maintained; a quarter bleed again, mostly for no reason anyone can find; and nobody knows what to do about it.[17]

Feminising. Two studies observed increased breast volume in birth-registered males — “although objectively breast volume was small.”[15] That is the whole of the adolescent evidence on the single most important intended effect of feminising treatment.

Note what has happened here. The most reliable finding in the field — that these drugs produce the physical changes they are given to produce — rests, for the individual endpoints that matter to a patient, on a handful of small studies using unvalidated instruments. Voice deepening: one study. Breast development in adolescents: two, one of which notes the effect is small. The adult evidence is far better (ER-005, ER-007) and is not a substitute.

Causal direction — strongly supported. Testosterone virilises; oestradiol feminises. Consistent across every appraisal, mechanistically certain, clinically universal.
Effect magnitude — 🔴 limited. Degree, timing and variability of each specific change in adolescents rest on very few studies, several using unvalidated measures.
Population directness — mixed. The fact of virilisation and feminisation is measured in adolescents. The detail is largely extrapolated from adults.

16. Gender dysphoria — a central intended outcome, measured once

Relief of gender dysphoria is one of the central intended outcomes named in the guidelines that recommend these drugs — alongside induction of desired secondary sex characteristics, and improved wellbeing. It is worth stating plainly what the evidence base contains on it.

Taylor’s systematic review identified one study.[15,18] One further cross-sectional study measured body satisfaction, finding lower dissatisfaction among birth-registered females on hormones than among those not treated.[15]

The single study is López de Lara et al., and an earlier draft of this review described it as uncontrolled. That was wrong, and the error came from reading it through NICE’s summary rather than reading the paper. And the summary is where the error lives. NICE’s GRADE footnote for this outcome states that the study “was assessed at high risk of bias (poor quality overall; lack of blinding and no control group)”. The paper reports thirty cisgender controls, and its lead finding is a comparison against them. Taylor, appraising the same paper, lists it among the six included studies that used matched controls.[15] Two appraisals of one study, disagreeing on whether it has a comparator. If what NICE meant is that there is no untreated trans comparator, that is correct, and it is the more important point — but it is not what the footnote says, and this review reproduced the footnote rather than the paper.

It is a prospective analytic study of 23 transgender adolescents aged 14–18 (16 transmasculine, 7 transfeminine) and 30 cisgender controls, assessed before treatment and after one year of hormone therapy.[18] Mean Utrecht Gender Dysphoria Scale scores changed from 57.1 (SD 4.1) at baseline to 14.7 (SD 3.2) at one year — a scale running from 12 to 60, with a stated cut-off of 40. That cut-off was devised to identify probable gender dysphoria, not to mark a clinically meaningful treatment response. It is a diagnostic discriminator, not a minimal clinically important difference, and it cannot be used as one — least of all across two different forms of the scale. Anxiety, depression (BDI-II), emotional symptoms, conduct problems and prosocial behaviour all improved significantly.

And this is the study on which the entire chain rests. Taylor’s systematic review contains exactly one study of gender dysphoria: “One cohort study measured gender dysphoria pre–post, and reported a reduction in dysphoria with no participants in clinical range at follow-up.”[15] That study is this one. NICE’s single positive conclusion — that hormones are “likely to improve symptoms of gender dysphoria” — rests on it.[14] NHS England inherits the characterisation.[8] Neither Taylor nor NICE reports what follows.

And here is the fact that governs everything that can be said about that number. The UGDS exists in two forms, male and female, which are selected by sex assigned at birth and contain differently framed items. The paper’s Methods state that the scale was “filled out for the sex assigned at birth at T0 and the self-identified gender at T1.”[18]

The respondents did not complete the same instrument twice. They completed one form before treatment and the other form afterwards. The numerical difference between 57.1 and 14.7 therefore cannot be read as a 42-point change on a measurement-invariant longitudinal scale: the items, their referents and their psychological framing all changed with the version. Longitudinal measurement equivalence was not demonstrated, and on this design it could not be. This does not make the result meaningless — a large pre–post contrast on a dysphoria instrument, in a cohort that also converged with cisgender controls, is a real observation and should not be waved away. But it is not a 42-point treatment-response estimate, and this review has until now presented it as one. That was wrong, and it is corrected here.

And the finding the authors lead with: the trans group differed significantly from the cisgender controls at baseline, and the scores equalised at one year.[18]

On this basis NICE concluded hormones are “likely to improve symptoms of gender dysphoria”, at very low certainty.[14]

That is the entire adolescent evidence base on this outcome: one prospective study of twenty-three young people.

Three things must be said about it, and each matters.

The observed change is numerically large. Roughly 42 points on a 48-point range, from near the ceiling to near the floor. Anyone describing this as “no evidence” is misdescribing it. But — see above — the two scores are not on the same form of the instrument, so the magnitude cannot be taken at face value as a change score.

It has a comparison group — but not the one that answers the question. The controls are cisgender adolescents. That design can show, and does show, that trans adolescents began markedly worse than their cisgender peers and were statistically indistinguishable from them a year later. It cannot show what would have happened to trans adolescents who were not treated, because no such group was followed. Convergence with a healthy comparison group is consistent with the treatment working; it is also consistent with regression to the mean, with the effect of a year in a supportive specialist service, with social transition, and with expectancy and other non-specific treatment effects — which are neither trivial nor avoidable in an unblinded intervention whose effects the participant can see in the mirror. The design cannot separate them.

And it cannot bear the weight placed on it. Twenty-three participants — sixteen transmasculine, seven transfeminine. Twelve months. One centre in Madrid. The reviewers who identified it also concluded that no conclusions could be drawn about gender-related outcomes.[15] A near-ceiling-to-near-floor shift on a self-report scale in a cohort that also underwent social transition and a year of specialist care is a striking observation, and it is not a causal effect estimate.

Causal attribution — 🔴 not established. The comparison group is cisgender, not untreated trans adolescents. Convergence with cisgender peers cannot distinguish treatment effect from regression to the mean, social transition, expectancy and other non-specific effects, or a year of specialist care.
Observed effect size — 🟡 numerically large, interpretability limited. A difference of roughly 42 points is reported, but different sex-specific versions of the UGDS were administered at baseline and follow-up and longitudinal measurement equivalence was not demonstrated. The number is not a change score on a fixed instrument.
Clinical importance — 🟡 uncertain. Mean scores moved across the study’s stated cut-off, but that cut-off was designed to identify probable dysphoria, not to mark a clinically meaningful treatment response — and it cannot bear that use across two different forms of the scale.
Population directness — direct, and very narrow. Twenty-three adolescents at one centre, of whom seven were transfeminine. The problem is not directness. It is quantity, and it is the fact that the whole field rests on it.

17. Mental health — the contested ground

This is where the argument actually happens, and where this review is obliged to be most careful. Four sources carry it, and all four have been read in full.

17.1 What the studies found

Chen et al. 2023 — 315 adolescents and young adults, four US sites, prospective, two years, no comparison group.[13] Over two years, the primary latent-growth models showed significant favourable slopes for all five outcomes: appearance congruence +0.48 per year (95% CI 0.42 to 0.54), depression −1.27 (−1.98 to −0.57), anxiety −1.47 (−2.15 to −0.80), positive affect +1.12 (0.37 to 1.89), life satisfaction +2.31 (1.65 to 2.99).[13] These are the values in the corrected Table 3, read in primary. The article as first published gave −1.46, +0.80 and +2.32 for anxiety, positive affect and life satisfaction; the correction of 18 October 2023 states that the numerical values throughout Table 3 were inaccurate and replaces them.[38] The effect sizes, and one disagreement between analyses, are the story.

Outcome, baseline to 24 months p Cohen’s d
Appearance congruence <0.001 −1.12 (large)
Life satisfaction <0.001 −0.39
Anxiety <0.001 0.25
Depression <0.001 0.20 (small)
Positive affect 0.39 −0.06

Provenance. These values are read from the replaced Supplementary Appendix at NEJM.org, whose Tables S5 and S6 are marked “(Corrected)”. The superseded author-manuscript appendix deposited in PubMed Central (NIHMS1877102) carries different values and is not used.[38]

From Table S5 — a paired baseline-to-24-month comparison of the same participants. The complete-case comparison shows no significant change in positive affect; the primary latent-growth model, which uses all available data, does. Both are reported because both are true, and the disagreement is informative: the complete-case analysis uses roughly 70% of the cohort and is vulnerable to attrition bias, while the growth model is the better-specified analysis. The 24 participants excluded from the longitudinal analysis were, if anything, more depressed at baseline than those retained (Cohen’s d 0.48, p=0.06), which is the direction that would flatter the complete-case result.[13] Cohen’s d signs follow each scale’s direction — a negative d for appearance congruence and life satisfaction indicates improvement, as does a positive d for depression and anxiety.[13]

Depression moves 2.16 points on a 63-point scale over two years — from a mean of 16.01 (SD 11.88) to 13.85 (SD 12.71) in the 211 participants with both measurements.[13] Two different quantities are in circulation here and they must not be conflated. The latent-growth model estimates a slope of about 1.27 points per year, which is roughly 2.5 points over two years; the paired complete-case comparison gives 2.16. They are close, they are not the same number, and they come from different analyses of different subsets. This review reports both and labels which is which. The proportion classified as severely depressed was 15.6% at baseline (48 of 307) and 13.6% at 24 months (30 of 220) among the relevant observations — note that these are not necessarily the same participants: roughly one-third of the cohort has no 24-month depression score at all (220 of 315 present, 70%).[13] Roughly one in seven of those assessed at two years remained in the severe range — in a cohort that excluded anyone presenting in acute suicidal distress at entry.

The companion paper, Olson-Kennedy et al. 2025, follows the same cohort on the NIH Toolbox and reports significant improvement across psychological well-being, self-efficacy, social satisfaction, negative social perception and negative affect.[19] Its T-scores are normed to a US population mean of 50. After two years of treatment, the cohort sits at: positive affect 43.7, life satisfaction 44.8, friendship 41.0, loneliness 58.8, fear 56.3, sadness 55.3. Below the population norm on ten of eleven measures. The paper reports, accurately, that every T-score improved. It does not say where the cohort ends up.

Tordoff et al. 2022 — 104 adolescents, Seattle, twelve months. Widely cited, and frequently misdescribed.[20] Intervention directness — mixed, and this is routinely lost. Tordoff’s exposure is initiation of puberty blockers, gender-affirming hormones, or both. It is a study of pharmacological gender-affirming care, not of hormones alone, and it cannot estimate the isolated association of testosterone or oestradiol with any outcome. Where this review refers to “medication” in describing it, that is what is meant.[20] Its headline — 60% lower odds of moderate-to-severe depression, 73% lower odds of self-harm or suicidal thoughts — is not a before-and-after improvement. The paper’s own words: “there were no statistically significant temporal trends”, and “we observed a transient and nonsignificant worsening in mental health outcomes in the first several months of care among all participants and that these outcomes subsequently returned to baseline by 12 months.”[20]

Over twelve months, the study found no statistically significant overall temporal trend in depression, anxiety or suicidality. The 60% and 73% are a comparison, at each time point, between those who had started medication and those who had not. The reported association appears to reflect worse short-term outcomes among participants not yet receiving medication (depression aOR 3.22 at three months; self-harm or suicidal thoughts aOR 2.76 at six months) rather than clear within-group improvement after initiation — although the exposure is time-varying and non-random, and the observational design prevents a simple decomposition.[20] Unadjusted, the depression association is null (aOR 0.67, 95% CI 0.33–1.34). Anxiety shows nothing in any model (aOR 1.01, 0.41–2.51).

And that reading is not a debunking. If, in a gender clinic, participants not yet receiving medication show markedly higher odds of depression at three months, that is relevant to any policy that delays or withholds it. Preventing deterioration would be a benefit. But the exposure here is time-varying and non-random, and the comparison is not between a treated cohort and an untreated one over time; this review offers the reading as an interpretation, not as a finding of the study. It is, in any case, not the benefit the study is usually quoted for.

Two cautions. “Not yet receiving medication” is not the same as “withheld from” — delay may reflect eligibility, readiness, preference, scheduling or contraindication, any of which could also predict outcome. And confounding by indication is not excluded: the paper reports no baseline comparison between those who did and did not go on to receive medication. The authors’ E-values (2.56 and 3.25) describe the minimum association an unmeasured confounder would need with both exposure and outcome, conditional on measured covariates, to move the estimate to the null. They do not classify real-world confounding as weak or strong, and they do not address bias from time-varying exposure, selection, measurement error or model specification.

17.2 The finding that must be stated first

From Chen et al., verbatim:

“Depression and anxiety symptoms decreased significantly, and life satisfaction increased significantly, among youth designated female at birth but not among those designated male at birth.”[13]

That sentence is the paper’s own. Note that it names anxiety — and that, on the corrected figures, the anxiety interaction no longer reaches significance. This review states the finding as depression and life satisfaction only, for the reason given below.

In the largest prospective cohort included in the appraisals, improvement in depression and life satisfaction differed significantly by sex designated at birth: favourable slopes were concentrated among participants designated female at birth, while the corresponding slopes among those designated male at birth were not statistically significant.

This is a moderation finding, and it must not be over-read in either direction. A non-significant within-group slope does not establish that there was no improvement. And sex designated at birth is not a randomised comparison of oestradiol versus testosterone: it is a proxy, and the groups may also differ in prior suppression, dose, age, baseline score and much else. This does not establish that oestradiol has no psychological benefit. It establishes that the benefit observed in this study was not evenly distributed, and that the group in which it was not observed is the transfeminine one.

This is not a subgroup fished from the data. It is a pre-specified moderator, tested as a time-invariant effect on the growth slope, and it is significant in the paper’s own analysis.[13]

In the corrected analysis, the designated-sex-at-birth effect on the growth slope remains significant for depression (1.86, 95% CI 0.27 to 3.43) and for life satisfaction (−1.92, −3.44 to −0.41).[13] Read in primary from the corrected Table 3. Designated sex at birth is coded 0 = assigned female at birth (reference), 1 = assigned male at birth, so a positive coefficient on the depression slope means the favourable slope is attenuated in participants designated male at birth.[13]

Anxiety is the outcome this review declines to claim. The paper’s own sentence, quoted above, names anxiety. But in the corrected Table 3 the anxiety interaction is 1.53 (95% CI −0.04 to 3.11): the interval crosses zero, and the finding does not stand.[13,38] This review therefore claims depression and life satisfaction, and does not claim anxiety. A version of this section written from the paper’s narrative, or from the figures as first published, would have claimed all three — and would have been wrong.

The authors offer two explanations, both explicitly speculative: that oestrogen-mediated changes take longer to manifest, and that transfeminine young people may face greater minority stress.[13] Both are plausible. Neither is evidence. And the companion paper’s composite measure does not show a designated-sex-at-birth difference in slope[19] — a tension between two papers from the same cohort that is unresolved and is stated here rather than smoothed over.

17.3 What none of it can do

None of the three appraisal exercises identified a randomised trial, or any controlled study capable of establishing causality, and no such study was identified in the additional literature reviewed here. NICE found none; the NHS England reviews, searching to June 2025, found none; Taylor found none.[8,14,15]

That is a statement about the evidence that exists. It is not a statement about the evidence that is being built — and a review published in 2026 must say so. The PATHWAYS Trial, sponsored by King’s College London and South London and Maudsley NHS Foundation Trust and funded by the Department of Health and Social Care and NHS England, is a randomised trial of GnRH-analogue treatment in 226 participants under sixteen. Participants are randomly assigned to begin treatment immediately or after a one-year delay, and the two groups are compared over two years across quality of life, mental health, physical development, cognitive function and gender-related distress. A parallel observational cohort of 300 runs alongside it.[44]

The delayed-start design is precisely the comparator this literature lacks. It constructs an untreated control group by randomising timing rather than access — which is the only ethically available route to the counterfactual whose absence every certainty rating in this review turns on. Two things must be said about it, and this review will say both. First: PATHWAYS tests the blocker, not the hormones. It will not answer the question this review is about, and nothing in it will resolve the certainty ratings graded here. Second: it has not reported, and this review does not know what it will find. Its existence changes the evidential future. It changes nothing about the evidential present.

Chen’s authors say it plainly: “our study lacked a comparison group, which limits our ability to establish causality.”[13] Tordoff’s say the same.[20]

And the study’s own data show selection at work, in the one comparison it makes. The 24 participants who had received early gender-affirming care — those started in Tanner stage 2 or 3, the closest thing the cohort has to a sequentially treated group — were already markedly better off at baseline than the other 291: depression 9.57 versus 17.00 (p<0.001, Cohen’s d 0.71), anxiety 51.54 versus 60.75 (d 0.79), appearance congruence 3.08 versus 2.31 (d 0.86).[13] In the growth models, early treatment is associated with much better baseline scores (depression −5.91, 95% CI −9.76 to −2.06) but with no difference in the rate of change (−0.68, −3.36 to 1.83).[13] Whatever is going on, it is not that early treatment made these young people improve faster. They arrived less depressed. Who gets treated early, and why, is not random, and this is the clearest measurement of that fact in the literature.

Participants in these cohorts were embedded in specialist gender-affirming care, often including mental-health support. This is not to say every participant was on hormones at every time point — Tordoff’s analysis expressly compares those who had initiated medication with those who had not yet done so. The care package was shared; hormone exposure was not.[20] In the uncontrolled hormone cohorts, the hormone cannot be separated from the care, from the relief of finally being treated, from regression to the mean, from selection effects, or from the natural course of adolescence. This is the exact mirror of ER-014’s finding about the blocker, and it is not a rhetorical point: it is the reason all three appraisals rate the certainty very low.

17.4 And nobody knows whether any of it is enough to matter

Return to section 9. Depression falls a little over two points on a 63-point scale over two years — 1.27 per year on the growth model, 2.16 on the paired comparison. Is that better?

The observed magnitude is known. What is not established is its clinical importance: no study here demonstrates that the mean change meets a validated threshold of clinically important improvement for this population, and NHS England’s own PICO states that “there are no known minimal clinically important differences” for these outcomes,[48] as NICE did six years earlier.[14]

This is not a hedge. It is one of the most consequential methodological limitations in the mental-health evidence, and it belongs to neither side. It does not show the treatment fails: the effect might be substantial. It does not show the treatment works: the effect might be trivial. It shows that the argument is being conducted over an instrument that was never calibrated.

17.5 Suicide

Two of the 315 participants in Chen’s cohort died by suicide during the study; eleven reported suicidal ideation at a study visit.[13] These are reported in the paper’s adverse-events table and in its abstract.

They must be reported here. A review of this treatment that omits two deaths in a two-year study of 315 young people is not an evidence review.

They do not establish harm. A published letter converts them into an annual rate of 317 per 100,000 — with a 95% confidence interval of 38 to 1,142.[21] Two events permit only an extremely imprecise crude rate, spanning a factor of thirty, and cannot support a reliable comparison or a causal inference. The authors’ response — that US and UK baseline rates differ, that their cohort included young adults, that the comparator clinic served a different population — is well made.[21]

They do not establish safety either. “Not statistically distinguishable” is not “reassuring” when n = 2.

One further fact, and it is documentary. The study protocol published by the journal alongside the article pre-specifies suicidality and self-injury as outcomes for the hormone cohort, in terms: “Patients treated with cross-sex hormones will exhibit decreased symptoms of anxiety and depression, gender dysphoria, self-injury, trauma symptoms, and suicidality…” (Hypothesis 2a). It specifies the instrument — an eight-item Suicidal Ideation Scale, “Eight yes/no questions … to capture participants’ suicide ideation and attempts” — and it schedules the survey battery containing it for administration “at baseline and 6, 12, 18, and 24 months.”[39] A protocol schedules an instrument; it does not by itself prove that it was administered at every point, or that usable data exist. Neither outcome is reported in Chen et al., nor in the companion paper. The same pre-specification appears in the study’s published protocol paper.[22] The suicidality data were requested by name in correspondence to the journal; the authors’ reply did not supply them.[21]

This review draws no inference from that. Non-reporting is not evidence of an adverse result — that is the absence-of-evidence error named in Part I, and committing it here would be indefensible. We do not know what those instruments showed. We know only that they were pre-specified, scheduled for collection at five time points over two years, asked for, and not published. Whether they were in fact administered at every point, and whether usable data exist, is not established by a protocol. The correct response to that is to ask for the data, not to assume the worst.

Causal attribution — 🔴 not established. No randomised or otherwise adequately controlled study exists. Every observed improvement remains compatible with regression to the mean, natural history, concurrent care, selection effects and residual confounding. Three formal appraisals rated the certainty very low, and none concluded an absence of benefit.[8,14,15]

Observed effect size — 🟡 small. Complete-case Cohen’s d 0.20 to 0.39 across depression, anxiety and life satisfaction; positive affect null on the paired comparison (d = −0.06, p = 0.39). No confidence intervals are reported for these estimates, and they are sensitive to attrition.

Clinical importance — 🔴 not estimable. Distinct from the above. The observed complete-case effect-size estimates are small; their causal and clinical interpretation remains uncertain. These are point estimates from a paired complete-case subset, sensitive to attrition, and the appendix reports no confidence intervals for them. They are not immutable quantities.

Moderation by sex designated at birth — 🟡 significant. The interaction is significant for depression and life satisfaction; the corresponding slopes among participants designated male at birth were not statistically demonstrated. A non-significant subgroup slope is not evidence of absence, and this does not establish that oestradiol has no psychological benefit — and it concerns the population most directly affected by the feminising half of the policy.

Suicide — not estimable in either direction. Two events, no comparator, extreme baseline risk, and a pre-specified instrument that was never reported.

18. Bone-density recovery after the sequential pathway

ER-014 left this unfinished. Bone density decelerates during puberty suppression; whether it recovers once hormones begin was, at the time of that review, unresolved. Four reports from the same clinical programme — with overlapping populations, a shared treatment protocol and shared measurement conventions — now characterise it. They are not four independent replications. The long-term follow-up, the dose–response study, the prospective development study and the untreated cross-sectional study all come from the Amsterdam programme; some explicitly contain overlapping participants. Counting them as four evidential votes overstates the breadth of the evidence, and an earlier version of this section did. The answer they give is more interesting than the summary tables carry: mean z-scores returned approximately to pre-suppression levels at the measured sites in the testosterone-treated group; a lumbar-spine deficit remained in the oestradiol-treated group. That deficit is plausibly explained in part by lower oestradiol exposure under the historical protocol, and by a cross-sectional age gradient in untreated participants that raises the possibility that z-scores may not track in this population at all.

18.1 What actually happens

A note on what “recovery” can and cannot mean here. Everything below is measured as an areal or height-adjusted bone-density z-score at three skeletal sites. A return of the mean z-score to its pre-treatment value at those sites is not the same as restored peak bone mass, normal microarchitecture, normal bone strength, normal fracture risk, or recovery in every individual. This review uses “returned to pre-treatment levels at the measured sites” and means no more than that.

And the primary itself does not observe that limit. The long-term follow-up study concludes that the pathway is “safe regarding bone health” in testosterone-treated participants. It measured DXA-derived density and z-scores. It did not measure fracture incidence with adequate power, bone microarchitecture, or mechanical strength — and the same paper notes that paediatric osteoporosis requires a low z-score together with a clinically meaningful fracture history. The data support recovery of the measured density outcomes. They do not support a general conclusion of bone safety, and this review does not adopt one.[28]

During suppression, z-scores fall. In the largest prospective adolescent series, areal and apparent bone density were within the normal range at the start of GnRH-analogue treatment, but z-scores declined in every group over two years.[30] That is not in dispute and this review does not dispute it.

On hormones, accrual resumes — and in trans boys the mean z-scores return to where they started. Over three years of gender-affirming hormones, bone mineral apparent density rose significantly in all groups; z-scores normalised in trans boys and remained below zero in trans girls.[30] The long-term follow-up is more precise. In 75 adults who had used suppression before eighteen and then hormones for a median of about twelve years, z-scores at every site had caught up with pre-treatment levels except the lumbar spine in those assigned male at birth. In participants assigned female at birth, the changes from the start of the blocker were 0.09 at the lumbar spine (95% CI −0.09 to 0.27), 0.10 at the total hip and −0.20 at the femoral neck — none significant. In those assigned male at birth, the total hip (−0.12; 95% CI −0.31 to 0.07) and femoral neck (0.01; −0.20 to 0.22) had recovered, but the lumbar spine had not: a mean z-score of −1.34, a change of −0.87 from the start of the blocker (95% CI −1.15 to −0.59).[28]

So the group-specific finding is real, it is measured, and it is the one this review will be quoted on. What follows is the part that is not being quoted.

18.2 Lower oestradiol exposure may contribute to incomplete lumbar-spine recovery

The Amsterdam protocol raised oestradiol gradually to 2 mg, reached on average two years into treatment. Median serum oestradiol on that dose was 126 pmol/L — below the centre’s own minimum target of 200 pmol/L, and far below the Endocrine Society range.[29]

When the dose was higher, the bone came back. In 87 trans girls, lumbar-spine height-adjusted z-scores fell by 0.69 during suppression, then over two years of hormones rose by only 0.14 on the regular 2 mg dose (95% CI −0.01 to 0.28), by 0.42 on 6 mg (0.13 to 0.72), and by 0.68 on ethinylestradiol (0.20 to 1.15). After two years, scores had returned to their pre-suppression baseline in the 6 mg and ethinylestradiol groups — but not in the regular-dose group, which remained 0.64 below baseline (95% CI −0.79 to −0.49).[29] The authors conclude that 2 mg is insufficient and that roughly 4 mg may be needed.

What this does and does not license. Boogers demonstrates a dose–response relationship. It does not demonstrate that low dosing caused the long-term deficit. The high-dose groups were small (n=13 and n=15), non-randomly allocated, younger at the start of suppression, and given those doses to reduce growth rather than to protect bone; dose, selection, age, growth velocity and unmeasured factors are not separated. Nor does it license ethinylestradiol, whose thrombotic profile is the least favourable in the feminising formulary and which the authors explicitly decline to recommend for this purpose. What can be said is narrower and still consequential: the incomplete lumbar-spine recovery in trans girls is consistent with an important contribution from an oestradiol dose the treating centre now regards as too low. A finding routinely presented as an intrinsic property of the pathway may be, in part, a property of how the pathway was dosed. That is a synthesis, not a demonstration, and it is offered as one.

18.3 And the comparator may not be stable

Every z-score argument in this field rests on an unstated assumption: that without treatment, a young person’s bone-density z-score would hold steady — the phenomenon paediatricians call tracking. If z-scores fall during suppression, the fall is attributed to the suppression.

In trans girls, there is reason to doubt that assumption — though the study that raises the doubt cannot settle it either. The paper is titled “the natural course of bone mineral density”, and that framing outruns its design: a cross-sectional comparison of different people at different ages cannot demonstrate a within-person course. This review adopts the finding, not the framing. In a cross-sectional study of 333 adolescents and young adults assigned male at birth, scanned before any medical treatment at all, bone-density z-scores were negatively associated with age between twelve and twenty-two: −0.13 per year at the lumbar spine (95% CI −0.17 to −0.09), −0.04 at the total hip (−0.08 to −0.01), −0.06 at the femoral neck (−0.09 to −0.03), and −0.12 at total body less head (−0.15 to −0.09). These are the figures in the paper’s Results. Its abstract prints three of them differently — the total hip as −0.05 (−0.08 to −0.02), the lumbar-spine interval as −0.17 to −0.10, and the femoral-neck interval as −0.10 to −0.03. The Results govern. An earlier version of this review quoted the abstract, having read the paper in full — which is a different failure from not reading it, and a more insidious one. Nothing in the finding or the conclusion turns on the difference. In 556 assigned female at birth, no such association existed at the lumbar spine, hip or femoral neck.[31]

Adjusting for height-adjusted lean mass eliminated the association at the hip and femoral neck and attenuated it at the spine — pointing at reduced physical activity and muscle mass rather than at anything hormonal.[31] This is consistent with what the earlier cohorts kept observing and could not explain: trans girls’ z-scores were already well below the population mean before a single injection.[28,30]

What this is, and what it is not. It is a cross-sectional age gradient: older untreated participants had lower z-scores than younger untreated participants. It is not a measured within-person trajectory, and it does not show that any individual’s z-score was falling. Age, cohort, referral patterns, body composition and secular change are all entangled in a design of this kind, and this review will not convert the association into a trajectory.

What it does do is raise a serious doubt about the comparator. Every uncontrolled before-and-after z-score inference in this literature assumes that, without treatment, z-scores would have tracked. That assumption has never been tested, and the one study that bears on it points against it in the group where the alleged harm is found. A cross-sectional age gradient in untreated participants raises the possibility that z-scores may not track in this population — and if they do not, then some unknown share of the observed decline is not attributable to treatment. How large a share, nobody can say. Neither can anyone say it is zero, which is what the standard inference assumes.

18.4 What is still missing

Fractures. Density is a surrogate. In the prospective adolescent cohort, no participant sustained a fracture during the study — but that cohort was neither large enough nor followed long enough to estimate fracture risk, and no study in this population is.[30] A z-score is not a broken bone, and this review will not pretend otherwise in either direction.

Attribution. Every one of these studies is uncontrolled and assesses a combined sequential pathway. Resumed sex-steroid exposure, continued maturation, weight gain, physical activity, vitamin D and dose are not separable. Duration of GnRH-analogue monotherapy was not associated with long-term z-scores in the follow-up cohort[28] — which is evidence against a simple dose-of-blocker mechanism, and is not the same as evidence of no harm.

And a boundary. These are Dutch cohorts who had puberty suppression from early puberty. They say nothing about a seventeen-year-old starting oestradiol or testosterone without a blocker — which, as §5 sets out, is now the common route. From the adult evidence, aromatised oestradiol matters for bone even under masculinising therapy, so driving oestradiol to zero is a bone risk rather than a treatment goal (ER-007).

Causal attribution — 🔴 limited, and weaker than the field assumes. The physiological direction is expected, but the studies are uncontrolled and the exposures are entangled. Two further problems bear on the oestradiol finding specifically: the historical protocol delivered lower oestradiol exposure than the treating centre now considers adequate, and a cross-sectional age gradient among untreated participants raises the possibility that the z-score comparator does not track in this group. Neither has been shown to account for the deficit; both make the standard causal reading less secure than it is usually presented.[29,31]
Observed effect size — 🟡 estimable, and group-specific. In the testosterone-treated group, mean z-scores returned to pre-suppression levels at all three measured sites. In the oestradiol-treated group they returned at the hip and femoral neck but not the lumbar spine, which remained −0.87 below the pre-suppression value at a median of about twelve years of hormones (95% CI −1.15 to −0.59). On 6 mg oestradiol or ethinylestradiol, mean lumbar-spine scores returned to the pre-suppression level within two years. None of this speaks to peak bone mass, microarchitecture, strength or fracture.
Clinical importance — 🔴 not estimable. No fracture data exist in this population. Density is a surrogate and is being asked to carry a conclusion it cannot.
Population directness — direct. Sequentially treated Dutch cohorts. Not transferable to hormones begun at sixteen or seventeen without prior suppression.
What this does not say. It does not say puberty suppression is safe for bone. It says the evidence that it is unsafe is weaker than it has been made to look, that part of the signal is plausibly attributable to lower oestradiol exposure under the historical protocol and to an untreated age gradient that had not previously been characterised, and that the outcome anyone actually cares about — fractures — has never been studied.

19. Fertility

Taylor: “no study assessed fertility in birth-registered females.”[15] That is the whole of the evidence on fertility in transmasculine adolescents. Nothing.

For transfeminine adolescents, there is exactly one study.

De Nie et al. analysed orchiectomy specimens from 214 transgender women who underwent genital gender-affirming surgery between 2006 and 2018, grouped by pubertal stage at the start of medical treatment (Tanner 2–3, Tanner 4–5, or adult) and by whether hormones were stopped before surgery.[23] All had used oestrogens plus testosterone-suppressing therapy.

The cohort is selected several times over before any tissue is examined, and this bounds everything below. The source population was 788 people undergoing orchiectomy with vaginoplasty at a single centre. Participants with certain medical histories were excluded, as were those on oestradiol monotherapy and those on spironolactone; groups were capped at 80 and randomly sampled where larger. What remains is adult surgical histology from people who reached, and chose, gonadal surgery. It excludes everyone who never pursued surgery, who discontinued, who retained their gonads, who used other anti-androgens, or who had contraindications. It is not a sample of adolescents who begin feminising treatment; it is a sample of those who later had an orchiectomy at one Amsterdam clinic.[23]

Most advanced germ cell present Proportion Who
Mature spermatozoa 4.7% All had begun treatment at Tanner 4 or later
Immature germ cells only 88.3% The great majority, across all groups
Complete absence of germ cells 7.0% All had begun treatment in adulthood

An earlier draft of this review, working from a systematic review’s one-line summary, read this as showing that early suppression forecloses fertility. Reading the paper inverts that.

Three findings that summary loses:

Mature spermatogenesis was rare in everyone, at every starting age, while on treatment — 4.7% overall. This is a histological snapshot taken at orchiectomy, in people who were at that moment on oestrogen and testosterone suppression. It says nothing about what spermatogenesis would have been had treatment never begun, and it must not be quoted as “timing does not matter.” It was absent in most Tanner 4–5 and adult starters too, because all of them were on oestrogen and testosterone suppression, which suppresses spermatogenesis regardless of when it began. The Tanner 4+ condition is necessary but nowhere near sufficient. It is not the case that later starters kept their fertility.

Every specimen with no germ cells at all came from someone who began treatment in adulthood. Adolescents — including those suppressed at Tanner 2–3 — retained germ cells, and mean Johnsen scores were lowest in the adult cohort.

19.1 And that difference is not about age. It is very probably about a different drug.

The adolescent and adult groups did not receive the same testosterone-suppressing agent. They received entirely different ones, with no overlap whatever.

Testosterone suppression Adolescent starters Adult starters
Triptorelin (GnRH agonist) 100% 0%
Cyproterone acetate 0% 100%

Age at initiation and anti-androgen class are perfectly confounded in this study. The histological difference cannot be attributed to when treatment began, because everyone who began in adolescence received a GnRH agonist and everyone who began in adulthood received cyproterone acetate. Treatment duration, age at surgery, cohort period, lifestyle and oestradiol dose also differ between the groups.[23]

And the authors say so themselves — which is the part that matters most. They write that the difference “might be explained by age, lifestyle …, higher dosages of estradiol or the use of cyproterone acetate instead of GnRHa as testosterone suppressing therapy,” and that whereas a GnRH agonist inhibits gonadotrophin secretion, cyproterone acetate also acts as a direct androgen-receptor antagonist and “might have more profound and irreversible effects on testicular tissue.”[23]

They then draw the conclusion that this review had missed entirely. Because of its side-effect profile, their clinic had already stopped giving cyproterone acetate to adults starting treatment after eighteen. The authors add: “The potential consequence of irreversible infertility might be an extra reason to not prescribe cyproterone acetate anymore.” And they call for a study comparing adults and adolescents when both receive a GnRH agonist.[23]

So the reading changes, and it changes in a direction this review did not expect. The finding is not “starting early protects fertility.” It is: in the only histological cohort available, the participants with no germ cells at all were the ones who received cyproterone acetate.

What the authors say, precisely. They offer four candidate explanations for the adult–adolescent difference — “age, lifestyle (a higher percentage of smokers and alcohol drinkers), higher dosages of estradiol or the use of cyproterone acetate instead of GnRHa” — and age is the first of them. They then elaborate a mechanism for one alone: cyproterone acetate acts as a direct androgen-receptor antagonist as well as suppressing gonadotrophins, and so “might have more profound and irreversible effects on testicular tissue.” And they draw a practice conclusion from that one alone: “the potential consequence of irreversible infertility might be an extra reason to not prescribe cyproterone acetate anymore.”[23] An earlier version of this review reported that the authors regarded the drug “rather than age” as the leading candidate. They did not say that — they listed age first, and this review quoted them doing so two paragraphs earlier. Elaborating a mechanism and acting on it is a strong signal of where they place their weight, but it is an inference, and the inference is ours.

And the dose matters, because it is not the dose used today. The adults in this cohort received cyproterone acetate at 25–100 mg daily; the adolescents received triptorelin. That is a protocol fact, stated in the paper’s Methods — which is why the confounding is total rather than incidental.[23] Contemporary trans-care practice has converged on roughly 10–12.5 mg. The fertility signal therefore has the same shape as the meningioma signal that reshaped cyproterone prescribing: historical high-dose exposure, with the magnitude at present-day doses unquantified. It differs in one respect, and that difference is why it belongs in a review at all — the meningioma excess recedes after treatment stops. Germ-cell loss does not.

This is a signal about a medicine, not about an age. Cyproterone acetate remains in use in adult feminising regimens, including in the United Kingdom. It belongs in the anti-androgen literature (ER-009), where the drug is graded against its alternatives, and this review flags it there.

What this leaves standing for adolescents. Adolescents suppressed with a GnRH agonist and then given oestradiol retained immature germ cells in the great majority of cases. That remains true, and it remains the basis for the authors’ argument about cryopreservation. It simply cannot be credited to the timing of treatment, because nobody in the comparison group was treated the same way.

And the paper is an argument for fertility preservation, not a report of foreclosure. Its stated implication is that because the vast majority still harbour immature germ cells, testicular tissue containing spermatogonial stem cells can be cryopreserved at the time of surgery — and that “if maturation techniques like in vitro spermatogenesis become available in the future, harvesting germ cells from orchiectomy specimens might be a promising option for those who are otherwise unable to have biological children.”[23]

One finding is often read as cutting the other way, and it does not support that reading. The study reports no difference in the presence of germ cells or their maturation stage between participants recorded as having briefly stopped hormonal treatment before surgery and those who had not, and none by duration of hormone treatment. The cessation was a short instructed interruption of roughly four weeks, adherence was not verified, and no serum hormone levels were available on the day of surgery.

What that design can establish is narrow. A null result after four unverified weeks tells us nothing about whether spermatogenesis recovers after several months, or a year, or permanently off treatment. It is not a biological non-recovery finding, and it must not be quoted as one. Recovery after a clinically sufficient cessation interval has not been tested.

For an adolescent who completes puberty and then starts oestradiol, spermatogenesis is suppressed during treatment; whether, how completely and how often it recovers on cessation is not quantified, particularly for treatment begun in adolescence (ER-005). Fertility preservation counselling before starting is recommended for exactly this reason.

Causal attribution — 🔴 not isolated. Every participant received both suppression and hormones; there is no untreated comparator. Which component suppresses spermatogenesis, and in what proportion, cannot be determined from this design.
Observed effect size — 🟡 estimable in one large histological cohort (n=214), unreplicated: mature spermatozoa 4.7%, immature germ cells only 88.3%, no germ cells 7.0%.
Clinical importance — high, and more nuanced than the debate allows. Mature spermatogenesis is rare on this treatment at every starting age, while on treatment; germ cells are retained in the great majority, including early starters. Complete germ-cell loss occurred in fifteen specimens, all from adult starters — but age at initiation is perfectly confounded with anti-androgen class by protocol (triptorelin for adolescent starters; cyproterone acetate 25–100 mg daily for adult starters), and the authors name four candidate explanations, of which cyproterone acetate is the one they give a mechanism and act on (§19.1). Whether retained immature germ cells can ever yield a biological child depends on techniques that do not yet exist.
Recovery after cessation — not estimable. A recorded four-week pre-operative interruption was not associated with a different germ-cell maturation stage; adherence and hormonal recovery were not verified, and no serum levels were available on the day of surgery. The study does not estimate whether spermatogenesis recovers after a longer, clinically adequate cessation, and it should not be read as though it were a recovery experiment. No study has tested that.
For birth-registered females — not estimable. No study exists at all.

20. Growth and adult height

Read in primary, height is one of the better-characterised long-term outcomes in this field — and what it shows is not what either side quotes. It is set out at length here because the three papers that carry it are routinely cited through systematic-review summary tables that lose their findings.

Trans boys are not stunted, and finish slightly taller than predicted. Willemsen et al. followed 146 trans boys through puberty suppression and testosterone to adult height. In the 61 with a bone age of 14 years or less at the start of suppression, growth velocity and bone maturation decelerated during the blocker and accelerated once testosterone began. Adult height was 172.0 ± 6.9 cm; the height standard-deviation score was unchanged from baseline (0.1; 95% CI −0.2 to 0.4); and adult height was 3.9 ± 6.0 cm above midparental height and 3.0 ± 3.6 cm above the height predicted from bone age at the start of suppression. The younger the bone age at start, the further above prediction they finished.[25]

Trans girls grow tall. Boogers et al. followed 161 trans girls on the same sequential pathway. Growth decelerated during the blocker and accelerated on oestradiol. Adult height after regular-dose treatment was 180.4 ± 5.6 cm — 1.5 cm below the height predicted at the start of suppression (95% CI 0.2 to 2.7), and not significantly different from target height (−1.1 cm; 95% CI −2.5 to 0.3).[26]

That number is the finding, and it is rarely reported. A trans girl started on a blocker in early puberty and then on oestradiol reaches a mean adult height not significantly different from her midparental target height. Mean adult height remained close to the male-referenced target and was not shifted into a typical female population range. Target height and predicted adult height are statistical estimates with error, and a cohort mean is not an individual outcome. The claim here is about group means against a male-referenced expectation, and nothing more. It is stated here because it is a durable, visible characteristic that the treatment does not alter, and a young person consenting to that treatment should learn it beforehand rather than afterwards.

Reducing adult height is a separate, deliberate and modest intervention. Against regular-dose treatment, high-dose ethinylestradiol reduced adult height by 3.0 cm (95% CI 0.2 to 5.8); high-dose oestradiol by 0.9 cm (95% CI −0.9 to 2.8 — not significant).[26] Three centimetres, from the preparation with the least favourable thrombotic profile in the feminising formulary (ER-005). The authors say the height gain must be weighed against possible adverse effects. This review says the same, and adds that the trade-off is not one a growth chart can adjudicate.

And early suppression may add height in trans boys — but the evidence for that is much weaker than its headline, and this review reported the headline. Persky et al. compared 32 transmasculine youth given a GnRH analogue in early-to-midpuberty before testosterone with 62 treated with testosterone alone in late or post-puberty.[27]

Start with the crude comparison. Final adult height was 166.6 ± 7.5 cm in the sequential group and 164.7 ± 6.7 cm in the testosterone-only group: a difference of 1.9 cm, which was not statistically significant (P = .25).[27]

The paper’s headline figure is the adjusted one: a modelled difference of 7.95 cm. Adjustment therefore expands the observed difference more than fourfold, across two groups that differ in the single variable that most determines final height — pubertal stage at presentation. The model is extrapolating between developmentally non-comparable cohorts.

Three further limitations, each of which an earlier version of this review omitted.

The abstract misstates its own direction. It reads: final adult height “was 7.95 cm greater in the GnRHa + T group than in T-only group (95% CI −10.85, −5.06).” A positive difference cannot carry a wholly negative confidence interval. The results section shows the coefficient was coded for the testosterone-only group (β = −7.95), meaning that group was estimated to be 7.95 cm shorter. The magnitude is unaffected; the abstract’s direction and reference group are wrong as printed.[27]

The significant midparental comparison rests on severe differential missingness. Final height minus midparental target height was +2.3 ± 5.7 cm in the sequential group and −2.2 ± 5.6 cm in the testosterone-only group (P < .01) — but midparental height was available for 26 of the 62 testosterone-only participants. The authors note that parental height may have been recorded selectively, for instance where a clinician was concerned a patient was short. If so, the −2.2 cm figure is biased downward and the between-group contrast with it.[27]

The dose–duration figure comes from twenty-two people. The much-quoted estimate that each additional month of GnRH-analogue monotherapy was associated with 0.59 cm of additional final height (95% CI 0.31 to 0.9) is a complete-case regression with n = 22, containing initial height, BMI z-score, Tanner stage, blocker duration, formulation and race. Duration of suppression, Tanner stage, age and remaining growth potential are biologically entangled, and there are very few observations per parameter. It is what the model prints. It is not a causal, linear, transportable return of 0.59 cm per month, and this review should not have presented it as one. The associated estimate that starting at breast stage 3 rather than stage 2 costs 6.5 cm comes from the same small model.[27]

20.1 What NHS England extracted from Persky, and why it cannot answer the question

An earlier draft of this review reported that Persky found no significant difference in final height, on the strength of the figure tabulated in NHS England’s evidence review. That figure has now been traced. It is not Persky’s comparison, and it is not informative for the question the review was asking.

URN 2417j reviews testosterone monotherapy, and by design excludes anyone who received a GnRH analogue for puberty suppression. Persky’s entire finding is a comparison between those two groups. So the sequential arm — the 32 adolescents whose blocker exposure is the whole point of the paper — falls outside the PICO and was discarded. What remained was the 62 testosterone-only participants.

From those 62, the review extracted this: “no statistically significant difference in final adult height … compared to baseline: mean 164.1 (SD 6.8) cm vs 164.7 (SD 6.7) cm.”[47]

The comparison is height at the start of testosterone against final adult height, in a cohort of late- and post-pubertal adolescents who had already substantially finished growing. It could not have shown a difference, and it does not. It is not evidence that testosterone leaves height unchanged; it is a measurement of the growth that was left to happen in people who had none left. And the figures as printed are impossible: final adult height (164.1 cm) is lower than baseline height (164.7 cm). A final adult height cannot be lower than the height recorded at baseline. The z-score line in the same table reads “mean 0.21 (SD 1.0) vs 21.0 (SD 1.0)” — and a z-score of 21 does not exist.[47] The review also states that “no statistical measures were reported,” while simultaneously describing the difference as not statistically significant.

The scope of this criticism, stated so that it cannot be over-read. What is at issue is a single extraction — how one study’s data were transcribed into one evidence table. It is not a criticism of URN 2417j’s search strategy, its inclusion criteria, its GRADE methodology, or its conclusion, none of which this review disputes. The certainty rating for this outcome was VERY LOW, and correcting the extraction does not change it.

This review will not build an argument on the extracted figure, and neither should anyone else. The extracted comparison is out of scope of the paper’s design, mis-specified as a before-and-after in people who had stopped growing, and printed with at least two numerically impossible values. It is the only place in this review where a figure has been found to be not merely weak but unusable — and it is a figure this review itself reproduced, from a secondary table, before checking.[27] The discrepancies in this passage were reported to NHS England on 12 July 2026.[45]

What none of this establishes. Every one of these studies is uncontrolled or non-randomly compared. Persky’s two groups differ in the thing that most determines final height — pubertal stage at presentation — so the adjusted between-group difference is not an effect estimate for the blocker; it is a comparison between an early-pubertal cohort and a post-pubertal one, adjusted for measured covariates only. Willemsen’s and Boogers’s comparators are predicted and midparental heights, which are statistical constructions with known systematic bias — Boogers attributes the 1.5 cm shortfall against prediction to exactly that. And all three describe the sequential pathway. None of them describes a seventeen-year-old who never had a blocker, which is the route most adolescents now reaching hormones have taken.[13]

And a synthesis problem that is this review’s, not the papers’. The three cohorts are not measuring the same quantity. Height against midparental target, height against predicted adult height, final height in early- versus post-pubertal cohorts, change in a sex-referenced z-score, and a duration-of-suppression regression are five different estimands. A person can finish above a prediction because treatment added height — or because the prediction underestimated them, or because parental heights were misreported or selectively recorded, or because bone-age methods are biased at younger ages, or because the sample was selected on remaining growth potential. Both Willemsen and Persky discuss prediction error and selective ascertainment explicitly. Stacking these estimands into a single claim that “trans boys finish taller” is not supported, and an earlier version of this section did exactly that.

The defensible conclusion. Across three sequential-pathway cohorts, no large reduction in final height was observed. In transmasculine cohorts, final height was modestly above predicted or midparental estimates in some analyses, particularly with earlier suppression — but the magnitude attributable to treatment cannot be isolated from prediction error, pubertal-stage differences and selective missingness. In trans girls, mean adult height was not significantly different from midparental target height and remained within a male-referenced range. Height is not among the characteristics this treatment changes; it is among those that, once the epiphyses have fused, nothing changes.

Whether earlier suppression itself increases final adult height remains uncertain, because none of the available studies provides a developmentally comparable counterfactual.

Causal attribution — 🔴 limited. The cohorts consistently show the direction of growth expected from delayed epiphyseal fusion, but none provides a randomised or developmentally comparable untreated counterfactual. In Persky, the crude between-group height difference was 1.9 cm and not significant (P = .25); the adjusted estimate of 7.95 cm depends heavily on modelling two developmentally different groups, and the significant midparental comparison uses 26 of 62 testosterone-only participants. The evidence supports “no evidence of marked stunting” far more comfortably than it supports a precise treatment-induced height gain.
Observed effect size — 🟡 estimable, and small. Trans boys: adult height ≈172 cm, some 3–4 cm above midparental height and above prediction, more so the earlier the start. Trans girls: ≈180 cm on regular-dose treatment, within about a centimetre of target height. High-dose ethinylestradiol reduces adult height by ≈3 cm; high-dose oestradiol does not significantly.
Clinical importance — 🟢 high, and routinely understated. Adult height in trans girls is unchanged by the full pathway. That is a permanent characteristic the treatment does not address, and its being unaddressed is not a neutral fact to the person it belongs to.
Population directness — indirect. All three cohorts were sequentially treated. None speaks to hormones begun at sixteen or seventeen without prior suppression, which is now the common route.

21. Cardiometabolic

The pattern here is instructive — and it is the reverse of what this review previously reported.

Valentine et al. is a cross-sectional analysis of electronic health records from six US paediatric centres, 2009–2019: 4,172 transgender and gender-diverse young people, propensity-score matched on eight variables to 16,648 controls. Roughly a third of the cohort had been prescribed any gender-affirming hormone. The study makes two distinct sets of comparisons, and they must not be collapsed.[46]

Trans youth against matched controls. Unadjusted, higher odds of overweight/obesity, dyslipidaemia, liver dysfunction, hypertension and PCOS. After adjustment for overweight/obesity, depression and antipsychotic prescription, only overweight or obesity remained significantly raised, at OR 1.2 (95% CI 1.1 to 1.3) — and two outcomes moved significantly the other way: liver dysfunction and dysglycaemia were lower in trans youth than in controls. The authors attribute that reversal to ascertainment: trans youth on hormones are tested routinely, so a control who gets tested has a higher pre-test probability of an abnormal result.[46]

Hormone regimens within the trans cohort. Participants prescribed testosterone alone had higher adjusted odds of dyslipidaemia (OR 1.7, 95% CI 1.3 to 2.3), liver dysfunction (1.5, 1.1 to 1.9), overweight or obesity (1.8, 1.5 to 2.1) and hypertension (1.6, 1.2 to 2.2) than participants not prescribed gender-affirming hormones. With a GnRH analogue added: dyslipidaemia (3.7, 2.0 to 6.7) and liver dysfunction (2.5, 1.4 to 4.3). Oestradiol alone, and a GnRH analogue alone, were not associated with greater adjusted odds of any of the cardiometabolic diagnoses examined. These are the adjusted estimates, they are published with confidence intervals, and the paper states that adjustment changed neither which outcomes were significant nor their direction.[46]

What URN 2417h extracted, and why it is not at fault. 2417h reviews oestrogen monotherapy, and took from Valentine only the 349 participants on oestrogen alone. For that arm, against those not on hormones, it reports statistically significant unadjusted odds of dyslipidaemia (OR 1.9, 95% CI 1.3 to 2.7), hypertension (2.3, 1.7 to 3.2) and liver dysfunction (1.6, 1.2 to 2.3) — and states of each that it “was not statistically significant after adjusting for confounders”, the adjusted results not presented. It rates the evidence very low.[49] In this instance the evidence review accurately reflected the oestrogen-monotherapy findings it extracted, and the omission it notes is the primary’s, not its own: Valentine gives the unadjusted oestradiol odds ratios and states only that they ceased to be significant after adjustment, without printing the adjusted figures.[46] The error arose in this review’s synthesis of that scoped extraction with the wider paper. It is not an error in 2417h and it is not a further item for the correction notice.

What this review did with those two sources, and it was wrong. An earlier version listed the testosterone outcomes (which come from the primary) and the oestrogen outcomes (which come from 2417h) in a single sentence, and then applied 2417h’s disclaimer — that the associations lost significance after adjustment and that the adjusted figures were not published — to both. It is true of the oestrogen arm. It is false of the testosterone arm, whose adjusted estimates are among the paper’s headline results and are published with intervals. The disclaimer travelled across the join and erased them. The claim that Valentine was “the only comparator study in NHS England’s evidence base” was also wrong: 2417h names three — Grannis, Kramer and Valentine.[49]

The restraint that travels with it. These are adjusted cross-sectional associations, not causal treatment effects. Exposure and outcome are measured at the same moment, so temporality is unrecoverable; overweight or obesity may sit both upstream and downstream of the other recorded diagnoses. A prescription record establishes neither use, adherence, dose nor duration. Residual confounding remains possible, and confounding by indication is not excluded — the authors name one themselves: young people prescribed combined oral contraceptives for menstrual suppression, who are largely the same people prescribed testosterone, had higher odds of overweight/obesity (1.7) and dyslipidaemia (2.1), which the paper says may be driving part of the testosterone association.[46] The within-cohort hormone comparisons were also not propensity-matched; only the trans-versus-control comparison was. The outcome is an ICD-coded diagnosis, and each outcome is defined as a diagnosis code or at least two abnormal measurements.[46]

And that definition meets a screening gradient the paper does not carry into this comparison. Valentine reports that trans youth were tested for lipids roughly three times as often as controls — total cholesterol in 39.1% versus 11.6%, HDL 38.8% versus 11.1%, triglycerides 39.1% versus 12.3% — and states plainly that this “could lead to more opportunities to receive a cardiometabolic-related diagnosis”. The authors use it to explain their own results: the lower adjusted odds of liver dysfunction and dysglycaemia in trans youth are attributed to routine testing, because a control who gets tested has a higher pre-test probability of an abnormal value.[46] They apply that reasoning to the trans-versus-control comparison. They do not apply it to the hormone comparison — and by their own account it must operate there too, because guideline-based practice tests those on hormones and does not routinely test those who are not. The consequence is specific rather than general: of the four testosterone associations, dyslipidaemia and liver dysfunction are defined by laboratory values and are exposed to the gradient; overweight/obesity and hypertension are not, because the excess testing was, in the paper’s own words, “laboratory (but not anthropometric)”. This is not an objection the paper failed to see. It is one it saw, stated, and did not carry across. And NHS England’s URN 2417h reproduces the outcome definition — “either a diagnosis … or at least two abnormal measurements” — while recording no ascertainment limitation anywhere.[49] The mechanism travelled; the caveat did not. And the authors state the ceiling themselves: they could not evaluate whether the diagnoses preceded or followed the prescription, and the finding “represent[s] an association, not causality”. Prevalence of every outcome other than overweight/obesity was under 15%, and the measured laboratory differences between groups, while statistically significant, were not clinically meaningful. What these results are is a real adjusted safety signal on testosterone. What they are not is proof of harm — and they cannot accurately be described as disappearing after adjustment.

Elsewhere: Taylor found body-mass index unchanged overall with inconsistent results by group; blood pressure showed no clinically significant change across eight studies; HbA1c, glucose and insulin showed no consistent changes.[15] A fall in HDL cholesterol is the most consistent lipid finding, in adolescents as in adults (ER-007). Mullins et al. found no venous thromboembolism or arterial thrombosis in 429 adolescents on testosterone at a median 577 days.[47] Millington et al. found haemoglobin and haematocrit rising progressively to 24 months[4] — the expected androgenic effect, and the one that is genuinely monitorable.[47]

Follow-up duration limits what this section can say. Almost every cardiometabolic study here follows participants for one to three years. Cardiovascular disease is a multi-decade outcome. A surrogate marker measured over two years in a cohort of sixteen-year-olds is not weak evidence about cardiovascular risk; for most purposes it is not evidence about cardiovascular risk at all. This applies symmetrically: the absence of events over three years is not reassurance, and the presence of adjusted cross-sectional associations is not proof of causation.

NICE’s caution stands and should not be softened: one study reported significant increases in blood pressure and body mass index and worsening lipids in transmasculine participants by age 22, and “longer term studies that report on cardiovascular event rates are required.”[14] Surrogate markers are not events. Event data in this population are too sparse and follow-up too short to estimate risk in either direction.

Causal attribution — 🟡 moderate for the haematocrit rise (mechanistic, consistent, dose-related); 🔴 limited for the diagnostic associations. Cross-sectional design; temporality unrecoverable; confounding by indication not excluded.
Observed effect size — 🟡 estimable, and adjusted. Testosterone alone: overweight/obesity 1.8, dyslipidaemia 1.7, liver dysfunction 1.5, hypertension 1.6 — all adjusted, all published with intervals. Oestradiol alone: no association. This is a firmer effect-size estimate than most of Part V, attached to one of its weakest causal designs. Two of the four — dyslipidaemia and liver dysfunction — are laboratory-defined and additionally exposed to a screening gradient the paper documents but does not apply to this comparison; the other two are not.
Clinical importance — not estimable. Cardiovascular event data in this population are too sparse and follow-up too short to estimate risk in either direction. Surrogate diagnoses are not events.

22. Skin

The most concrete adverse effect in the corpus, in this reviewer’s own clinical domain, and a worked example of everything Part III describes.

The mechanism is not in doubt, and it is worth stating properly because it is the clearest case in this review of firm biology sitting on top of absent measurement. The sebaceous gland is an androgen-responsive organ: sebocytes express androgen receptors, perform local androgen metabolism, and convert testosterone to dihydrotestosterone, which binds the androgen receptor with substantially greater affinity than testosterone itself.[33] Androgen signalling drives sebocyte differentiation and sebum production; acne arises from the interaction of that signalling with follicular hyperkeratinisation, the follicular microbiome and an inflammatory response that is present from the earliest lesions.[32,33] Acne is therefore an expected biological consequence of testosterone therapy, not a surprising one. That expectation is exactly what makes the measurement failure below so consequential: nobody needed a study to predict that acne would happen, and so nobody built one capable of saying how often, how badly, or with what lasting mark.

Three studies of birth-registered females on testosterone reported an increase in acne.[15] The adult evidence is more complete and is consistent with them: in trans men, acne appears within months of starting testosterone, peaks in the first year and improves thereafter in most — but that is an adult trajectory, in adults, and it is not a substitute for the adolescent measurement that does not exist.[12] Inside NHS England’s evidence base, Laurenzano et al. report “progression of acne” in 77 of 119 (64.7%), with 23 of those 77 requiring oral treatment, dermatology referral, or both.[47]

Two words that are not interchangeable. Incidence counts people who did not have acne and then got it — it requires a clean denominator of unaffected people at baseline. Progression counts people whose acne got worse, and it can include someone who already had acne before treatment began. A cohort in which most participants were already acneic can post a high progression figure without a single new case.

That figure is a progression proportion, not an incidence. In a footnote to its own evidence table, 2417j records that “sixty one of 106 (57.5%) participants were documented to have acne at baseline, with the majority being mild to moderate.”[47] More than half the cohort had acne before treatment began. The headline figure is therefore reported against a majority-acneic baseline, and the two numbers appear in different parts of the document. This review does not claim to know how any given reader interprets the layout — only that the headline is labelled progression, the baseline appears in a footnote, and the two are not presented together.

A study that does separate incident from prevalent acne exists. Chu et al. followed 60 adolescents under 18 on testosterone: 23% had acne at baseline; of the 46 who did not, 54% developed it within one year.[24] It was retrieved at full text by NHS England’s reviewers and excluded, with the stated reason: “Larger, higher level evidence identified reporting on incidence of acne.”[47]

The exclusion note describes the retained evidence as reporting incidence, although the retained result is labelled progression elsewhere in the same review. That is a documentary inconsistency and it stands on its own. The public record does not show how sample size, outcome definition and clinical informativeness were weighed against each other, and this review does not assert that it knows. The discrepancy may reflect loose terminology rather than a deliberate adjudication. What can be said is that the study which separates incident from prevalent acne was read at full text and excluded, and the one retained does not make that separation.

What none of the adolescent studies measures is severity, or sequelae. The PICO asks for severe acne and receives an ungraded binary. And not one study reports post-inflammatory hyperpigmentation, post-inflammatory erythema, or atrophic, hypertrophic or keloidal scarring — despite these being, for many patients, the longest-lasting consequence of the acne rather than the acne itself.

This is not a niche omission. Post-inflammatory hyperpigmentation is one of the commonest sequelae of inflammatory skin disease, occurs more frequently and more severely in more deeply pigmented skin, and can persist for months to years after the inflammatory lesion has resolved.[35] In acne specifically, post-inflammatory hyperpigmentation is reported in 65% of African American, 48% of Hispanic and 25% of White patients; it persists beyond a year in more than half of those affected, and beyond five years in 22.3%. [34] And the review reporting those figures says something this review will recognise: acne-induced hyperpigmentation is “still inconsistently studied as an outcome in clinical trials, especially in patients with skin phototypes III to VI.”[34] An outcome that matters to patients, and is not measured. The finding is not ours, and it is not about trans people. It is what dermatologists say about their own literature. Susceptibility to hypertrophic and keloidal scarring varies by ancestry and individual predisposition. Post-inflammatory erythema is a different phenomenon: it reflects persistent vascular change rather than pigment deposition, and is more readily visible and more often described in lighter skin; its comparative epidemiology by pigmentation is poorly quantified.

The frequency and persistence of these sequelae vary with constitutive pigmentation, ancestry and individual biology. The adolescent studies provide little or no stratified reporting by ancestry or measured skin pigmentation, and not one of them stratifies any dermatological outcome by either. (The adult Kaiser cohort in §22.1 is not: it matched on self-reported race and ethnicity and is 54.7% non-Hispanic White, 20.2% Hispanic, 9.0% Asian and 8.3% non-Hispanic Black.[37] The design is possible. It has not been done in adolescents.) Two cautions on how that sentence should be read. First, it matters for consequences, not incidence: the same episode of acne can leave a mark that fades in weeks or one that persists for years, and this evidence base cannot distinguish them. Second, Fitzpatrick phototype is an imperfect proxy for any of this. It was devised in 1975 to select ultraviolet dosing for photochemotherapy — originally, and explicitly, to classify people with white skin — and it classifies burning and tanning propensity, not constitutive pigmentation, ancestry, or the tendency to hyperpigment or scar.[40] It is widely used as a proxy for skin colour and for race, and it performs poorly at both. It should not be pressed into service as a research variable for a question it was never designed to answer — which is a reason to measure the right thing, not a reason to measure nothing.

22.1 The comparator study exists — in adults

Everything above is a complaint about a missing design: no comparator, no clean denominator, no severity grading. That design has now been executed, at scale, and it is worth setting beside the adolescent evidence precisely because of how badly it shows it up.

A retrospective matched cohort across four Kaiser Permanente regions followed 280,997 people without baseline acne — 11,234 transmasculine and 9,486 transfeminine individuals, matched to 132,462 cisgender men and 127,815 cisgender women on age, race and ethnicity, enrolment year and region, with up to five years of follow-up. Of these, 12,156 initiated hormone therapy after the index date.[37]

Five-year cumulative acne incidence was 15.8% in transmasculine individuals, against 3.8% in matched cisgender men and 10.5% in matched cisgender women. In the first year after starting testosterone the hazard ratio against matched cisgender men was 8.29 (95% CI 7.11 to 9.68) and against matched cisgender women 2.63 (2.33 to 2.97). It does not then subside to baseline: in subsequent years the risk remained higher than in cisgender men (HR 5.29; 4.45 to 6.28) and in cisgender women (HR 1.69; 1.46 to 1.96).[37]

And one result worth reading twice. Five-year cumulative incidence in transfeminine individuals was 6.0%, against 2.9% in matched cisgender men and 8.4% in matched cisgender women. After starting oestradiol, their acne risk was higher than matched cisgender men (HR 1.56; 1.31 to 1.84) and lower than matched cisgender women (HR 0.53; 0.46 to 0.62).[37] That is a statement about diagnosed acne relative to matched comparison groups. It does not establish a “skin phenotype”, and it does not isolate the contributions of oestradiol, anti-androgen choice, adherence, residual androgen levels, baseline physiology or differential healthcare use. What it does show is that feminising treatment neither abolishes androgen-driven skin disease nor reproduces the acne pattern of matched cisgender women. That belongs in the feminising literature (ER-006), and no adolescent study carries it.

A note on sourcing, and it is not a small one. The figures above are taken from the study’s published abstract; the full text is paywalled and has not been read, and the severity-stratified proportions (reported only as following “similar patterns”) are therefore not given here. Secondary coverage of this paper reported the first-year hazard ratio as 8.56. The abstract says 8.29. An earlier version of this section carried 8.56, taken from that coverage. It has been corrected. A figure reproduced from a secondary source has again proved wrong against the primary record.

Two limits, and they are the reason this does not close the section. First, and decisively: this is an adult cohort. Mean age at the index date was 27.7 years (SD 10.0) in the transmasculine group and 33.2 (SD 13.5) in the transfeminine group.[37] It tells us what happens to adults starting hormones. It does not tell us what happens to a sixteen-year-old in the middle of endogenous puberty, in whom acne is common anyway and the counterfactual is entirely different. Second, the outcome is an ICD code. It measures acne that was diagnosed — which is a function of care-seeking, insurance, dermatological access and how much the acne bothered someone — not acne that occurred. Against a matched comparator that shares those pressures, that is a defensible proxy; it is still a proxy.

So the state of play in this reviewer’s own specialty is as follows. A 281,000-person matched cohort with cisgender controls, an incident-acne endpoint and an abstract-level moderate-to-severe definition exists — in adults. The adolescent literature, which is what the policy question is about, has a progression proportion reported against a majority-acneic baseline, an incidence study that was read at full text and excluded, no severity grading, no sequelae, and no stratification by ancestry or phototype. The design is not impossible. It has been done. It has simply not been done in the people the argument is about. A Kaiser adolescent cohort (5,638 transmasculine adolescents, mean age 13.6 years, matched to more than 58,000 cisgender adolescents) was presented as a conference abstract in 2025 and has not been published in full. If it appears, it answers this section.

See the companion reviews for the clinical detail: Acne in Gender-Affirming Care (CR-001) and Skin Effects of Oestrogen (ER-006).

Causal attribution — 🟡 moderate to strong. Androgen-driven; consistent across three studies and the adult literature.
Observed effect size — 🔴 limited. The most-cited figure is a progression proportion, not an incidence, and the study that reports incidence was excluded.
Clinical importance — 🔴 not estimable. No severity grading; no sequelae data; no stratification by phototype.

23. Pelvic pain

This section exists because a study read for an entirely different reason turned out to contain a common, often severe, poorly treated adverse effect that appears in none of the three appraisals, in none of the certainty tables, and in no public argument about this treatment on either side. It was found while checking a dosing footnote.

In the Melbourne cohort of 158 transmasculine adolescents started on testosterone, 37 (23.4%) had documented pelvic pain — almost one in four. The median interval from starting testosterone to the onset of pain was 1.6 months (IQR 3.4, range 0.3–6.4). Intensity was recorded in only eleven of the thirty-seven; among those eleven, ten were rated severe (7–10/10) and one mild. Documentation of intensity is unlikely to be random — a clinician is more likely to score pain that is severe — so the severity distribution of the other twenty-six cases is unknown, and no statement about the severity of pelvic pain in general can be made from this. The commonest descriptions were “cramps” (45.9%) and “similar to previous period pain” (21.6%). Ten of the thirty-seven participants with documented pain reported pain associated with sexual activity — after orgasm, during penetrative sex, on arousal. Eleven of the thirty-seven had pain associated with breakthrough bleeding.[5] Both denominators are the 37 with recorded pain, not the 158 in the cohort.

And it was not treated well. A wide and unstandardised range of approaches was used — paracetamol, NSAIDs, codeine, danazol, progestogens, a levonorgestrel intrauterine device, GnRH analogues, amitriptyline, physiotherapy, laparoscopy. Improvement or resolution was documented in eight of the thirty-seven, after a median of 7.5 months; outcomes were not adequately documented for the remainder. Absence of documentation is not persistence of symptoms. Norethisterone and medroxyprogesterone were the commonest hormonal agents tried and helped two of ten. The intrauterine device helped four of six.[5]

The one association the study found is also the one it cannot interpret. Pelvic pain was more than five times commoner in those on additional menstrual-suppression agents (36 of 137, 26.3%) than in those who were not (1 of 21, 4.8%) — a risk difference of 21.5% (95% CI 9.8% to 33.2%, p=0.028). The authors give the obvious confounder themselves and this review endorses it: adolescents with a history of dysmenorrhoea were more likely to be prescribed menstrual suppression in the first place, and past dysmenorrhoea was too poorly documented to adjust for. This is confounding by indication, and it means the association cannot be read as menstrual suppression causing pain. It may equally be a marker for who was already prone to it.

A second, weaker signal points the other way and should be stated because it is the only place in this review where prior suppression looks protective. Pelvic pain occurred in 1 of 15 adolescents (6.7%) who had used a GnRH analogue, against 36 of 143 (25.2%) who had not — but the difference is not significant (p=0.196) and the exposed group has fifteen people in it. The authors’ speculation is that avoiding menarche may avoid the central sensitisation that follows repeated dysmenorrhoea.[5] It is a hypothesis with a plausible mechanism and almost no data behind it.

What this cannot establish, and it is a long list. The study is retrospective and single-centre. It has no untreated comparator, so the counterfactual — how much pelvic pain these adolescents would have had anyway — is unknown, and pelvic pain is common in the general population of people assigned female at birth. The authors could not distinguish new-onset pain from pain that pre-dated testosterone. Pain was ascertained from chart review, so chart ascertainment may underestimate symptom occurrence, because a retrospective record cannot capture pain that was never mentioned or never documented. Those with pain had significantly longer follow-up, which may mean longer observation finds more pain, or may mean those in pain were retained in the service. And 23.4% is a prevalence in a treated cohort, not an attributable risk.[5]

Not because 23.4% is a reliable estimate — it is not, and the section says so four times. It is here because an adverse effect documented in roughly a quarter of transmasculine adolescents, rated severe in ten of the eleven cases in which intensity was scored, associated with sexual activity in ten of the thirty-seven with documented pain, and with improvement recorded in only eight of them, was never named as an outcome worth looking for. It was found anyway, by one study, and duly recorded. And having no name, it has no weight. That is the deepest version of the point made in section 7: “no evidence identified” is a fact about a search, not about a drug — and here the search never looked. This is the clearest illustration in the review of why Taylor’s recommendation is not simply for more research but for agreement on core outcomes first.[15]

That claim has been checked directly against the primary, and it holds. The PICO for testosterone monotherapy (URN 2417j) lists its outcomes explicitly. Critical to decision-making: gender incongruence, mental health, quality of life. Important to decision-making: masculinising physical changes, psychosocial impact, fertility, feasibility of genital surgery, cognitive outcomes, detransition, regret. And under safety, a list of striking specificity: thromboembolic disease, cardiovascular events, polycythaemia, pulmonary oil microembolism, pre-diabetes and diabetes, anaemia, breast, ovarian and endometrial cancer, migraine, seizures, impaired liver function, sleep apnoea, sexually transmitted infections, gynaecomastia, skin reactions, severe acne.[48]

Pelvic pain appears nowhere in it. A document that thought to ask about pulmonary oil microembolism, seizures and sleep apnoea did not think to ask about a symptom affecting roughly one in four of the people it is about.

An earlier version of this section then overreached, and the overreach is corrected here rather than quietly removed. It said pelvic pain was absent from the appraisals. It is absent from the PICO. It is not absent from the evidence review: URN 2417j captured Moussaoui under the generic heading “frequency of adverse events” and reports the 23.4% figure in its executive summary and its conclusion.[47] The reviewers found it because the study reported it, not because anyone asked.

An outcome that is never named in a PICO does not vanish — it survives, as an orphan. It is picked up incidentally if a study happens to report it. It gets no outcome-specific certainty rating, no synthesis, no comparison across studies, and no place in the structure of the argument. 2417j goes further and downgrades it, listing pelvic pain among the outcomes assessed with “unvalidated questionnaires” — a description that does not fit a retrospective chart review at all.[47] So the finding survives the search and dies in the architecture. It is in the document. It is in no debate, no summary, no headline and no policy proposition. That is how an outcome disappears: not by being excluded, but by never having been asked for, and therefore never having anywhere to go.

Causal attribution — 🔴 not established. No untreated comparator; new-onset and pre-existing pain not separable; pelvic pain is common in this population at baseline. The temporal association is striking (median onset 1.6 months) and the mechanism is plausible, but neither is a causal estimate.
Observed effect size — 🟡 estimable once. Documented pelvic pain in 37 of 158 (23.4%) in one retrospective cohort; chart ascertainment may miss unreported pain.
Clinical importance — 🟡 potentially substantial, but poorly characterised. Ten of the eleven cases with a recorded intensity score were rated severe, and ten of the thirty-seven with documented pain reported pain associated with sexual activity — but severity was unrecorded for twenty-six cases and treatment outcome for twenty-nine, so the distribution across all cases is unknown.
Population directness — direct. Adolescents, on testosterone, median age 16.6 — the exact population this policy argument is about.
Prior puberty suppression — not estimable. One event in fifteen exposed.

24. Discontinuation and regret

Politically loaded, methodologically thin, and thin in both directions.

The largest dataset on this question is not in NHS England’s evidence base, because it was excluded from it. Boskey et al. followed 1,050 transgender adolescents prescribed gender-affirming hormones between 2007 and 2022. At last contact, 973 (93%) were still taking them. Twenty (2%) stopped for three months or more and restarted. Thirty-seven (4%) stopped without restarting — and of those, five did so because they reidentified with the gender associated with their sex assigned at birth. That is 0.5% of the cohort.[43]

The documented reasons for stopping were, in the authors’ account: having achieved their gender-expression or embodiment goals; difficulty accessing or taking the medication; and an evolution of gender identity such that the drug was no longer needed — in people who still identified as transgender.[43]

Now set that against the outcome NHS England commissioned. The PICO’s detransition outcome is defined as “important to patients because they may choose to discontinue treatment”, with nine listed synonyms: detransitioner, desistence, discontinuation, cessation, termination, reversion, reversal, disidentification, reidentification.[48] Every one of those terms describes the 0.5%. Not one describes goal attainment. Not one describes access failure. Not one describes a person who remains transgender and no longer needs the medicine. The taxonomy is not incomplete. It is mis-shaped — it can record only the reason that accounts for half a percent of the cohort, and is structurally incapable of recording the reasons that account for the rest.

And Boskey wrote the warning against precisely this error, in her own conclusion: that future research must “evaluate the reasons for discontinuation and not make assumptions about detransition and/or regret.”[43] Her study was excluded from URN 2417h on the stated ground that its intervention was out of scope.

What the NHS England evidence base shows. In Chen’s cohort of 315, nine discontinued hormone therapy over two years — approximately 2.9%.[13] In Laurenzano’s series of 119, three stopped testosterone: two because they were satisfied with the effects achieved, and one after reassessing their gender identity, citing concerns about body changes and fertility.[47] That last case is a real detransition and should be reported as one.

What the appraisals found. Taylor’s review reports no synthesised evidence on regret. NHS England’s reviews list regret after receipt of masculinising/feminising medicines as an important outcome in every PICO and record, for every one of them: no evidence identified.[48]

What follows, and what does not. Low observed discontinuation over two years is not evidence of low lifetime regret — follow-up is far too short, and cohorts recruited in specialist clinics under-represent those who leave. Equally, the absence of regret data is not evidence that regret is common. Both inferences are the same error, and both are made constantly.

Causal attribution — not applicable.
Observed effect size — 🔴 limited. Discontinuation ≈3% at two years in one cohort; three of 119 in another.
Clinical importance / lifetime regret — not estimable. No study of adequate duration exists. This is among the largest evidence gaps in the field, and among the most confidently asserted.

25. Cognition

Four studies, all small, and none designed to answer the clinically relevant question — which is not whether these drugs shift brain activity toward sex-typed patterns, but whether they help or harm cognitive development in an adolescent taking them.[15] Two cohort studies of birth-registered females examined whether testosterone shifts brain activity toward male-typical patterns — finding no difference in visuospatial working-memory performance, and slight changes in amygdala lateralisation. A cross-sectional study found better executive functioning, cognitive flexibility and working memory in those receiving hormones than in those not.

Taylor’s own observation is the honest summary: the rationale for these studies varied, several were primarily interested in sex-typed brain differences rather than in whether the treatment helps or harms, and “few studies examined whether hormones influence cognitive development in adolescence, which is identified as a key area of uncertainty.”[15]

Causal attribution — 🔴 not established.
Observed effect size — 🔴 not estimable.
Clinical importance — unknown. The question of whether these drugs affect cognitive development during adolescence has, in effect, not been studied.

26. Cancer

And here the absence has an obvious and legitimate explanation, which should be said before anything else. Cancer is a latency outcome: the interval between exposure and diagnosis is measured in decades, not years. Adolescents treated in the era of modern protocols are not yet old enough to have generated the events. The absence of adolescent cancer data is therefore not a scandal of under-investigation; it is a fact about arithmetic. It is also, for exactly that reason, not reassurance.

No evidence in adolescents, in either direction. The adult data (ER-005, ER-007) are reassuring against a large excess — breast cancer in trans women is more common than in cisgender men and less common than in cisgender women; prostate cancer is rare under androgen suppression; endometrial and ovarian histology in trans men is predominantly benign — but every one of those findings comes from adults with decades of exposure, and none of it speaks to lifetime risk after treatment begun at sixteen.

Not estimable. No adolescent data exist. The adult evidence is reassuring and is not transferable to lifetime risk from adolescent initiation.

Part VI — Evidence quality, and the bottom line

27. The grading, gathered

Outcome Does the treatment cause it? How large? Does it matter to a patient?
Physical virilisation / feminisation 🟢 Yes 🔴 Poorly measured in adolescents 🟢 It is the point of the treatment
Gender dysphoria 🔴 Not established. One study, n=23; comparator is cisgender, not untreated trans 🟡 Numerically large (≈42 UGDS points) — but different sex-specific versions of the scale were used at baseline and follow-up, so this is not a change score on a fixed instrument 🟡 Uncertain. Mean scores crossed the study’s cut-off, but it is a diagnostic cut-off, not a validated response threshold, and the instrument changed between measurements
Mental health 🔴 Not established 🟡 Small (d 0.20–0.39); significant moderation by sex designated at birth: the favourable depression and life-satisfaction slopes were not statistically demonstrated in participants designated male at birth (not evidence of absence) 🔴 Not estimable — no validated MCID for this population
Bone 🔴 Not isolated — and a cross-sectional age gradient in untreated participants casts doubt on whether the z-score comparator tracks 🟡 Mean z-scores returned to pre-suppression levels at measured sites on testosterone; lumbar spine −0.87 on oestradiol at ~12 years, but returning within 2 years on 6 mg 🔴 No fracture data
Spermatogenesis on treatment (transfeminine) 🔴 Not isolated; suppression and hormones inseparable 🟡 One cohort, n=214: mature spermatozoa 4.7%; immature germ cells 88.3%; none 7.0% (all adult starters) High — germ cells retained in most, including early starters; recovery after a sufficient cessation interval untested
Fertility (transmasculine) No study exists.
Adult height 🔴 Limited — no developmentally comparable counterfactual; crude difference in Persky 1.9 cm, P=.25 🟡 No large reduction in final height. Trans boys modestly above midparental/predicted in some analyses; trans girls ≈180 cm, not significantly different from target — height is not feminised. High-dose ethinylestradiol reduces height ≈3 cm 🟢 High — a permanent characteristic the treatment does not change
Cardiometabolic 🟡 Haematocrit; 🔴 the rest 🟡 Adjusted, testosterone alone: overweight/obesity 1.8, dyslipidaemia 1.7, liver dysfunction 1.5, hypertension 1.6. Oestradiol alone: no association 🔴 No event data
Skin / acne 🟡 Androgen-driven 🔴 Most-cited figure is a progression proportion; the study reporting incidence was excluded 🔴 No severity, no sequelae, no phototype data
Pelvic pain 🔴 Not established (no comparator) 🟡 Documented in 23.4% of one cohort; onset ≈1.6 months 🟡 Potentially substantial — severe in 10 of the 11 cases scored, but severity unrecorded in 26 of 37
Discontinuation / regret 🔴 ~3% at two years 🔴 Lifetime regret: not estimable
Cognition 🔴 No study was designed to answer the clinically relevant question.
Cancer No adolescent data.

28. The policy landscape

28.1 The trial that is being funded while the treatment is being withdrawn

Two things are happening at once in the United Kingdom, and a review of this evidence has to hold both in view.

Puberty blockers have been unavailable for gender incongruence in under-18s, on the NHS and privately, since an indefinite ban took effect in December 2024, to be reviewed in 2027. NHS England has paused new hormone prescriptions and consulted on removing them from routine commissioning — on the ground, correctly stated, that the evidence is of very low certainty.[41,42]

And the Department of Health and Social Care and NHS England are jointly funding a randomised trial designed to generate that evidence. The PATHWAYS Trial was paused in February 2026 following safety concerns raised by the MHRA, and resumed under a strengthened protocol. On 23 June 2026 the House of Commons debated and defeated an opposition motion calling on the Government to halt it.[44] The Hansard record of the debate has been read; the division figures are reported in secondary coverage and are not stated here.

This review offers no view on the trial and takes no position on the vote. It records the configuration because the configuration is a fact about the policy environment in which this evidence is being appraised: the same institutions are withdrawing a treatment on the ground that the evidence is insufficient, and funding the study intended to make it sufficient. Those are not contradictory positions. They are what a body may reasonably do when it judges that the evidence does not currently support routine commissioning but might be made to.

One further observation, and it belongs in this review rather than anywhere else. The PATHWAYS protocol provides for a participant’s treatment to be stopped if adverse effects on bone density or cognition emerge, and states that they will receive medical and psychological support if it is.[44] A clinical trial has stopping rules, withdrawal criteria and support for the participants who come off treatment. The ten evidence reviews that will inform the policy have no outcome for what stopping does at all (§11.3). The trial is more careful about withdrawal than the appraisal architecture that surrounds it.

28.2 The consultation

This section describes a fast-moving position as it stood at the time of writing and will date faster than the evidence sections. It is quarantined deliberately: nothing in Parts I to V depends on it.

At the time of writing, NHS England has paused new prescriptions of masculinising and feminising hormones for sixteen- and seventeen-year-olds through the Children and Young People’s Gender Service, with immediate effect from 9 March 2026,[41] and has consulted on a clinical policy that would remove them as a routine commissioning option. The consultation ran for ninety days, from 9 March to 7 June 2026.[42] The final policy has not been published. Status re-checked 13 July 2026: the consultation closed on 7 June 2026 and the final policy has not been published. Jurisdictions elsewhere have reached different conclusions from substantially the same evidence.

28.3 Sweden, read in the original

An earlier version of this review said Sweden had recommended hormone treatment “only within a research framework.” That is a compression of the Swedish position, and reading the primary shows it is the wrong compression. The National Board of Health and Welfare’s national guidance, published in December 2022, says several things at once, and every one of them is dropped by the one-line summaries in circulation.[36]

It does recommend research. The Board notes that the Swedish Agency for Health Technology Assessment concluded that the scientific evidence is insufficient to assess the effects of puberty-suppressing and gender-affirming hormone therapy on gender dysphoria, psychosocial health and quality of life, and it recommends that these treatments be provided in the context of research.[36]

But it does not confine them to research. At group level the Board assesses that the risks of puberty blockers and gender-affirming treatment are likely to outweigh the expected benefits, and on that basis issues a weak, negative recommendation: treatment with GnRH analogues, gender-affirming hormones and mastectomy can be administered in exceptional cases.[36] Exceptional-case provision outside a trial is therefore explicitly retained. “Research only” is not what the document says.

And it distinguishes the two drugs. For puberty suppression, the Board records the participating experts’ view that treatment can in some cases be of great benefit even at Tanner stages 4 and 5, particularly for young people registered male at birth whose later-pubertal masculinisation will make adult presentation very difficult. There is no equivalent carve-out for hormones.[36] A summary that collapses blockers and hormones into a single Swedish “restriction” loses this.

The reasons the Board gives for the shift are the most important thing in the document, and they are the reason it belongs in this review at all. Three factors are said to have moved the balance: the continuing rise in gender-dysphoria diagnoses, particularly among 13–17-year-olds registered female at birth; the documented prevalence of medical detransition among young adults, which the health-technology agency says cannot be quantified; and the fact that the experience-based knowledge of the participating experts is less uniform than it was in 2015.[36]

Not one of those three is is a finding about whether the treatment works. Rising incidence is an epidemiological fact whose significance depends entirely on what you think is causing it. The inability to quantify medical detransition was treated as a reason for precaution. That is a policy response to uncertainty, not a claim that detransition is common, and it is not the absence-of-evidence error named in Part I — though it is exactly the point at which values, rather than evidence, begin to do the work. And expert disagreement is a statement about experts. These are precautionary weights applied under uncertainty, and Sweden is unusually candid in saying so. This review does not criticise them for it. It observes that they are the clearest published illustration of §1’s central claim: the evidence supplied the uncertainty; the values supplied the decision.

What Sweden retains is also instructive, and is almost never quoted: the Board continues to recommend, as measures of high expected benefit and comparatively low risk, sexology counselling, fertility preservation, voice and communication therapy, and hair removal.[36] A jurisdiction can restrict hormones and expand everything else, and this one did.

Only a summary of the 111-page Swedish guidance has been published officially in English; this review has read that summary in the original and has not read the full Swedish text. Where the summary is silent, this review is silent.

Other systems continue provision with consent.

That divergence is not only a disagreement about values. The appraisal exercises broadly agree that certainty is low. They differ in eligibility criteria, in how they classify study quality, in how they synthesise, in the language they use about benefit, and in whether they think the available signals justify any conclusion at all. So jurisdictions differ both in how they interpret the signals and in the policy threshold they apply under uncertainty. Where a system lands depends on which error it more fears, and on how it weighs a persistent consequence of treating against a persistent consequence of not treating — but it also depends on how it reads the same studies.

This review takes no position on that question. It notes only that it is the question, and that the evidence does not answer it.

29. The bottom line

What is firm. These drugs do what they physically do. Testosterone virilises; oestradiol feminises. Some of those changes substantially persist after treatment stops — laryngeal enlargement, clitoral growth, glandular breast tissue — and so do some of the changes produced by a puberty allowed to continue. Neither treating nor deferring is free of lasting consequence, and which consequences fall due depends on pubertal stage, not on age.

What is better characterised than the argument allows. Bone mineral accrual resumes once sex steroids are introduced. In the testosterone-treated group, mean z-scores returned to pre-suppression levels at all three measured sites; in the oestradiol-treated group they returned at the hip and femoral neck but not the lumbar spine. That residual deficit is plausibly explained in part by lower oestradiol exposure under the historical protocol; and a cross-sectional age gradient among untreated trans girls raises the possibility that z-scores may not track in this population, which would undermine the comparator on which the whole inference rests. The bone signal is weaker evidence of harm than it has been made to carry — but “returned to pre-treatment mean z-scores at the measured sites” is not the same as restored bone strength, and fracture data do not exist. Adult height, on the sequential pathway, is among the better-characterised long-term outcomes here and among the least quoted: trans boys finish modestly above midparental and predicted heights in some analyses — though how much of that is attributable to treatment rather than to prediction error, pubertal stage and selective missingness cannot be established — and trans girls reach a mean adult height of around 180 cm, not significantly different from midparental target height. The standard sequential pathway does not materially shift mean adult height toward a typical female range. Reducing it requires high-dose ethinylestradiol, which buys roughly three centimetres at a thrombotic cost. In the one histological cohort available, mature spermatogenesis was rare on this treatment regardless of when it began (4.7%), the great majority retained immature germ cells that could in principle be preserved, and complete germ-cell loss occurred only in those who began treatment as adults. Stopping hormones before surgery did not restore maturation. Whether retained germ cells can ever produce a biological child depends on techniques that do not yet exist.

What is weak. Everything the treatment is principally given for. Gender dysphoria — a central outcome named in every guideline — has been measured, in adolescents, in one study of twenty-three people, whose only comparison group was cisgender, and which administered one version of the dysphoria scale before treatment and a different version afterwards. Its headline figure is not a change score on a fixed instrument. Favourable mental-health changes are observed in the data and are small in complete-case standardised comparisons; causality and clinical importance are uncertain. In the largest prospective cohort, the favourable depression and life-satisfaction slopes were not statistically demonstrated among participants designated male at birth — which is not the same as showing an absence of benefit. And the changes remain uninterpretable as to clinical importance because no minimal clinically important difference exists for any of these outcomes — a fact NHS England wrote into the specification of its own reviews.

And one thing that is neither firm nor weak nor absent, but simply unasked. Roughly one in four transmasculine adolescents on testosterone had documented pelvic pain, beginning a median of six weeks after starting. Where intensity was scored — in only eleven of the thirty-seven — ten were severe. Ten reported pain associated with sexual activity. Improvement was documented in eight. It appears in no appraisal, no certainty table and no public argument, because no PICO asked for it. It was found in this review by accident, while checking a dosing footnote.

What is absent. No randomised trial, and no controlled study capable of establishing causality, was identified by any of the three appraisals or in the further literature reviewed here. Fertility in transmasculine adolescents: unstudied. Cognitive development: effectively unstudied. Lifetime regret: unstudied. Cardiovascular events: unstudied. Cancer: unstudied.

And what follows from that is less than either side would like.

“The evidence is weak” is true, and it does not entail that the treatment should be withdrawn — weak evidence can justify caution, research-only provision, restriction, or continued provision with informed consent, depending on values the evidence cannot supply. “Young people report relief” is also true, and it does not entail that access should be unrestricted, because it rests on evidence that cannot establish causation and cannot tell us whether the relief is large enough to matter.

Both of those sentences will be quoted against each other. They are both in this review, at the same size, on purpose.

The most useful thing this review can say is the thing that helps neither camp: the argument is being conducted over instruments that were never calibrated, about a treatment whose central intended outcome has been measured once, in twenty-three people, against a comparison group that could never answer the question — and measured with one version of the scale at the start and another at the end. That is not a reason to conclude anything. It is a reason to find out.

Declarations, corrections and versions

Competing interests. The author is a trans woman and supports access to gender-affirming hormone therapy for adolescents. This interest was declared to the external reviewer before drafting began. The review was written under adversarial review with symmetric bias-checking as its explicit standard, and it reaches several conclusions the author did not want.

Preparation. Prepared with the assistance of a large language model (Claude, Anthropic), used for source retrieval, arithmetic verification, structural editing and adversarial review. Every claim, every number and every citation was verified against the primary source by the author, and the author is responsible for all of them.

Provenance of sources. Unless otherwise stated in the reference list, cited documents were obtained in their original form and reviewed in full. Where only a published abstract was available, this is stated explicitly. Of 49 references, 42 were obtained and read in full; 5 in published abstract only; 2 in an official extract or translation; and none were not obtained. This is no longer a claim. It is a test. Every reference asserting primary reading has its document in a source archive, keyed by DOI and checked weekly for corrections, and the build refuses to produce a deposit if the archive cannot corroborate it. How we check our work →

Correspondence. Sixteen apparent errors in NHS England evidence reviews URN 2417h and URN 2417j, and in the associated PICO and consultation documents, were reported to the NHS England Clinical Effectiveness Team on 12 July 2026. Five further items were reported on 14 July, completing the findings. At the time of writing they have neither been accepted nor disputed, and the published documents are unchanged.

Method and limitations. This is a narrative evidence review, not a systematic review, and it has not been peer reviewed. Its certainty marks are editorial judgements, not formal GRADE ratings. It is not a prescribing protocol and not a policy position.

Version history

Version Date DOI What changed
1.4 (current) 14 July 2026 10.5281/zenodo.21364084 The archive gate is on. This is the first version whose Verification declaration is a test rather than a claim: every reference asserting a source was read in full now has that document in a source archive, keyed by DOI and checked weekly for corrections, and the build refuses to produce a deposit if the archive cannot corroborate it. Switching it on found five things. Hembree (ref 1) is superseded — two corrigenda were issued and the article was never amended in place; neither affects Table 13, from which the masculinising timetable is taken. Tordoff (ref 20) was corrected in 2022; the correction DOI is now cited and its scope stated. Taylor SC 2002 (ref 34) could not be obtained — the claim it carried was true, and it was removed anyway, because a claim that cannot be evidenced is a claim that cannot be made; it is replaced by an open-access source a reader can check. §14’s scarring claim carried two citations, neither of which supported it; it now cites nothing, and says why. And §15 is enriched from Grimstad’s own abstract: in 46 of the 58 who bled, no cause was identified, and no management method was superior. Correction 14 added. No certainty grading changed.
1.3 13 July 2026 10.5281/zenodo.21329685 §19.1 and §19: the de Nie fertility findings corrected. The authors name four candidate explanations and age is the first; the inference that cyproterone acetate is the leading one is ours, and is now marked as ours. The dose — 25–100 mg daily, the historical high-dose range — is stated. Cross-reference redirected to ER-009. Correction 13 added.
1.2 13 July 2026 10.5281/zenodo.21328838 §18.3 and reference 31: bone-density slopes had been taken from van der Loos’s abstract, which disagrees with the paper’s Results at three of four sites. Correction 12: “read in primary” is not the same as “quoted from the primary.”
1.1 13 July 2026 10.5281/zenodo.21328493 §21: the screening gradient Valentine documents but does not apply to its own hormone comparison. §16: NICE’s footnote asserts the sole gender-dysphoria study had “no control group”; it reports thirty cisgender controls. Correction record restated with no tally; correction 11 added.
1.0 13 July 2026 10.5281/zenodo.21324880 First deposit.

Correction record — 14 entries

Substantive corrections made during production, and the source of each. This record is the review’s only credential for method, and it is published in full. On the direction of corrections. The standard is symmetry: a correction is retained because it is faithful to the source, not because of which way it cuts. Where a correction’s direction is unambiguous, it is stated in the row. Where a correction cuts both ways, that is said, and it is not scored. No ratio is given, and none should be inferred. An earlier version of this record asserted that nine corrections tightened the case for caution and one did the opposite. That count was not derived. It was asserted, and it was wrong — see correction 11.

What was wrong Where it came from What changed
1. Fertility interpretation reversed. An early draft stated that early puberty suppression forecloses fertility. A one-line characterisation in a systematic review, read in place of the paper. §19 rewritten from the primary. Germ cells were retained in the great majority; complete loss occurred only in adult starters — and that difference is fully confounded with anti-androgen class, which the authors state themselves.
2. A controlled study described as uncontrolled. NICE’s summary, read in place of the paper. It has 30 cisgender controls. §16 rewritten — and the same section then missed the instrument-version switch, which was disclosed in the paper’s Methods.
3. An uninterpretable comparison reproduced as a finding. NHS England’s evidence table, read in place of the paper. §20 rewritten and §20.1 added. The extracted comparison was not capable of estimating a treatment effect; the review now separates crude observation from adjusted modelling.
4. Adult-height interpretation overstated. An adjusted estimate read without fully accounting for developmental non-comparability between groups. §20 rewritten. The conclusion is limited to “no large reduction in final height was observed”; any treatment-attributable gain is described as uncertain.
5. Bone findings overclaimed. A dose–response relationship stated as a causal contribution; a cross-sectional age gradient described as a longitudinal trajectory. §18 downgraded; causal attribution reduced from moderate to limited. Identified by the external reviewer.
6. Figures replaced with superseded values. The author-manuscript deposit of Chen et al. held by PubMed Central. The journal’s correction did not reach the version held by PubMed Central. Figures were restored to the values reported in the final published article and its replaced supplementary appendix.
7. A fabricated citation. An author list written from memory rather than from the record. Not verified before writing. Corrected against the primary record. It carried an unverified flag, which is why it did not survive.
8. Ten citations silently renumbered. A reference inserted mid-list. The integrity check reported no orphaned or dangling citations while ten citations pointed at the wrong paper. A claim-to-reference map was built. The check now tests that a citation resolves to a source supporting the claim, not merely that the number resolves.
9. The same renumbering, recurring — and undetected by the check built to catch it. The claim-to-reference map of correction 8 was built but never read. It recorded, in plain sight, eleven claims about a testosterone evidence review filed under a written ministerial statement. References 40–48 were misaligned in the deposited v1.0: nine references, thirty-four citations, every one resolving to the wrong source — and a numeric integrity check that passed cleanly, with no orphans and no dangling citations. The list was rebuilt against the map. A check that is built and not read is not a check.
10. Valentine mischaracterised. The adjusted testosterone associations were erased, and an oestrogen association was wrongly attributed. A splice. The testosterone findings were taken from the primary; the oestrogen findings were taken from URN 2417h; the two were combined into a single narrative sentence; and 2417h’s qualification — “not significant after adjustment; results not presented” — was then applied to the whole of it. The qualification was true only of the oestrogen arm 2417h had extracted. §21 rewritten from the primary, read in full. The paper’s two comparisons are separated; the adjusted testosterone associations are restored (overweight/obesity 1.8, dyslipidaemia 1.7, liver dysfunction 1.5, hypertension 1.6); the absence of an adjusted oestradiol association is stated; and the causal limitations, including the authors’ own, are retained. Direction: this correction restores an adverse association, and so tightens the case for caution. The error it corrects ran the other way. That distinction — between the direction of an error and the direction of its correction — was itself got wrong, and is correction 11.
11. The direction of these corrections was asserted rather than worked out. Three documents signed the same correction three different ways. The de Nie fertility correction (1) was recorded in this review as removing a claim of foreclosure — which takes away a harm. It was recorded in the companion Evidence Note as showing that early suppression cannot be credited with preserving fertility — which takes away a benefit. A third reading, drafted for this record, signed it a third way. All three describe the same rewrite of §19. The correction cuts both ways, and cannot be scored. It removed a foreclosure claim and it removed the attribution of preserved fertility to early timing. Both are true, and they point in opposite directions. This record now states direction only where it can be derived, says so where a correction cuts both ways, and gives no ratio. The standing rule is that direction of effect is recorded, not selected. It was broken here, in the act of writing the record that exists to enforce it.
12. “Read in primary” is not the same as “quoted from the primary.” Van der Loos’s bone-density slopes were taken from the paper’s abstract, not its Results — in a paper this review had read in full. The abstract and the Results of that paper disagree at three of four skeletal sites. The abstract was nearer to hand. The reference carried the annotation Read in primary, and that annotation was true. §18.3 and reference 31 now carry the Results figures, and state the disagreement. A source can be opened and still be quoted from its abstract. The provenance convention was sharpened accordingly: reading a paper in full does not license quoting its abstract; figures come from the Results. Nothing in the finding or the conclusion turned on it — which is exactly why it survived every check.
13. De Nie’s authors were reported as regarding cyproterone acetate “rather than age” as the leading candidate explanation. They did not say that. And the cyproterone dose used in that cohort — 25–100 mg daily, the historical high-dose range — was omitted. The authors list four candidates — age, lifestyle, higher oestradiol doses, or cyproterone acetate — and age is the first. This review quoted that sentence correctly, and then, two paragraphs later, reported the opposite. The inference (they elaborate a mechanism for cyproterone alone, and change their practice on it) is defensible; presenting it as the authors’ own characterisation is not. §19.1 and the §19 confidence box now state what the authors said and mark the inference as ours. The dose is added: the signal comes from the high-dose era, and its magnitude at the 10–12.5 mg now standard is unquantified — the same shape as the meningioma signal, differing in that germ-cell loss does not recede after stopping. The cross-reference is redirected from ER-005 to ER-009, where the drug is graded against its alternatives. Direction: both errors ran toward the more striking reading of a harm.
14. Four references could not be corroborated, and one could not be evidenced at all. The Verification declaration asserted that every unmarked source had been obtained in its original form and read in full. Switching on the archive gate tested that claim for the first time. Until 14 July 2026 the claim rested on memory. There was no artefact behind it, and no check that could have detected its absence. The gate now refuses to build a deposit whose references the archive cannot corroborate. Hembree (ref 1) is superseded — two corrigenda were issued and the article was never amended in place. Neither affects Table 13, from which the masculinising timetable is taken; the reference now says so. Tordoff (ref 20) was corrected in 2022; the correction DOI is now cited and its scope stated. Taylor SC 2002 (ref 34) could not be obtained. The claim it carried was true, and it was removed anyway: a claim that cannot be evidenced is a claim that cannot be made. It is replaced by an open-access source a reader can check. And §14’s scarring claim carried two citations, neither of which supported it — a claim with two irrelevant citations is worse than a claim with none. It now cites nothing, and says why. Every reference now carries a resolvable identifier, and every source claiming primary reading is held in the archive. The build refuses to produce a deposit otherwise.

The correction that produced most of the others. Systematic-review summaries and secondary evidence tables were treated as adequate substitutes for the primary papers. Corrections 1, 2, 3, 6 and 10 have a single cause: a summary was read instead of a source. That is the same failure this review documents in the evidence chain it examines.

References

Population flags mark sources whose primary evidence is from adult or cisgender populations and is extrapolated to adolescents. Provenance of every source is stated in the Verification declaration; exceptions are marked on the reference itself.

  1. Hembree WC, Cohen-Kettenis PT, Gooren L, et al. Endocrine treatment of gender-dysphoric/gender-incongruent persons: an Endocrine Society clinical practice guideline. J Clin Endocrinol Metab. 2017;102(11):3869–3903. doi:10.1210/jc.2017-01658. Read in primary. (Masculinising timetable = Table 13.) The article as published is superseded, and was never amended in place. Two corrigenda have been issued, and the original text still stands uncorrected on the publisher’s site: Corrigendum. J Clin Endocrinol Metab. 2018;103(2):699. doi:10.1210/jc.2017-02548 — corrects a garbled testosterone target in Table 14, item 2a (“the target level is 400–700 ng/dL to 400 ng/dL” should read “the target level is 400–700 ng/dL”). And Corrigendum. J Clin Endocrinol Metab. 2018;103(7):2758–2759. doi:10.1210/jc.2018-01268 — corrects Recommendation 1.1 (diagnosis by mental health professionals and/or trained physicians) and restores a citation omitted from §4.0 on bone health. Neither corrigendum affects Table 13, from which the masculinising timetable used here is taken. Both were read in primary. Population flag: adult consensus tables. The guideline states that its timelines are clinical-observation estimates, not prospective measurement.
  2. Deutsch MB, ed. Guidelines for the Primary and Gender-Affirming Care of Transgender and Gender Nonbinary People. 2nd ed. University of California, San Francisco (UCSF Gender Affirming Health Program); 2016. Population flag: adult.
  3. Baines HK, Connelly KJ. A prospective comparison study of subcutaneous and intramuscular testosterone injections in transgender male adolescents. J Pediatr Endocrinol Metab. 2023;36(11):1028–1036. doi:10.1515/jpem-2023-0237. Read in published abstract; full text not obtained. The review cites this study only for its existence and design — that subcutaneous and intramuscular testosterone have been compared directly and prospectively in adolescents. No figure is taken from it, and the abstract supports the claim in full.
  4. Millington K, Lee JY, Olson-Kennedy J, Garofalo R, Rosenthal SM, Chan Y-M. Laboratory changes during gender-affirming hormone therapy in transgender adolescents. Pediatrics. 2024;153(5). doi:10.1542/peds.2023-064380.
  5. Moussaoui D, Elder CV, O’Connell MA, Mclean A, Grover SR, Pang KC. Pelvic pain in transmasculine adolescents receiving testosterone therapy. Int J Transgend Health. 2024;25(1):10–18. doi:10.1080/26895269.2022.2147118. PMID 38323021. Open access. Read in primary. (n=158; median age at testosterone initiation 16.6 years; 15 [9.5%] with prior puberty suppression; pelvic pain in 37 [23.4%], median onset 1.6 months; improvement documented in 8 of 37 [21.6%]. Reports the testosterone esters correctly — undecanoate 3-monthly, enanthate or mixed esters 3-weekly — contrary to the transcription in NHS England URN 2417j.)
  6. de Vries ALC, Steensma TD, Doreleijers TAH, Cohen-Kettenis PT. Puberty suppression in adolescents with gender identity disorder: a prospective follow-up study. J Sex Med. 2011;8(8):2276–2283. doi:10.1111/j.1743-6109.2010.01943.x. Read in primary.
  7. de Vries ALC, McGuire JK, Steensma TD, Wagenaar ECF, Doreleijers TAH, Cohen-Kettenis PT. Young adult psychological outcome after puberty suppression and gender reassignment. Pediatrics. 2014;134(4):696–704. doi:10.1542/peds.2013-2958. Read in primary.
  8. NHS England / Solutions for Public Health. Evidence reviews: feminising and masculinising medicines in the management of gender incongruence in children and young people. URNs 2417h–2417q; publication reference PRN02421. Completed January 2026; published 9 March 2026.
  9. de Blok CJM, Klaver M, Wiepjes CM, et al. Breast development in transwomen after 1 year of cross-sex hormone therapy: results of a prospective multicenter study. J Clin Endocrinol Metab. 2018;103(2):532–538. doi:10.1210/jc.2017-01927. Read in primary. Population flag: adults.
  10. de Blok CJM, Dijkman BAM, Wiepjes CM, et al. Sustained breast development and breast anthropometric changes in 3 years of gender-affirming hormone treatment. J Clin Endocrinol Metab. 2021;106(2):e782–e790. doi:10.1210/clinem/dgaa841. Read in primary. Population flag: adults.
  11. Dimakopoulou A, Seal LJ. Testosterone and other treatments for transgender males and non-binary trans masculine individuals. Best Pract Res Clin Endocrinol Metab. 2024. doi:10.1016/j.beem.2024.101908. Read in primary. (Clitoral growth onset ~3–4 months, complete by ~1 year; final length ~4–5 cm.) Population flag: adults.
  12. Wierckx K, Van de Peer F, Verhaeghe E, et al. Short- and long-term clinical skin effects of testosterone treatment in trans men. J Sex Med. 2014;11(1):222–229. doi:10.1111/jsm.12366. Read in published abstract; full text not obtained. The review cites this study for its headline trajectory only — acne appears within months, peaks in the first year, and improves thereafter in most — which the abstract states. No figure is taken from the full text. Population flag: adults.
  13. Chen D, Berona J, Chan Y-M, et al. Psychosocial functioning in transgender youth after 2 years of hormones. N Engl J Med. 2023;388(3):240–250. doi:10.1056/NEJMoa2206297. Corrected 19 October 2023 (N Engl J Med 2023;389:1540); all figures cited here are the corrected values. The author manuscript deposited in PubMed Central (NIHMS1877102, 19 July 2023) and the article as first published predate the correction and carry inaccurate values throughout Table 3; they are not relied on here except where the correction leaves a figure unchanged. The Supplementary Appendix used is the replaced version. See reference 38. (n=315; 25 [7.9%] with prior puberty suppression; 262 [83.2%] at Tanner stage 5 at hormone initiation.)
  14. National Institute for Health and Care Excellence (NICE). Evidence review: gender-affirming hormones for children and adolescents with gender dysphoria. Prepared for NHS England; the document states that it was prepared by NICE in October 2020, and NHS England’s own consultation documents and the written ministerial statement of 10 March 2026 refer to it as published in 2021. Both dates are given here because both are used in the source material, and because this is a separate document from NICE’s evidence review on gonadotrophin-releasing hormone analogues. (Ten observational studies; none with a control group; all outcomes very low certainty on modified GRADE; hormones “likely to improve symptoms of gender dysphoria”; no agreed minimal clinically important difference for most outcomes.)
  15. Taylor J, Mitchell A, Hall R, Langton T, Fraser L, Hewitt CE. Masculinising and feminising hormone interventions for adolescents experiencing gender dysphoria or incongruence: a systematic review. Arch Dis Child. 2024;109(Suppl 2):s48–s56. doi:10.1136/archdischild-2023-326670. Read in primary. Two things about this review that are routinely lost, and that this review lost until 12 July 2026. (1) Quality is a study-level score on an adapted Newcastle–Ottawa Scale, converted to percentages: low ≤50%, moderate >50–75%, high >75%, with low-quality studies excluded from synthesis. It is not GRADE certainty in an effect estimate, and the two must not be set against one another (§7). (2) The headline of “40,906 participants” comprises 22,192 adolescents with gender dysphoria — of whom 8,164 received hormones and 14,028 did not — plus 18,714 comparators. The hormone-exposed denominator is 8,164. Studies published after the April 2022 search are discussed narratively in the Limitations, without a reported updated search, flow diagram or appraisal. (Commissioned for the Cass Review; 53 studies — 12 cohort, 9 cross-sectional, 32 pre–post; one rated high quality, measuring side effects only; gender dysphoria measured in one study, body satisfaction in one, fertility in one — and none in birth-registered females.)
  16. Stoffers IE, de Vries MC, Hannema SE. Physical changes, laboratory parameters, and bone mineral density during testosterone treatment in adolescents with gender dysphoria. J Sex Med. 2019;16(9):1459–68. doi:10.1016/j.jsxm.2019.06.014.
  17. Grimstad F, Kremen J, Shim J, et al. Breakthrough bleeding in transgender and gender diverse adolescents and young adults on long-term testosterone. J Pediatr Adolesc Gynecol. 2021;34(5):706–16. doi:10.1016/j.jpag.2021.04.004. Read in published abstract; full text not obtained. Every figure cited here appears in the published abstract: of 232 patients, 58 (one-fourth) had breakthrough bleeding; onset at a mean of 24.3 ± 17.2 months; no identifiable cause in 46 of 58 (79.3%); and no management method was found to be superior. The abstract carries the claim in full.
  18. López de Lara D, Pérez Rodríguez O, Cuellar Flores I, et al. Psychosocial assessment in transgender adolescents. An Pediatr (Engl Ed). 2020;93(1):41–8. (n=23; the single cohort study measuring gender dysphoria before and after hormones.) doi:10.1016/j.anpedi.2020.01.019.
  19. Olson-Kennedy J, Wang L, Wong CF, Chen D, et al. Emotional health of transgender youth 24 months after initiating gender-affirming hormone therapy. J Adolesc Health. 2025;77(1):41–50. doi:10.1016/j.jadohealth.2024.11.014.
  20. Tordoff DM, Wanta JW, Collin A, Stepney C, Inwards-Breland DJ, Ahrens K. Mental health outcomes in transgender and nonbinary youths receiving gender-affirming care. JAMA Netw Open. 2022;5(2):e220978. doi:10.1001/jamanetworkopen.2022.0978. Read in primary. Corrected 26 July 2022Data errors in eTables 2 and 3. JAMA Netw Open. 2022;5(7):e2229031. doi:10.1001/jamanetworkopen.2022.29031. Participant counts in eTables 2 and 3 were off by one, with associated percentages corrected. The corrected supplement, not an earlier download, must be used; the supplement used here reconciles against the main article and is the corrected version. The correction does not affect the adjusted odds ratios cited above, which are reported in the main article and not in the supplement. Exposure is initiation of puberty blockers, gender-affirming hormones, or both — not hormones alone.
  21. Correspondence and authors’ reply. N Engl J Med. 2023;389(16):1536–39. (Letters from Biggs; Hare; Jorgensen; Thompson and Barker. Authors’ reply and editorialists’ reply.) doi:10.1056/NEJMc2302030.
  22. Olson-Kennedy J, Chan Y-M, Garofalo R, et al. Impact of early medical treatment for transgender youth: protocol for the longitudinal, observational Trans Youth Care study. JMIR Res Protoc. 2019;8(7):e14434. doi:10.2196/14434. (The published protocol paper. The study protocol document hosted by NEJM alongside Chen et al. is a separate source and is listed at reference 39.)
  23. de Nie I, Mulder CL, Meißner A, et al. Histological study on the influence of puberty suppression and hormonal treatment on developing germ cells in transgender women. Hum Reprod. 2022;37(2):297–308. Read in primary. (n=214; mature spermatozoa 4.7%; immature germ cells only 88.3%; no germ cells 7.0%, all adult starters.) doi:10.1093/humrep/deab240.
  24. Chu L, Gold S, Harris C, Lawley L, Gupta P, Tangpricha V, Goodman M, Yeung H. Incidence and factors associated with acne in transgender adolescents on testosterone: a retrospective cohort study. Endocr Pract. 2023;29(5):353–5. doi:10.1016/j.eprac.2023.02.002. (Retrieved at full text and excluded from NHS England URN 2417j; stated reason: “Larger, higher level evidence identified reporting on incidence of acne.”)
  25. Willemsen LA, Boogers LS, Wiepjes CM, Klink DT, van Trotsenburg ASP, den Heijer M, Hannema SE. Just as tall on testosterone; a neutral to positive effect on adult height of GnRHa and testosterone in trans boys. J Clin Endocrinol Metab. 2023;108(2):414–421. doi:10.1210/clinem/dgac571. (n=146; pubertal subgroup n=61; adult height 172.0 ± 6.9 cm; +3.9 cm vs midparental height; +3.0 cm vs predicted adult height.) Read in primary.
  26. Boogers LS, Wiepjes CM, Klink DT, Hellinga I, van Trotsenburg ASP, den Heijer M, Hannema SE. Transgender girls grow tall: adult height is unaffected by GnRH analogue and estradiol treatment. J Clin Endocrinol Metab. 2022;107(9):e3805–e3815. doi:10.1210/clinem/dgac349. (n=161; regular-dose adult height 180.4 ± 5.6 cm; high-dose ethinylestradiol reduced adult height by 3.0 cm, 95% CI 0.2–5.8.) Read in primary.
  27. Persky RW, Apple D, Dowshen N, et al. Pubertal suppression in early puberty followed by testosterone mildly increases final height in transmasculine youth. J Endocr Soc. 2024;8:bvae089. doi:10.1210/jendso/bvae089. (n=94; GnRHa+T vs T-only; final adult height minus midparental target height +2.3 vs −2.2 cm, P<.01.) Read in primary. Confirmed against the primary: J Endocr Soc. 2024;8:bvae089; advance access publication 2 May 2024. There is no further volume/issue pagination.
  28. van der Loos MATC, Vlot MC, Klink DT, Hannema SE, den Heijer M, Wiepjes CM. Bone mineral density in transgender adolescents treated with puberty suppression and subsequent gender-affirming hormones. JAMA Pediatr. 2023;177(12):1332–1341. doi:10.1001/jamapediatrics.2023.4588. (n=75; ≥9 years of hormones; lumbar-spine z-score change from start of GnRH agonist in those assigned male at birth −0.87, 95% CI −1.15 to −0.59; all other sites and both groups recovered.) Read in primary.
  29. Boogers LS, van der Loos MATC, Wiepjes CM, van Trotsenburg ASP, den Heijer M, Hannema SE. The dose-dependent effect of estrogen on bone mineral density in trans girls. Eur J Endocrinol. 2023;189(2):290–296. doi:10.1093/ejendo/lvad116. (n=87; lumbar-spine height-adjusted z-score gain over 2 years of hormones: 0.14 on 2 mg, 0.42 on 6 mg, 0.68 on ethinylestradiol; median serum oestradiol on 2 mg = 126 pmol/L.) Read in primary.
  30. Schagen SEE, Wouters FM, Cohen-Kettenis PT, Gooren LJ, Hannema SE. Bone development in transgender adolescents treated with GnRH analogues and subsequent gender-affirming hormones. J Clin Endocrinol Metab. 2020;105(12):e4252–e4263. doi:10.1210/clinem/dgaa604. (51 trans girls and 70 trans boys on GnRHa; 36 and 42 on GnRHa plus hormones; z-scores normalised in trans boys, remained below zero in trans girls; no fractures during the study.) Read in primary.
  31. van der Loos MATC, Boogers LS, Klink DT, den Heijer M, Wiepjes CM, Hannema SE. The natural course of bone mineral density in transgender youth before medical treatment; a cross sectional study. Eur J Endocrinol. 2024;191(4):426–432. doi:10.1093/ejendo/lvae126. (333 AMAB and 556 AFAB, aged 12–25, scanned before any treatment; in those assigned male at birth, BMD z-score fell with age at every site — lumbar spine −0.13/year, 95% CI −0.17 to −0.09; total hip −0.04; femoral neck −0.06; total body less head −0.12 — indicating that z-scores do not track in this group.) Read in primary; figures taken from the Results. The abstract disagrees with the Results at three of the four sites. Per the deposit standard, the Results text governs and the disagreement is stated.
  32. Zouboulis CC, Eady A, Philpott M, et al. What is the pathogenesis of acne? Exp Dermatol. 2005;14(2):143–152. doi:10.1111/j.0906-6705.2005.0285a.x
  33. Zouboulis CC. An update on the role of the sebaceous gland in the pathogenesis of acne. Dermatoendocrinol. 2011;3(1):41–49. doi:10.4161/derm.3.1.13900. (Sebocyte androgen-receptor expression; local androgen metabolism; DHT binds the androgen receptor with greater affinity than testosterone; acne as the interaction of androgen signalling, follicular keratinisation, microbiome and inflammation.)
  34. Auffret N, Leccia MT, Ballanger F, Claudel JP, Dahan S, Dréno B. Acne-induced post-inflammatory hyperpigmentation: from grading to treatment. Acta Derm Venereol. 2025;105:adv42925. doi:10.2340/actadv.v105.42925. Read in primary. Open access (CC BY-NC). Population flag: general dermatological populations, not trans-specific. (Prevalence of acne-induced post-inflammatory hyperpigmentation reported as 65% in African American, 48% in Hispanic and 25% in White patients (Perkins et al., n=2,895); present in 87.2% of 262 patients of phototype IV and above in a Middle East cohort; persists beyond one year in more than half of those affected and beyond five years in 22.3%. The review states that acne-induced hyperpigmentation is “still inconsistently studied as an outcome in clinical trials, especially in patients with skin phototypes III to VI” — the same structural omission this review documents elsewhere, arrived at independently by dermatologists writing about their own field. Replaces Taylor SC et al. 2002, which could not be obtained and therefore could not be relied upon; see the correction record.)
  35. Davis EC, Callender VD. Postinflammatory hyperpigmentation: a review of the epidemiology, clinical features, and treatment options in skin of color. J Clin Aesthet Dermatol. 2010;3(7):20–31.
  36. Socialstyrelsen (Swedish National Board of Health and Welfare). Care of children and adolescents with gender dysphoria — summary of national guidelines, December 2022. Article number 2023-1-8330; published 11 January 2023. Read in the original English summary. Only a summary of the full 111-page Swedish guidance (Vård av barn och ungdomar med könsdysfori, December 2022) is officially available in English. All Swedish claims in §28.3 are drawn from that official summary and from no secondary characterisation.
  37. Smith CA, Kaabi O, Manatunga AK, Lash TL, Silverberg MJ, Getahun D, Vupputuri S, McCracken CE, Chen SC, Tangpricha V, Goodman M, Yeung H. Acne incidence and severity in transgender individuals. JAMA Dermatol. 2026;162(3):255–263. doi:10.1001/jamadermatol.2025.5597. PMID 41563779. PMCID PMC12824851 (embargoed until 21 January 2027). Population flag: adults. Mean age at index 27.7 years (transmasculine) and 33.2 (transfeminine). The outcome is ICD-coded acne — diagnosed acne, not acne. Cited from the published abstract; the full text is paywalled and has not been read. Every figure in §22.1 is stated in the abstract. Severity-stratified proportions are reported in the abstract only as following “similar patterns” and are therefore not quantified in this review. Senior author Howa Yeung is also senior author of Chu et al. 2023 — the adolescent incidence study read at full text and excluded from NHS England URN 2417j (§22). The adolescent counterpart from the same programme (Kaiser Permanente multi-centre adolescent cohort; 5,638 transmasculine adolescents, mean age 13.6 years, matched to more than 58,000 cisgender adolescents) was presented as a conference abstract in J Invest Dermatol 2025 and has not been published in full.
  38. Correction: Psychosocial functioning in transgender youth after 2 years of hormones. N Engl J Med. 2023;389(16):1540. doi:10.1056/NEJMx230007. Published 18 October 2023. Read in primary. (Corrects Chen et al., N Engl J Med 2023;388:240–250. States that “numerical values throughout” Table 3 were inaccurate; corrects the analytic sample to 6,138 observations and 237 [75.2%]; corrects the annual slopes for positive affect to 1.12 [95% CI 0.37 to 1.89], life satisfaction to 2.31 [1.65 to 2.99] and anxiety to −1.47 [−2.15 to −0.80]; corrects the severe-depression figures; and confirms that Table 3 and the Supplementary Appendix have been replaced at NEJM.org.)
  39. Trans Youth Care Study. Impact of early medical treatment in transgender youth — study protocol. Published by the New England Journal of Medicine alongside Chen et al. doi:10.1056/NEJMoa2206297#protocol. (nejmoa2206297_protocol.pdf). Read in primary. (Hypothesis 2a pre-specifies self-injury and suicidality as outcomes for the cross-sex hormone cohort; the eight-item Suicidal Ideation Scale is named as the instrument; the survey battery containing it is scheduled at baseline and 6, 12, 18 and 24 months. Neither outcome is reported in Chen et al. or the companion paper.)
  40. Goon P, Banfield C, Bello O, Levell NJ. Skin cancers in skin types IV–VI: does the Fitzpatrick scale give a false sense of security? Skin Health Dis. 2021;1(3):e40. doi:10.1002/ski2.40. (The Fitzpatrick scale was devised in 1975 to select ultraviolet dosing for photochemotherapy and, as Fitzpatrick himself stated, to classify people with white skin; it classifies burning and tanning propensity rather than constitutive pigmentation.)
  41. Written statement to Parliament: NHS England policy on masculinising and feminising hormones. 10 March 2026 (HCWS1391 / HLWS1395). (NHS England paused its existing clinical policy with immediate effect; no new prescriptions initiated through the Children and Young People’s Gender Service at least until a final policy is determined.)
  42. NHS England. Clinical policy: prescribing of masculinising and feminising hormones for children and adolescents who have gender incongruence or dysphoria — public consultation. Consultation open 9 March to 7 June 2026; final policy and consultation report to be published following consideration of responses. Re-verified 13 July 2026: consultation closed 7 June 2026; final policy and consultation report not yet published.
  43. Boskey ER, Scheffey KL, Pilcher S, Barerra EP, McGregor K, Carswell JM, Kant JD, Kremen J. A retrospective cohort study of transgender adolescents’ gender-affirming hormone discontinuation. J Adolesc Health. 2025;76(4):584–591. doi:10.1016/j.jadohealth.2024.11.002. Read in abstract; full text not obtained. (n=1,050 adolescents prescribed gender-affirming hormones 2007–2022; 973 [93%] continuing at last contact; 37 [4%] discontinued without restarting; 5 [0.5%] because they reidentified with the gender associated with their sex assigned at birth. Documented reasons for discontinuation were predominantly goal attainment, access difficulty, or an evolution of gender identity in people who remained transgender. No physiological outcome of discontinuation is reported. Excluded from NHS England URN 2417h as out of scope.)
  44. PATHWAYS Trial. Sponsors: King’s College London and South London and Maudsley NHS Foundation Trust. Funded by the Department of Health and Social Care and NHS England. Health Research Authority research summary; trial protocol and participant FAQ published by King’s College London. Read in the HRA research summary and in secondary coverage. The trial protocol has not been read in full. (Randomised trial of GnRH-analogue treatment in 226 participants under 16, randomly assigned to immediate or one-year-delayed start, compared over two years across quality of life, mental health, physical development, cognitive function and gender-related distress; parallel observational cohort of 300. Paused February 2026 following MHRA safety concerns; resumed under a strengthened protocol. An opposition motion to halt the trial was debated and defeated in the House of Commons on 23 June 2026 — the division figures are reported in secondary coverage and are not stated here. The protocol provides for treatment to be stopped if adverse effects on bone density or cognition emerge.) The Hansard record of the debate of 23 June 2026 has been read. The trial protocol has not been read in full.
  45. Correspondence: Eden Openly to NHS England Clinical Effectiveness Team ([email protected]), 12 July 2026. Sixteen verified discrepancies in URN 2417h, URN 2417j and the associated PICO and consultation documents, each checked against the primary source. The notice states that, in this review’s assessment, correcting the errors does not change the certainty ratings and that several corrections would strengthen rather than weaken the case for caution. Status of the published documents as at 13 July 2026: unchanged.
  46. Valentine A, Davis S, Furniss A, Dowshen N, Kazak AE, Lewis C, Loeb DF, Nahata L, Pyle L, Schilling LM, Sequeira GM, Nokoff N. Multicenter analysis of cardiometabolic-related diagnoses in transgender and gender-diverse youth: a PEDSnet study. J Clin Endocrinol Metab. 2022;107(10):e4004–e4014. doi:10.1210/clinem/dgac469. Read in primary. (Retrospective cross-sectional analysis of electronic health records, 2009–2019; 4,172 transgender and gender-diverse youth at six PEDSnet sites, propensity-score matched on eight variables to 16,648 controls; 1,412 [33.8%] had any gender-affirming hormone prescription. Adjusted, testosterone alone vs no gender-affirming hormone: overweight/obesity 1.8 [95% CI 1.5–2.1], dyslipidaemia 1.7 [1.3–2.3], liver dysfunction 1.5 [1.1–1.9], hypertension 1.6 [1.2–2.2]; the paper states there were no differences in which outcomes were significant, or in directionality, between the unadjusted and adjusted models. Testosterone with a GnRH analogue: dyslipidaemia 3.7 [2.0–6.7], liver dysfunction 2.5 [1.4–4.3]. Oestradiol alone, oestradiol with a GnRH analogue, and a GnRH analogue alone: no outcome significant after adjustment. Against matched controls, after adjustment only overweight/obesity remained significantly raised [1.2, 1.1–1.3], while liver dysfunction and dysglycaemia were significantly lower. The paper reports the unadjusted oestradiol odds ratios and states only that they were no longer significant after adjustment, without giving the adjusted figures — so the “results not presented” noted in URN 2417h is the primary’s omission, not the evidence review’s. Abstract and Results disagree on one interval: the abstract gives testosterone+GnRHa dyslipidaemia as 3.7 [2.1–6.7], the Results text as 3.7 [2.0–6.7]. The Results text is used here.)
  47. NHS England / Solutions for Public Health. Evidence review: testosterone monotherapy for children and young people with gender incongruence — binary transition. URN 2417j; publication reference PRN02421i. Completed January 2026; published 9 March 2026. Read in primary.
  48. NHS England. PICO documents, URN 2417h–2417q (ten documents; publication references PRN02422i–x). Signed off May 2025; published 9 March 2026. All ten read in primary. The outcome sets are a shared template. The wording quoted in this review — the hypogonadism-mitigation sentence, the detransition synonyms, the safety framing of withdrawal, and the statement that no minimal clinically important differences are known — appears in identical form across all ten.
  49. NHS England / Solutions for Public Health. Evidence review: oestrogen monotherapy for children and young people with gender incongruence — binary transition. URN 2417h; publication reference PRN02421vi. Completed January 2026; published 9 March 2026. Read in primary.

Cross-references within the corpus: ER-005 (Feminising Hormone Therapy) and ER-007 (Masculinising Hormone Therapy) grade the adult evidence for every permanence claim in Part IV. ER-014 (Puberty Suppression) grades the blocker phase. ER-006 (Skin Effects of Oestrogen) and CR-001 (Acne in Gender-Affirming Care) carry the dermatological detail referenced in Part V.