Eden Openly

Reading the Evidence

How we read the evidence

This page explains why the evidence in transgender medicine — and in much of clinical medicine — has the shape it does, and how to read it well.

Our Evidence Reviews and Evidence Notes refer back here rather than re-explaining the same cautions each time. If you have ever wondered why a review says “associated with” instead of “causes,” why so much rests on a handful of cohort studies, or why a figure is flagged as borrowed from research on cisgender people, this is the explanation. You can read it straight through or dip into the section you need.

Why this page exists

Every review we publish does the same quiet work in the background: it separates what is well established from what is reasonable inference, and both of those from what has simply been repeated until it sounds true. The reason that work is necessary is that the evidence in this field has a characteristic shape — relatively few randomised trials, a heavy reliance on observational cohorts, and a good deal of careful borrowing from adjacent populations. Understanding that shape is the single most useful thing a reader can bring to any individual review. Rather than re-argue it in every piece, we set it out once, here, and link back.

None of what follows is unique to transgender medicine. It is ordinary clinical epistemology. But the pressures are unusually concentrated in this field, and the consequences of misreading the evidence — in either direction — are unusually high.

Why randomised controlled trials are scarce

The randomised controlled trial is usually called the gold standard of medical evidence, and readers are right to ask why so few exist here. The answer is mostly structural, not a failure of diligence.

To run a trial you generally need equipoise — genuine uncertainty about whether a treatment helps. For some questions that uncertainty is genuinely gone: you cannot ethically randomise people who need a treatment to a placebo for years to measure an outcome. But that does not apply to every question. Many remain open to experiment — comparing routes, doses and formulations, monitoring intervals, anti-androgen choices and other comparative regimens, or specific risk-mitigation strategies — and some have been studied exactly that way. The scarcity is concentrated on the long-term, whole-treatment questions, not on everything. On top of this, the populations are comparatively small and clinically varied, which makes adequately powered trials hard to assemble; several of the outcomes that matter most — cardiovascular events, cancer, fracture — take years or decades to appear; and research funding in this area is limited and, at times, politically contested.

The consequence is an evidence base built largely from observational studies rather than trials. That is worth understanding clearly, because it cuts both ways. It means firm causal claims should be made sparingly. But the lack of trial evidence does not, by itself, show that a treatment is ineffective or unsafe: absence of trial evidence is not evidence of absence of effect. Benefit and safety have to be judged from the whole body of evidence — observational data, clinical experience, biological mechanism — set against a patient’s own priorities.

Why cohort studies dominate — and what they can and cannot show

Most of what is known in this field comes from cohort studies: groups of people followed over time, with their treatments and outcomes recorded. These studies have real strengths. They describe what actually happens to real patients over real timescales, often years, in a way no short trial could. When several independent cohorts point the same way, that consistency carries genuine weight.

What they cannot do as cleanly as a randomised trial is rule out bias as the explanation for what they find. A cohort can show that two things travel together — that people on a given treatment have a particular outcome more often — without that settling whether the treatment produced it. This is not an absolute barrier: strong prospective design, a clear time order (cause before effect), a dose-response pattern, and agreement across several independent lines of evidence can all strengthen a causal reading. But several traps pull the other way. Selection bias means the people who enter a clinic or registry cohort, or who stay in care, often differ systematically from those who do not. Survivorship and loss to follow-up mean those who drop out may differ from those who remain, quietly skewing the result. And cross-sectional snapshots, which capture exposure and outcome at the same moment, cannot establish which came first. A good review reads cohort evidence for what it reliably gives — direction, trajectory, association, and sometimes more — while naming the ways it could be misleading.

Confounding, in plain terms

A raised number is not, by itself, a verdict on a treatment. The reason is confounding: some third factor that travels alongside the treatment and drives the outcome independently.

Take a concrete example. The long-term Dutch cohort found higher overall mortality among trans women on hormone therapy than in the general population. Read carelessly, that becomes “hormones are dangerous.” Read carefully, it is the opposite of a smoking gun: the excess appeared largely attributable to causes such as cardiovascular disease, lung cancer, HIV-related illness and suicide, and the authors found no indication of a direct harmful effect of the hormone treatment itself. Smoking, comorbidity, and the social determinants of health travel with the exposure, and it is those — more than the medication — that the pattern points to. Confounding is why we resist reading “a higher rate in people taking X” as “X causes it,” and why an honest review keeps asking what else might explain a number before crediting it to the treatment.

Absolute versus relative risk

How a risk is expressed can do as much work as the risk itself. A treatment that “doubles” the chance of an event sounds alarming, but a doubling of a very small baseline is still very small — twice one-in-ten-thousand is two-in-ten-thousand. Conversely, a modest-sounding relative increase can matter a great deal when the baseline risk is already high. Wherever we can, we give the absolute risk, or at least the baseline it is measured against, alongside any relative figure, because “twice as likely” means entirely different things depending on where it starts.

Which outcome, and over what time

Not all outcomes carry the same weight. Hard outcomes — clots, heart attacks, fractures, cancers, deaths — are what ultimately matter. Surrogate outcomes — cholesterol, bone density, hormone levels, scores on a psychological scale — are easier and quicker to measure, and often stand in for the hard ones, but a change in a surrogate does not guarantee a change in the outcome that counts. Time compounds this. A study running a year or two can describe early effects well, but it cannot settle questions about cancer, fracture, cardiovascular disease or mortality, which play out over decades. When a review is cautious or quiet about a long-term outcome, it is often because the evidence has simply not had time to accumulate — which is itself worth knowing.

Extrapolation, and who the evidence is about

Frequently, the most directly relevant evidence does not yet exist, and the best available data come from cisgender women — menopausal hormone therapy, combined contraception — or from cisgender men. Borrowing from those bodies of research is sometimes reasonable, because some physiology is shared. But it is an inference, not a finding, and it can mislead.

The cautionary tale is clotting risk. Much of the historical alarm around oestrogen and blood clots traces to ethinylestradiol, especially at the older, higher doses — a synthetic oestrogen far more thrombogenic than the 17β-oestradiol used in modern therapy. Extrapolating those figures to today’s treatment, the wrong molecule, overstates the risk substantially. That said, modern oestradiol is not risk-free: the risk is real and is shaped by route, dose, age, smoking, inherited clotting disorders and other comorbidity. The same care applies in the other direction — the advice to prefer transdermal oestradiol for people at higher clot risk is well supported in cisgender women but has not been shown head-to-head in trans women, so we present it as sound, borrowed guidance rather than settled fact.

Applicability also varies within trans populations. Evidence gathered in one group may apply differently by age, pubertal status, how long someone has been treated, route and dose, surgical status, smoking and HIV status, and broader social circumstances. A figure that holds on average need not describe a particular person, so “who was this actually measured in?” is always a fair question. Wherever a claim rests on extrapolation — across populations or between subgroups — we flag it, so you can see the difference between a finding and a loan.

Why we write “associated with”

Our reviews lean on careful verbs, and “associated with” is the one that does the most work. It is not timidity. It is a precise claim about the type of evidence: we can see that two things occur together, but the available data do not let us say that one caused the other. When the evidence does support a causal statement, we make it — and say why we can. The habit to watch for in other resources is the reverse: causal language (“X improves Y,” “X causes Z”) laid over evidence that only shows association. The verb is doing quiet but important work, and we try to make ours mean exactly what they say.

How we grade certainty

Where it helps, our reviews rate the strength of the evidence behind a claim using a simple traffic-light scheme: 🟢 reasonably robust, 🟡 moderate, 🔴 limited. The rating reflects four things taken together — the type of evidence (randomised trial, prospective cohort, expert consensus), how consistent the findings are, how directly they apply to trans populations, and the size of the effect relative to its uncertainty.

Formal evidence-grading systems go considerably further. GRADE, the most widely used, weighs risk of bias, inconsistency, indirectness, imprecision and publication bias, and can even raise confidence in observational evidence when the effect is large, shows a dose-response, or where any plausible residual confounding would only have reduced the effect that was seen. Our traffic lights are a lighter-touch signal, not that full apparatus, and they are editorial judgements rather than formal ratings. We use them to let you calibrate — to see at a glance which claims you could lean on and which are still provisional. A 🔴 does not mean a claim is wrong; it means the evidence is thin, and the claim should be held more lightly.

Reading the whole body of evidence

One further caution underlies everything above: certainty is not determined by study design alone. Randomised trials can be small, short, indirect or biased; observational studies can be large, consistent, prospective and clinically informative. A systematic review is only as good as the studies inside it — pooling weak, inconsistent or biased work still yields a low-certainty answer, however formal the method. By the same token, clinical guidelines — such as WPATH’s Standards of Care or the Endocrine Society’s — are syntheses of evidence and practice, not primary evidence in themselves; they are valuable navigational aids, but they consolidate and interpret the underlying studies rather than replace them. We therefore read the whole body of evidence rather than treating the “evidence pyramid” as a shortcut. The questions we keep returning to are always the same: compared with what, in whom, over what timescale, measuring which outcome, and with how much uncertainty?

What this means for you as a reader

A few practical habits make our reviews — and most clinical evidence — easier to read well.

Finally, evidence is only ever half of a real decision. It can estimate benefits, harms and uncertainty; it cannot weigh them for you. What matters most — relief of dysphoria, fertility, tolerance of risk, what care you can actually access — belongs to you. A good clinician helps you set the evidence against those priorities, not in place of them.

This page is about method, not medical advice. It explains how we weigh evidence so that our reviews can show their working — what is known, what is inferred, what is borrowed from another population, and what is not yet known. Wherever those distinctions matter, our Evidence Reviews and Evidence Notes link back here. It is a living page, and we revise it as our thinking sharpens.