Medical research coverage generally reports the conclusion without the design, which is where all the information about how much to believe it actually sits.

The phases

Early phase studies, in small numbers of participants, primarily assess safety and dosing rather than whether the treatment works.

Middle phase studies, in larger groups with the condition, assess whether there is an effect and continue safety monitoring.

Late phase studies, in large populations, compare against existing treatment or placebo and provide the evidence for approval.

Post-approval surveillance continues afterwards, detecting rare effects that trials were too small to find.

Which means a promising early-phase result establishes very little about whether something works, and most treatments that look good early fail later.

Randomisation

Allocating participants to groups by chance rather than by choice.

Which balances known and unknown differences between groups, and is the feature that allows causal conclusions.

Without it, the groups differ systematically and any outcome difference could reflect that rather than the treatment.

Blinding

Concealing which group a participant is in, from the participant and ideally from those assessing outcomes.

Which prevents expectation from influencing reported outcomes, and prevents assessors from interpreting ambiguous findings in the expected direction.

Some interventions cannot be blinded, which is a genuine limitation rather than a failure, and it should be weighed when reading results.

Endpoints

Where a great deal of confusion originates.

A hard endpoint is something that matters directly — death, hospitalisation, a defined clinical event.

A surrogate endpoint is a measurement expected to predict those — a blood marker, a scan finding, a score.

Surrogates allow smaller and shorter trials, and treatments improving surrogates have repeatedly failed to improve outcomes, and in some cases have worsened them.

Which means a trial reporting a surrogate improvement has not established clinical benefit.

Composite endpoints

Combining several events into one count to increase statistical power.

Which works and can mislead, if the composite is driven by the least serious component.

A result reported as reducing a combined outcome may reflect fewer hospital admissions with no effect on death, and the breakdown is generally published even when the headline is not.

Relative and absolute

The most consequential reporting distinction.

A relative reduction describes proportional change. An absolute reduction describes how many people are affected.

A large relative reduction on a small baseline risk is a small absolute effect, and headlines almost always report the relative figure because it is larger.

The number needed to treat — how many people must receive the treatment for one to benefit — is the most interpretable form and is rarely quoted.

Registration and publication

Trials are required to be registered before starting in most jurisdictions, specifying the planned outcomes.

Which allows detection of outcome switching, where reported outcomes differ from planned ones, and studies auditing this have found it common.

Publication bias — positive results being published more readily than negative ones — distorts the literature, and registration was intended partly to make unpublished trials visible.

Reading coverage

Ask what phase, how many participants, what the comparison was, whether the endpoint was clinical or surrogate, and what the absolute effect was.

Those five questions filter out most overstated coverage, and the answers are generally in the paper even when they are absent from the article.

Who participates

A structural limitation on what trials establish.

Eligibility criteria typically exclude people with multiple conditions, those on other medications, older people and, historically, women of childbearing age.

Which produces trial populations healthier and simpler than the patients who will actually receive the treatment.

Regulators have pushed for broader inclusion and representation, and progress has been uneven and the gap remains substantial for several groups.

Non-inferiority trials

A design worth recognising because it is easy to misread.

Rather than showing a new treatment is better, it aims to show it is not meaningfully worse than an existing one.

Which is legitimate where the new treatment has other advantages — fewer side effects, easier administration, lower cost.

The margin defining not meaningfully worse is chosen by the researchers, and a generous margin makes the bar easy to clear.

Interim analysis and early stopping

Trials are sometimes stopped early for benefit, and studies have found that early-stopped trials tend to overestimate effects.

Which is a statistical consequence of stopping when results look favourable, and it is why stopping rules are specified in advance.

Funding and conflicts

Who funds a trial affects how it is designed, analysed and reported.

Analyses comparing industry-funded and independently funded trials have found systematic differences in reported conclusions.

Which does not mean industry trials are fabricated — they are generally well conducted — and the design choices available legitimately favour a sponsor's product.

Comparator choice, dose selection and endpoint selection are all legitimate decisions that shape the result.

Declaration of funding and of author conflicts is now standard, and reading it is a two-second check.