The trialsEntry 05.4

Randomisation as evidence

What was established

Allocating treatment by chance sounds like carelessness; it is, in fact, the most powerful device medicine has for distinguishing what works from what appears to work.

A sealed envelope on a desk beside a logbook and a pen, close, office light
The mechanism is mundaneA sealed allocation is what makes two groups comparable, and what lets a trial overturn a conviction.

Why chance is the point

For most of the twentieth century, the treatment a breast cancer patient received was determined by where she was, who was treating her, and when she presented. A surgeon convinced of the radical mastectomy's necessity would operate radically. A more cautious clinician might not. The women who reached each surgeon were not the same population — they differed in age, in how advanced their disease was, in general health, in the particular biology of their tumours. When outcomes were compared across practices, the comparison was already corrupted before it began. Any apparent difference between treatments could as easily be explained by a difference between the patients as by anything the treatment itself had done.

Randomisation removes that problem. When a patient enters a trial and the allocation of treatment is made by a mechanism she and her clinician cannot predict or influence — a random sequence, a sealed envelope, a computer — the two groups being compared are expected to differ only by chance. Prognostic factors known and unknown, measured and unmeasured, distribute roughly equally. No one selects themselves, or is selected, into the arm they expect to do better in. The groups are, within the limits of probability, the same group — which means that any difference in outcome can be attributed to the difference in treatment.

Lifted out of the flow

Chronology

  1. 1920sRonald A. Fisher develops randomisation theory in agricultural statistics
  2. 1948MRC streptomycin trial; first landmark randomised clinical trial
  3. 1971NSABP B-04 launched; radical vs. lesser surgery for breast cancer
  4. 1981Milan I results published; quadrantectomy shown equivalent to mastectomy
  5. 1985Early Breast Cancer Trialists' Collaborative Group begins Oxford overviews
  6. 1993Cochrane Collaboration founded

This logic was not invented for oncology. Randomisation as a formal statistical principle is most often dated to the work of Ronald A. Fisher in agricultural experimentation during the 1920s, and its first widely recognised clinical application was the Medical Research Council's 1948 trial of streptomycin in tuberculosis. What made it powerful in cancer medicine — and specifically in breast cancer — was what it could do to a century of settled conviction.

Conviction overturned

By the 1960s, William Halsted's radical mastectomy — removing the breast, the underlying chest muscles, and the axillary lymph nodes — had been the standard operation for roughly seventy years, not because a trial had established its superiority but because it followed from a coherent theory of how the disease spread, and because no one had tested it against anything else under conditions that would make the test fair. Surgeons who used the operation believed it worked. Patients who survived attributed their survival to it. The operation had cultural and institutional momentum that no amount of observational evidence could simply displace.

Bernard Fisher, working through the National Surgical Adjuvant Breast and Bowel Project — the NSABP, headquartered in Pittsburgh, Pennsylvania — understood that only a randomised trial could break that momentum honestly. NSABP B-04, launched in 1971, assigned women by random allocation to radical mastectomy, simple mastectomy with radiotherapy, or simple mastectomy alone. The results, when they came, showed no meaningful difference in survival across the arms. The radical operation's survival advantage, taken for granted for decades, did not exist in the data from a fair test.

A hospital records room of bound patient files on metal shelving
The same answer by another routeMilan compared a different operation and reached a consistent conclusion, which is why it held.See Milan I

The Italian contribution ran in parallel. Umberto Veronesi and colleagues in Milan launched what would become known as Milan I, a randomised trial comparing radical mastectomy with the much more limited quadrantectomy plus radiotherapy. When published in the New England Journal of Medicine in 1981, it showed equivalent survival — a finding that, alongside Fisher's work, made the surgical argument impossible to sustain on evidential grounds. Neither result could have meant much without randomisation. With it, each was decisive.

The accumulation of trials and what to do with it

Single trials carry their own uncertainties. They are sized to detect a certain magnitude of effect; a smaller real benefit may escape them. They run at one time, in one group of institutions, with patients who may not represent everyone. When a trial produces a striking result, the honest response is not immediate adoption but the question of whether the result would hold in a different setting and a different sample.

This is where meta-analysis enters. The Early Breast Cancer Trialists' Collaborative Group, based in Oxford, England, developed a method of assembling not just the published summaries of every relevant trial but the individual patient data from each — tens of thousands of women followed across dozens of trials over decades. The Oxford overviews, produced periodically since 1985, extracted reliable estimates of treatment effects that no single trial was large enough to see. The benefit of adding chemotherapy to surgery, the survival advantage of tamoxifen in oestrogen-receptor-positive disease, the effect of radiotherapy on long-term mortality: each of these was quantified, and quantified convincingly, not from any single randomised experiment but from their systematic assembly.

NSABP B-04, launched in 1971, assigned women by random allocation to radical mastectomy, simple mastectomy with radiotherapy, or simple mastectomy alone.

The method is not without critics. Combining trials that differ in patient selection, in dosing, in staging criteria, in follow-up duration is a decision that requires justification at every step. The Cochrane Collaboration, founded in 1993 in part to systematise exactly this kind of synthesis, has been a venue for both the advocacy and the critique of pooled analyses. The screening debate — whether population mammography reduces mortality enough to outweigh the harms of overdiagnosis — has been argued, in significant part, through competing meta-analyses of randomised screening trials, with organisations including the World Health Organization and the US Preventive Services Task Force reaching different weightings of the same data.

The structure beneath the result

What randomisation provides, finally, is not certainty but a particular kind of honesty. A trial that assigns treatment fairly, defines its primary outcome before anyone knows which arm is doing better, and reports what it finds regardless of direction gives medicine something it cannot obtain from accumulated clinical experience: a result that is unlikely to be explained by the preferences of the people who set it up. The NSABP trials, the Milan trials, the Oxford overviews — all of it depends on this structural fact. It is why a century of surgical conviction could be revised in a decade, and why the revision held.

Lifted out of the flow

Key concepts in this piece

  • Randomisationassigning treatment by chance so groups differ only randomly, not by selection
  • Meta-analysispooling data across multiple trials to detect effects too small for any single study
  • Overdiagnosisthe identification of disease that would not have caused harm if left undetected; a term central to the screening debate
  • Individual patient datathe Oxford method of collecting raw records from every trial, not just published summaries

Elsewhere in the trials