Clinical Development

Synthetic Control Arms vs Randomized Trials

The central question in comparing synthetic control arms with randomized trials is not whether a statistical model can produce a comparator. It can.

Synthetic Control Arms vs Randomized Trials

The harder question is whether that comparator represents the patients who would have entered the trial, received the same standard of care, been assessed on the same schedule, and experienced the same clinical pathway had they been randomized.

That distinction matters because the control arm is not a technical accessory to a clinical study. It is the reference point against which benefit, harm, progression, and survival are interpreted. When the reference point is unstable, even an apparently persuasive treatment effect may carry more uncertainty than the headline result suggests.

For our industry, this is where the patient burden and the methodological ambition of synthetic control arms meet. External data can reduce recruitment pressure, shorten development timelines, and make research more feasible in rare diseases or narrowly defined populations. But the use of real-world data, registries, electronic health records, or historical trial data does not remove the need for a credible counterfactual. It makes the work of establishing one more demanding.

A synthetic control arm can reduce the burden of participation, but it cannot reduce the burden of proof.

The Gold Standard: Why Randomization Remains Unrivaled

Randomized controlled trials remain the gold standard because randomization addresses a problem that statistical adjustment can only partly manage: the presence of factors we do not know, do not measure, or do not measure reliably.

When eligible patients are randomly assigned to treatment or control, known prognostic variables such as disease severity, age, prior treatment, and biomarker status can be compared between groups. More importantly, randomization also helps balance unmeasured prognostic factors. A patient’s underlying frailty, adherence capacity, access to specialist care, or subtle differences in disease biology may affect outcomes without appearing cleanly in the baseline dataset. Randomization does not make those factors disappear, but it makes their distribution between study arms less dependent on clinical judgment or data availability.

That is the foundation on which the treatment effect is estimated.

In a conventional randomized trial, we also have a shared protocol. Patients are assessed against defined eligibility criteria, clinical endpoints are collected according to a prespecified schedule, and investigators work within the same operational and safety framework. The control group is therefore not merely a set of untreated patients. It is a group observed within the same study environment, under comparable definitions and expectations.

This alignment becomes especially important when the endpoint is sensitive to timing or assessment practice. Tumour response, progression-free survival, symptom improvement, hospitalization, and treatment discontinuation can all be influenced by when and how they are measured. A difference between trial arms may reflect treatment biology, but it may also reflect differences in imaging intervals, clinical review, endpoint adjudication, or the threshold for recording an event.

Randomization does not eliminate every source of bias. Protocol deviations, missing data, loss to follow-up, and differences in subsequent treatment can still complicate interpretation. Yet it gives the analysis a strong starting point: the two groups were created through a process that did not depend on the investigator’s or patient’s expectation of outcome.

That starting point is difficult to reproduce retrospectively.

Why an external comparator is attractive

There are situations in which a traditional randomized trial creates substantial practical and ethical difficulty. The eligible population may be very small. The disease may progress quickly. The available standard of care may be poorly tolerated or regarded as inadequate by patients and clinicians. Recruitment may be concentrated in a limited number of specialist centres, leaving people far from a meaningful care pathway.

In these settings, an externally controlled design may appear to offer a more proportionate route to evidence generation. A patient receiving an investigational treatment is compared with a cohort assembled from data outside the current trial rather than with patients concurrently randomized to a control arm.

The potential advantages are real:

  • fewer patients may need to be assigned to a control condition;
  • recruitment may be more feasible in rare or highly selected diseases;
  • historical clinical information may help characterize outcomes where prospective enrolment is difficult;
  • development teams may be able to use data already generated through routine care, registries, or previous studies.

But those advantages are not automatic. They depend on whether the external cohort is sufficiently comparable to the treated population and whether the outcome data are collected with enough consistency to support the intended claim.

The question is therefore not simply whether an external comparator is available. It is whether it is exchangeable with the randomized control arm we would have observed under the same circumstances.

What a Synthetic Control Arm Actually Represents

Synthetic control arms use external data sources to construct a comparator cohort. These sources may include electronic health records, disease registries, real-world databases, historical clinical trials, and other structured collections of patient-level information.

The term can make the method sound more artificial than it is. In practice, the patients in the comparator may be real people whose treatment histories and outcomes were documented in routine care or earlier research. What is synthetic is the way those records are selected, aligned, weighted, or combined to represent a control population for the current study.

That distinction is important when discussing patient burden. An external control does not necessarily mean that no one bears the burden of research. Data may have been collected through earlier studies, routine visits, diagnostic procedures, or long-term follow-up. The relevant ethical question is how responsibly those data are used and whether the resulting comparison is strong enough to support decisions that affect future patients.

A synthetic control arm generally requires close alignment across several dimensions:

  • eligibility criteria and baseline disease characteristics;
  • prior and concomitant treatments;
  • disease stage and severity;
  • biomarker status where relevant;
  • calendar period and standard of care;
  • follow-up duration and censoring rules;
  • endpoint definitions and assessment schedules;
  • clinical setting and geographic distribution;
  • availability and completeness of patient-level data.

The more selective the investigational trial, the more difficult this alignment can become. Trial participants often receive care in specialist centres, undergo frequent assessments, and meet narrow inclusion and exclusion criteria. Patients in a real-world dataset may be older, more medically complex, less consistently monitored, or treated under a different standard of care. These differences are not defects in routine-care data. They are features of the lived experience that must be understood before the data are used as a comparator.

A useful comparison looks like this:

DimensionRandomized controlled trialSynthetic control arm
AllocationTreatment is assigned by randomizationComparator patients are identified from external data
Balance of prognostic factorsRandomization balances measured and unmeasured factors in expectationBalance depends on the quality and completeness of available data
Treatment settingShared protocol and contemporaneous trial environmentMay reflect different centres, periods, and care pathways
Endpoint collectionPrespecified assessments and definitionsMay involve variable documentation and assessment timing
BlindingMay be possible depending on intervention and designUsually absent in the external data source
Main uncertaintyProtocol deviations, missing data, and generalizabilitySelection, temporal, measurement, and unmeasured confounding
Regulatory postureEstablished evidence frameworkRequires design-specific justification and early discussion

This is why the phrase “synthetic control” should never be treated as a synonym for “randomized control without the inconvenience.” It is a different evidentiary structure, with a different pattern of strengths and vulnerabilities.

The Statistical Limits of Matching and Weighting

Propensity score matching, inverse probability weighting, and related approaches can improve comparability between treated and external control populations. They are valuable tools, but their role is often misunderstood.

A propensity score estimates the probability that a patient receives a particular treatment based on observed characteristics. Matching or weighting can then create groups that look more similar with respect to those measured characteristics. In an indirect comparison, methods such as Matching-Adjusted Indirect Comparison may also be used to align baseline populations across studies.

These methods can address imbalance in observed variables. They cannot guarantee that the groups are exchangeable in every clinically relevant respect.

That limitation is not a minor statistical footnote. It is the defining vulnerability of a non-randomized comparison.

Suppose the external dataset captures age, disease stage, previous therapy, laboratory values, and performance status. It may still fail to capture the patient’s access to care, treatment adherence, clinician selection patterns, frailty, symptom burden, or the reasons a treatment was started or withheld. A variable can be technically present in a database and still be too inconsistently recorded to serve as a reliable adjustment factor.

The analysis can only adjust for what has been measured with sufficient quality.

What matching can and cannot do

Matching and weighting can help us:

1. Reduce visible baseline imbalance.

If one cohort is younger, less heavily pretreated, or more likely to have a particular biomarker, adjustment may bring the measured distributions closer together.

2. Clarify the target population.

The analysis can define whether the treatment effect is intended to apply to all eligible patients, the trial population, or a narrower subgroup represented in both datasets.

3. Expose data limitations.

A failed match, extreme weighting, or substantial loss of eligible controls may reveal that the populations are not naturally comparable.

4. Support sensitivity analyses.

Different model specifications can show whether the estimated result changes materially when reasonable assumptions are altered.

But these methods cannot:

  • restore randomization after treatment has already been selected;
  • eliminate confounding from variables that were not measured;
  • correct an endpoint that is defined differently across data sources;
  • make a historical standard of care equivalent to a current one;
  • compensate for systematic differences in follow-up intensity;
  • guarantee that the remaining patients represent the same clinical population.

The most responsible analyses make those limits visible rather than presenting adjustment as a final solution. A well-balanced table of baseline characteristics is not proof that the study populations share the same prognosis. It is evidence that certain measured characteristics have been aligned.

Statistical balance is not the same as clinical comparability.

The distinction between measured and unmeasured confounding should remain at the centre of the protocol discussion. When the expected treatment effect is large and the disease has few effective options, an external comparator may provide useful supportive evidence, particularly when the outcome is objective and the patient population is clearly defined. When the expected effect is modest, however, even a relatively small residual bias can materially change the conclusion.

This is why externally controlled trials are generally considered unsuitable when only a modest treatment effect is expected. The smaller the true difference between treatment and control, the less room there is for uncertainty arising from selection, time, measurement, or missing variables.

Regulatory Expectations: FDA and EMA Perspectives

Regulatory engagement should begin before the database has been selected and before the statistical analysis plan has become difficult to change. The question is not merely whether a synthetic control arm can be built retrospectively. It is whether the proposed design can answer the specific clinical question with enough credibility for the intended development decision.

The U.S. Food and Drug Administration’s February 2023 draft guidance on externally controlled trials emphasizes that the treated and external control populations should be as similar as possible with respect to factors known to affect outcomes. That principle sounds straightforward, but it has practical consequences for protocol design, data governance, endpoint definition, and the selection of the target estimand.

An external control should not be assembled only after the treatment cohort has been observed. The design should prespecify, as far as possible:

  • the eligibility criteria that will be applied to the external population;
  • the index date from which follow-up begins;
  • the treatment line and prior therapy requirements;
  • the outcome definitions and assessment windows;
  • the handling of treatment switching and subsequent therapy;
  • the censoring rules;
  • the approach to missing data;
  • the covariates used for matching or weighting;
  • the sensitivity analyses that will test the robustness of the result.

The EMA’s evolving work on single-arm trials and external controls reflects the same underlying concern: a persuasive narrative around unmet need cannot substitute for a credible comparison. An important disease with limited options still requires an evidence base that can distinguish treatment effect from differences in patient selection and clinical context.

Regulatory acceptance is therefore not a universal property of the method. It depends on the indication, the treatment effect being sought, the natural history of the disease, the quality of the external data, and the consequences of uncertainty.

When an external control may be more defensible

A synthetic control arm may be more defensible when several conditions are present:

  • the disease is rare or recruitment to a randomized trial is genuinely constrained;
  • the treatment effect is expected to be substantial rather than modest;
  • the endpoint is objective and clinically meaningful;
  • disease progression is well characterized;
  • eligibility criteria can be applied consistently to the external dataset;
  • the standard of care has remained reasonably stable across the relevant period;
  • patient-level data are available rather than only aggregate summaries;
  • outcome assessment is sufficiently comparable between cohorts;
  • the analysis plan is agreed early with regulators and other key stakeholders.

Even in these circumstances, the external control should usually be understood as part of a broader evidence package rather than as an automatic replacement for a randomized trial.

The distinction between a pivotal efficacy claim and supportive evidence is particularly important. A synthetic control may help identify a strong signal, contextualize outcomes in a single-arm study, or support a development programme in a rare condition. That does not mean it can replace randomized evidence in every phase or therapeutic area.

The Biases That Shape the Final Result

The most significant risks in a synthetic control comparison often arise before the statistical model is fitted. They are embedded in how patients entered the datasets, how care was delivered, and how outcomes were recorded.

Selection bias

Selection bias occurs when the patients included in the treatment and external control groups differ in ways that affect outcomes. The trial may enrol patients who are fit enough to meet narrow criteria, have reliable access to specialist services, or are willing to attend frequent visits. The external dataset may contain a broader and more clinically diverse population, or it may capture only those who were treated at centres with stronger documentation.

Selection can also operate within the trial. Investigators may choose treatment for patients they believe are more likely to benefit, while clinicians in routine care may reserve an established treatment for patients with different risk profiles. If those pathways are not represented in the available covariates, adjustment cannot fully correct the difference.

Temporal bias

A historical control may have received care when the standard of care, diagnostic technology, supportive treatment, or treatment sequencing was different. Even a relatively short interval can matter in rapidly changing fields such as oncology and immunology.

A patient treated several years earlier may have had less access to molecular testing, different criteria for progression, fewer effective subsequent therapies, or a different threshold for hospitalization. Comparing that patient with someone treated today can create an apparent treatment advantage that partly reflects progress in care rather than the investigational product itself.

Temporal bias is not solved by adding calendar year as a covariate. The underlying care pathway may have changed in ways that are difficult to describe numerically.

Assessment and measurement bias

Randomized trials usually specify when imaging, laboratory assessments, symptom evaluations, and safety reviews occur. External data may be collected according to ordinary clinical need, which is entirely appropriate for care but not necessarily comparable for research.

A patient assessed more frequently has more opportunities for progression, adverse events, or treatment modification to be recorded. Another patient may appear to have a longer period without progression simply because the next assessment occurred later.

The same issue applies to patient-reported symptoms and functional outcomes. A protocol-defined instrument administered at regular intervals is not equivalent to an occasional clinical note describing how a patient feels. Both may be clinically useful, but they answer different questions.

Lack of blinding

External controls are generally not blinded in the way a randomized trial may be. Treatment decisions, clinical assessments, and documentation can therefore be influenced by expectations about benefit. This is particularly relevant for subjective endpoints, investigator-assessed response, symptom scores, and decisions about treatment discontinuation.

For that reason, objective outcomes may be more suitable than outcomes heavily dependent on unblinded clinical judgment, although no endpoint is entirely independent of the care context in which it is collected.

Missingness and data provenance

Missing data are not simply empty cells in a spreadsheet. They may reflect the fact that a test was not clinically indicated, a patient was lost to follow-up, a treatment was delivered elsewhere, or a symptom was never documented because the patient did not raise it during a routine visit.

The reason for missingness can itself be related to prognosis. If patients with worsening disease are more likely to leave care or have incomplete records, the apparent outcome in the external cohort may be distorted.

The provenance of the data therefore matters as much as the volume. A large database with uncertain endpoint definitions may be less useful than a smaller registry with carefully adjudicated clinical outcomes.

Designing the Comparison Around Meaningful Endpoints

The endpoint should not be chosen because it is the easiest variable to find in the external dataset. It should be chosen because it reflects a meaningful improvement in the patient’s care pathway and can be measured with reasonable consistency across the two populations.

This is where clinical strategy and statistical design need to remain connected. A treatment that delays radiographic progression but does not improve symptoms, function, treatment-free time, or survival may have a different meaning for patients than a treatment that changes how they live with the disease. Conversely, a survival endpoint may be affected by subsequent therapies and access to care in ways that complicate external comparisons.

Before building the synthetic arm, the development team should ask:

  • What outcome would represent a meaningful difference for patients?
  • Is that outcome recorded in the same way in both populations?
  • Are the assessment intervals comparable?
  • Can progression or response be defined without relying on undocumented clinical judgment?
  • Is follow-up long enough to observe the endpoint?
  • Could differences in subsequent therapy distort the comparison?
  • Does the estimand reflect the treatment effect the trial is intended to establish?

The endpoint and the comparator cannot be designed independently. A highly reliable external control for overall survival may be less reliable for symptom burden. A registry may capture treatment start and death well but provide limited information about imaging-based response. An electronic health record may contain rich clinical detail but inconsistent documentation across sites.

The practical consequence is that synthetic control development is not simply a data-science exercise. It is a clinical operations problem, a protocol design problem, and a patient interpretation problem at the same time.

A Practical Framework for Deciding Between Designs

There is no universal threshold that tells us when a synthetic control arm is acceptable. The decision should be made through a structured assessment of the clinical question and the evidence available to answer it.

A useful sequence is:

1. Define the decision the study must support.

A design intended to generate an early signal has different requirements from one intended to support a major efficacy claim.

2. Describe the natural history and current care pathway.

If outcomes vary substantially according to referral patterns, subsequent therapy, or access to specialist care, an external comparison becomes more difficult.

3. Map the external data to the protocol population.

Apply the same inclusion and exclusion criteria where possible, rather than relying on broad database labels.

4. Compare the timing of care.

Establish whether the external patients were treated during a period in which diagnostic methods, standard therapy, and supportive care were sufficiently comparable.

5. Check endpoint compatibility.

Confirm that the same clinical event is being measured, not merely similarly named.

6. Identify unmeasured prognostic factors.

Ask which variables clinicians use when making treatment decisions but the dataset does not capture reliably.

7. Run sensitivity analyses before interpreting the primary result.

Results that change substantially under plausible assumptions should be presented as uncertain, not simplified into a single definitive estimate.

8. Engage regulators early.

The FDA’s guidance and the EMA’s work in this area both point toward early, design-specific discussion rather than retrospective justification.

This process may conclude that an external control is appropriate, inappropriate, or useful only as supplementary evidence. That is not a failure of the method. A design review that prevents an overstated conclusion is a contribution to clinical development and to patient protection.

The Patient Meaning Behind the Method

The debate about synthetic control arms is sometimes framed as a contest between innovation and traditional methodology. That framing is too narrow. The real question is how to generate trustworthy evidence without imposing unnecessary patient burden or weakening the standards that protect patients from ineffective or unsafe treatment.

Randomized trials can ask patients to accept uncertainty, frequent visits, invasive procedures, and the possibility of receiving a control treatment. Synthetic controls may reduce some of that burden, particularly where the disease is rare or the clinical situation is urgent. But they can also produce uncertainty that patients and clinicians must carry later if the comparison is not credible.

That uncertainty has a lived experience. It may appear as delayed access to an effective treatment, exposure to a treatment that does not work, disagreement about the value of an endpoint, or difficulty deciding whether a benefit seen in a selected study population applies to the person in front of us.

Our responsibility is therefore not to defend one design in the abstract. It is to match the design to the clinical question, the disease, the expected treatment effect, and the quality of the available data. When randomization is feasible and ethically acceptable, it remains the most reliable way to balance both measured and unmeasured prognostic factors. When an external control is proposed, the burden shifts toward demonstrating exchangeability, protecting against temporal and selection bias, and making residual uncertainty explicit.

Final Perspective

To understand how to check synthetic control arms vs randomized trials, we should begin with the counterfactual rather than the algorithm: would these external patients have had a comparable outcome if they had entered the same trial, at the same point in the disease course, under the same standard of care and assessment schedule?

If the answer is uncertain, propensity scores and weighting cannot make that uncertainty disappear. They can improve the comparison, test assumptions, and show where the populations differ. They cannot recreate the protection offered by randomization against factors that were never measured.

Synthetic control arms have a meaningful role in clinical development, particularly in rare diseases, tightly defined populations, and situations where recruitment to a conventional control group presents an exceptional burden. Their value is greatest when their limitations are treated as part of the design rather than as an inconvenient qualification at the end.

In the end, the quality of a comparator is measured not by how sophisticated its construction appears, but by whether it helps us make a sound decision for the patient whose care pathway comes next.

FAQ

What is the difference between a synthetic control arm and a randomized control arm?
A randomized control arm consists of patients assigned to control treatment within the same trial protocol and clinical environment. A synthetic control arm is constructed from external data, such as electronic health records, registries, historical trials, or real-world databases.
Can propensity score matching replace randomization?
No. Matching and weighting can improve balance in measured characteristics, but they cannot restore randomization or eliminate confounding from variables that were not measured or were recorded unreliably.
When may a synthetic control arm be appropriate?
It may be more defensible when the disease is rare or recruitment is genuinely constrained, the expected treatment effect is substantial, the endpoint is objective and clinically meaningful, and the external data can be aligned with the trial population and care period.
What are the main sources of bias in synthetic control arms?
Key risks include selection bias, temporal bias, differences in endpoint assessment and measurement, lack of blinding, and missing data or uncertain data provenance.
Why are synthetic control arms generally unsuitable for modest treatment effects?
When the true treatment difference is small, even a relatively small residual bias from patient selection, timing, measurement, or missing variables can materially change the conclusion.

Read also