Clinical Development

Surrogate endpoints versus clinical outcomes in oncology trials

In oncology, a trial can reach its primary endpoint long before we know whether patients will live longer, feel better, or remain able to function in daily life.

Surrogate endpoints versus clinical outcomes in oncology trials

That is the attraction of surrogate endpoints: progression-free survival and objective response rate can shorten development timelines, reduce the time patients spend waiting for an answer, and allow a treatment signal to emerge before overall survival data are mature.

But the same speed creates a clinical responsibility. A surrogate endpoint is not the benefit itself. It is a substitute that we hope will predict the benefit that matters to patients. Learning how to check surrogate endpoints versus clinical outcomes in oncology trials therefore means asking more than whether a biomarker changes, a tumour shrinks, or progression is delayed. We must ask whether those changes reliably translate into longer survival or a better lived experience across patients, treatments, and trials.

That distinction is no longer a technical footnote in protocol design. It shapes the evidence package, the interpretation of an interim result, the design of confirmatory studies, and the confidence with which a treatment enters the care pathway.

Why oncology development relies on surrogate endpoints

Clinical outcomes measure what patients ultimately experience: whether they live longer, feel better, or function better. Overall survival is the clearest example because it captures length of life directly. Patient-reported symptoms, physical functioning, and health-related quality of life can also provide evidence of how treatment affects the person receiving it.

Surrogate endpoints are different. They are used as substitutes for those outcomes because they are expected to predict clinical benefit. In oncology, this may include:

  • Progression-free survival, which measures the time until disease progression or death.
  • Objective response rate, which reflects the proportion of patients whose tumours meet predefined criteria for response.
  • Biomarker response, such as a molecular change that is believed to reflect treatment activity.
  • Pathological or molecular measures, often used when a biological effect can be observed before the ultimate clinical outcome.

The practical appeal is obvious. Overall survival may require prolonged follow-up, particularly when patients receive several subsequent lines of treatment. A study designed around survival can remain open for years after the treatment effect on the tumour has already become visible. By contrast, a surrogate can produce a result sooner and may allow a promising treatment to move through development more quickly.

The FDA’s historical experience illustrates how central this approach has become. Between 2010 and 2012, 45% of new drugs approved by the agency were approved on the basis of a surrogate endpoint. That figure does not mean that surrogate endpoints are inherently unreliable, nor that approvals based on them lack value. It does show that substitutes for direct clinical benefit occupy a substantial part of the modern regulatory landscape.

The ethical question is what happens after the early signal has been accepted. A shorter development programme may reduce patient burden in one sense, but it can also leave uncertainty in place if the surrogate has not been convincingly linked to outcomes that matter at the bedside. A treatment that changes radiographic progression without improving survival, symptoms, or function may still have a role, but that role must be understood honestly.

A surrogate endpoint can accelerate an answer, but it cannot remove the need to ask whether the answer matters to the patient.

Clinical outcomes and surrogate endpoints are not interchangeable

The difference between the two endpoint families becomes clearer when we follow the care pathway rather than the statistical table.

Suppose a new therapy produces a higher response rate. That result may indicate that tumours are shrinking more often. It may support a plausible biological mechanism and provide an early signal that the treatment is active. Yet response does not automatically tell us whether patients live longer. Nor does it tell us whether the treatment delays deterioration, preserves independence, reduces symptoms, or introduces toxicities that offset the apparent benefit.

Progression-free survival often sits closer to the patient’s experience than response rate because it captures a period without documented progression. Even so, its meaning depends on how progression is defined, how frequently imaging is performed, how treatment after progression is managed, and whether delaying radiographic progression changes the subsequent course of illness.

Overall survival is also not a perfect endpoint in every setting. It can be influenced by later therapies, crossover, differences in supportive care, and the biology of the disease. But it remains a direct clinical outcome: the endpoint tells us whether people lived longer. That directness is precisely why a surrogate requires validation before it can stand in for survival with confidence.

DimensionSurrogate endpointClinical outcome
What it measuresAn intermediate biological or disease-related signal expected to predict benefitWhether patients live longer, feel better, or function better
Typical oncology examplesObjective response rate, progression-free survival, molecular responseOverall survival, symptoms, physical functioning, quality of life
Main advantageUsually available earlier and may shorten trial follow-upDirectly reflects patient benefit
Main limitationThe relationship with patient benefit can vary by disease and treatment mechanismOften requires longer follow-up and can be affected by later treatment
Evidence neededBiological rationale plus empirical validation across relevant trialsClear definition, reliable measurement, and clinically meaningful interpretation
Main risk in interpretationMistaking treatment activity for meaningful clinical benefitUnderestimating the value of benefit because follow-up is immature

This is why the question of how to check surrogate endpoints versus clinical outcomes in oncology trials cannot be answered by selecting the endpoint that produces the earliest statistically significant result. The task is to establish whether the surrogate consistently predicts the clinical outcome in the specific context in which it will be used.

The Buyse framework: two levels of surrogacy

A persuasive biological story is a beginning, not a validation strategy. The framework developed by Buyse and colleagues requires evidence at both the individual patient level and the trial level.

Individual-level surrogacy

Individual-level surrogacy asks whether patients who do better on the surrogate also tend to do better on the true clinical outcome. In practical terms, the analysis examines the relationship between a patient’s surrogate result and that patient’s subsequent outcome.

For example, patients who experience a longer period without progression may also appear to have longer overall survival. That relationship can support the idea that progression-free survival contains meaningful information about survival.

However, a patient-level association alone is not enough. Patients with less aggressive disease may naturally have better results across several endpoints, regardless of whether the treatment effect on the surrogate is responsible for the survival benefit. A biomarker may also correlate with prognosis without correctly predicting how a particular treatment will alter the course of illness.

This distinction matters because prognostic value and predictive surrogacy are not the same thing. A marker can identify patients with a better or worse expected outcome without demonstrating that an intervention improving the marker will improve the clinical outcome.

Trial-level surrogacy

Trial-level surrogacy asks a different question: when a treatment produces an effect on the surrogate across multiple randomized trials, does the size of that effect predict the treatment effect on the clinical outcome?

This is the level at which a surrogate begins to earn the right to guide development decisions. We are no longer asking only whether patients with a favourable surrogate result survive longer. We are asking whether interventions that improve the surrogate also improve survival, consistently and across trials.

The evidence generally comes from meta-analyses of randomized controlled trials. Researchers compare the treatment effect on the surrogate with the treatment effect on the true endpoint, often reporting a coefficient of determination, or R², that describes how much of the variation in the clinical outcome effect can be explained by the surrogate effect.

Both levels are necessary. A surrogate that performs well at the patient level but poorly at the trial level may be a useful prognostic measure without being a reliable substitute for survival in treatment comparisons. Conversely, a trial-level relationship that is not supported by a credible patient-level connection may be difficult to interpret clinically.

The framework also protects us from a common error in pharmaceutical development: treating a successful endpoint in one drug class, disease stage, or mechanism of action as universally transferable. Surrogacy is contextual. An endpoint validated for one tumour type or treatment strategy is not automatically valid for another.

What the empirical evidence tells us

The evidence base is more cautious than the language of surrogate validation sometimes suggests.

A review of 15 surrogate analyses conducted by the FDA between 2005 and 2022 found that only one demonstrated a strong correlation between a surrogate outcome and overall survival. This does not mean that the remaining endpoints had no clinical relevance. It means that strong evidence of a reliable relationship with overall survival was uncommon in that body of analyses.

A separate review of 65 sets of oncology trials examined correlations between treatment effects on surrogate endpoints and treatment effects on true endpoints. In that review, 34 sets had correlation coefficients below 0.7, 16 fell between 0.7 and 0.85, and 15 exceeded 0.85. The distribution is a useful reminder that the word surrogate covers a wide range of evidentiary performance. Some endpoints may offer a strong prediction in a narrowly defined context, while others provide only a limited or inconsistent signal.

These results should change how protocols describe the endpoint. If the evidence supports a surrogate as an indicator of treatment activity, the protocol should not silently convert that into a claim of proven patient benefit. The distinction may appear semantic, but it affects how investigators explain the study, how participants interpret results, and how later evidence is judged.

PFS and ORR are not equivalent substitutes

In advanced gastroesophageal cancer, a meta-analysis covering 87 trials found that progression-free survival had stronger trial-level surrogacy for overall survival than objective response rate. The reported R² was 0.45 for progression-free survival compared with 0.21 for objective response rate.

Neither value should be read as a universal threshold for every oncology setting. The comparison is valuable because it shows how endpoint choice can alter the credibility of the development argument. A response rate may capture whether tumours shrink, but it does not capture how long that response lasts or whether the disease remains controlled. Progression-free survival incorporates duration and can therefore contain more information about the disease course, although it remains vulnerable to assessment practices and subsequent treatment effects.

EndpointWhat it can tell usWhat it cannot establish by itself
Objective response rateWhether tumours shrink according to prespecified criteriaWhether responses are durable or improve overall survival
Progression-free survivalWhether progression or death is delayedWhether the delay translates into longer survival or better quality of life
Overall survivalWhether patients live longerWhether additional months are accompanied by preserved function or acceptable toxicity
Patient-reported outcomeHow symptoms, functioning, or quality of life change from the patient perspectiveWhether the result predicts survival or applies across all treatment contexts
Biomarker responseWhether a biological process changes after treatmentWhether the change is a validated surrogate for clinical benefit

There is no hierarchy in which one endpoint is always sufficient and another is always inadequate. The endpoint must be matched to the disease, mechanism, line of therapy, subsequent treatment landscape, and patient priorities. A rapidly fatal disease may create a different evidentiary context from a chronic malignancy in which patients live for years with several treatment options.

How to check surrogate endpoints during protocol design

The most useful assessment happens before the primary endpoint is fixed. Once a study has been built around an unvalidated surrogate, later statistical sophistication cannot fully repair the underlying uncertainty.

We can approach the assessment through several connected questions.

1. What is the clinical outcome that the surrogate is supposed to predict?

The protocol should name the true endpoint directly rather than relying on broad language about efficacy. If the intended benefit is longer survival, the surrogate must be assessed against overall survival. If the intended benefit is preserved function or reduced symptoms, the validation question must include those outcomes.

2. Is there a credible biological pathway from the surrogate to the clinical outcome?

The mechanism should be more specific than the general statement that tumour control is good. We need to understand how the intervention is expected to alter the disease and why the measured change should influence survival, symptoms, or function.

3. What is the evidence at the individual patient level?

A correlation within one dataset can be informative, but it does not establish trial-level surrogacy. The analysis should clarify whether the relationship is prognostic, predictive, or both, and whether it remains credible after accounting for disease burden and other relevant factors.

4. What is the evidence across randomized trials?

The central question is whether treatment effects on the surrogate predict treatment effects on the clinical outcome. Meta-analysis across multiple relevant trials is essential because a relationship observed in one study may reflect the characteristics of that study rather than a dependable property of the endpoint.

5. Does the evidence apply to this treatment context?

Validation should be considered in relation to tumour type, disease stage, treatment mechanism, line of therapy, comparator, and background care. A surrogate cannot be treated as portable simply because it has performed well elsewhere.

6. How will progression and outcomes be measured?

Imaging schedules, adjudication, missing assessments, censoring rules, crossover, and post-progression treatment can all affect the apparent relationship between a surrogate and overall survival. Operational details are not separate from endpoint validity; they shape the data from which validity is inferred.

7. What will the patient experience while the surrogate is being measured?

A study can show an improvement in progression-free survival while participants experience substantial toxicity, treatment interruptions, hospital visits, or loss of daily function. Meaningful endpoints should be considered as a connected set rather than as isolated statistical targets.

8. What uncertainty will remain after the primary analysis?

If a treatment is advanced on the basis of a surrogate, the confirmatory evidence should be designed from the outset. The remaining question is not merely whether the surrogate result is positive, but whether the treatment ultimately demonstrates the clinical benefit that justified reliance on the surrogate.

The protocol language matters

A protocol that calls an endpoint a measure of efficacy may leave too much unsaid. We should describe what the endpoint measures, why it is expected to predict benefit, what validation exists, and which uncertainties remain.

This is particularly important in studies where participants may understand a tumour response as a direct promise of longer life. The clinical team carries a responsibility to explain the difference without diminishing the value of early evidence. A response can be meaningful. A delay in progression can be meaningful. But meaningful does not mean interchangeable with overall survival.

For investigators, the distinction also affects the clinical study report. The CSR should preserve the separation between observed treatment activity and demonstrated patient benefit, including the impact of subsequent therapy, missing assessments, treatment discontinuation, and patient-reported outcomes. A polished narrative should not allow a favourable surrogate result to eclipse evidence that is neutral, immature, or uncertain.

Regulatory acceleration and the burden of confirmation

Surrogate endpoints can support efficient development, including pathways that allow a treatment to reach patients while confirmatory evidence is still being collected. That possibility can be valuable when patients have limited options and the disease imposes a high patient burden.

Yet accelerated access does not transform an uncertain surrogate into a confirmed clinical outcome. The treatment may enter practice with an obligation to generate further evidence, and that obligation should be treated as part of the original development strategy rather than an administrative afterthought.

Our industry sometimes speaks about faster development as though speed were an uncomplicated good. For people with advanced cancer, speed can mean access to a treatment when no acceptable alternative exists. It can also mean living with uncertainty about whether the treatment changes the length or quality of life. Both realities belong in the discussion.

Confirmatory studies should therefore be designed around the unresolved clinical question. If the initial evidence is based on response rate, later work may need to establish durability, progression-free survival, overall survival, symptoms, and function. If progression-free survival is the surrogate, the study should continue to assess whether the treatment changes survival and how patients feel and function during the delay.

A surrogate endpoint is most defensible when it is part of a coherent evidence package rather than a shortcut around the evidence package.

The patient burden hidden behind endpoint selection

Endpoint debates can become abstract very quickly. Correlation coefficients and R² values are necessary tools, but they do not describe the full lived experience of trial participation.

A participant may attend repeated imaging visits, undergo invasive assessments, manage adverse effects, wait through treatment interruptions, and reorganise family or working life around the care pathway. If the study is designed around a surrogate, the burden of participation should be weighed against the certainty and relevance of the knowledge it is likely to produce.

This does not mean that every oncology trial must wait for overall survival before it can be useful. Waiting can itself carry a cost, especially when disease is aggressive and available treatments are limited. It means that we should be precise about what the study can establish and careful not to overstate an early endpoint.

Patient-reported outcomes have a central role here. They can show whether a treatment that improves a radiographic measure also preserves quality of life, reduces symptoms, or simply extends time spent receiving a burdensome therapy. These outcomes do not replace survival, and they do not automatically validate a surrogate. They place the surrogate in the context where its meaning can be judged.

The same principle applies to safety oversight. A data safety monitoring board and the clinical trial team must consider not only whether the surrogate is moving in the intended direction, but whether the overall balance of benefit and harm remains acceptable. A faster endpoint does not make toxicity less consequential.

What a credible comparison should conclude

When comparing surrogate endpoints with clinical outcomes, the conclusion should not be that one category is good and the other is bad. The meaningful comparison is between what each endpoint can establish, how quickly it can establish it, and how much uncertainty remains.

Surrogate endpoints are often valuable because they allow us to detect treatment activity sooner. They can support efficient trial design and may be particularly useful when a direct outcome would take many years to mature. But their credibility depends on empirical validation, not on biological plausibility alone.

Clinical outcomes remain the anchor because they describe the patient’s actual experience: longer life, better function, fewer symptoms, or improved quality of life. Overall survival is not immune to complexity, but it is not a proxy. It measures the outcome directly.

The strongest protocol designs make the relationship visible. They specify the clinical benefit of interest, justify the surrogate in the relevant disease and treatment setting, draw on both patient-level and trial-level evidence, and preserve long-term follow-up. They also include the patient’s voice in the interpretation, so that a statistically favourable result is not mistaken for a meaningful result without examining toxicity, burden, and daily life.

The question is not whether a surrogate endpoint is convenient. It is whether the endpoint has earned our confidence that its movement represents a benefit patients can actually feel, sustain, or live longer to receive.

In the end, validation is a promise about translation: from a radiographic or biological signal to the bedside reality of cancer care. We should use surrogate endpoints when they help us learn sooner, but we should keep asking the harder question they are meant to answer. Has the treatment changed the course of illness for the person in front of us, or have we only measured a change that looks promising from a distance?

FAQ

What is the difference between a surrogate endpoint and a clinical outcome?
A clinical outcome directly measures what a patient experiences, such as overall survival or quality of life. A surrogate endpoint is an intermediate biological or disease-related signal, such as tumor shrinkage, used as a substitute to predict those clinical benefits.
Why are surrogate endpoints used in oncology trials?
They allow researchers to detect treatment activity sooner than overall survival data, which can take years to mature. This helps shorten development timelines and provides earlier access to potentially promising treatments.
Does a positive result on a surrogate endpoint guarantee a clinical benefit?
No. A surrogate endpoint is not the benefit itself, and a change in a biomarker or tumor size does not automatically translate into longer survival or improved daily functioning for the patient.
What is the Buyse framework for surrogate validation?
It is a strategy requiring evidence at two levels: the individual patient level, where patients with better surrogate results also have better outcomes, and the trial level, where treatment effects on the surrogate consistently predict treatment effects on the clinical outcome across multiple studies.
Why is overall survival considered the gold standard in oncology?
Overall survival is a direct clinical outcome that measures whether patients live longer. Unlike surrogate endpoints, it does not rely on predictions or biological proxies to define patient benefit.

Read also