
That number — drawn from an analysis of 791 trials covering over half a million patients — should reframe every endpoint selection conversation you are having right now. If your current protocol design strategy treats surrogate endpoints as a reliable proxy for definitive clinical value, you are building on a foundation that regulators and health technology assessment bodies are increasingly testing.
The August 2025 FDA draft guidance on overall survival assessment is part of that shift. It is not a binding rule, but its recommendations signal where the agency expects sponsors to strengthen randomized oncology studies used to support marketing approval. In particular, the guidance recommends prespecified OS assessment even when overall survival is not the primary efficacy endpoint. That changes how trial architecture should be planned, how follow-up should be funded, and how the final dossier must explain the relationship between early disease control and patient benefit.
Your organization cannot afford to treat endpoint selection as a speed-to-market decision alone. The bottleneck has shifted: it is no longer about how fast you can reach a statistical milestone. It is about whether that milestone holds up under simultaneous regulatory and health technology assessment scrutiny. The gap between accelerated approval and reimbursement has become one of the most expensive vulnerabilities in clinical development. This breakdown maps the problem, the regulatory evolution, and the protocol framework you need to operationalize now.
The Statistical Gap: Why Surrogate Endpoints Often Fail to Predict OS
The evidence is unambiguous. A meta-analysis presented at the 2025 ASCO Annual Meeting and published in JAMA Oncology examined 791 randomized controlled oncology trials involving 555,580 patients and spanning 2002 to 2024. The core finding: only 28% of trials that achieved statistical significance on an alternative primary endpoint ultimately demonstrated an overall survival improvement. Fewer still — 11% — showed meaningful patient-reported quality-of-life gains. A mere 6% demonstrated superiority on both overall survival and global quality of life.
These numbers expose a fundamental scalability problem in how the industry operationalizes surrogate endpoints. Progression-free survival and objective response rate can accelerate development timelines, but their predictive power for what patients and payers actually care about — living longer, living better — is structurally limited. A surrogate endpoint can capture an important aspect of disease control without capturing the full effect of treatment on a patient's life.
That distinction is not semantic. A therapy may delay radiographic progression while introducing treatment-related toxicity, limiting later treatment options, or failing to alter the disease course in a way that extends survival. In other settings, subsequent-line therapies can dilute or obscure an overall survival effect even when the experimental treatment has produced a genuine early benefit. The relationship between the surrogate and OS therefore depends on the treatment landscape surrounding the trial, not only on the statistical result generated inside it.
The limitations are also not uniform. An integrated analysis of trial data shows that PFS consistently demonstrates a stronger patient-level correlation with OS than best overall response, particularly in first-line therapy settings and immuno-oncology trials. That distinction matters enormously when you are drafting a protocol: the choice between PFS and ORR as a primary endpoint is not interchangeable. Each carries a different risk profile for downstream regulatory and HTA outcomes.
ORR is often attractive because it can produce an earlier and more visible signal, particularly in diseases where tumor shrinkage is clinically meaningful. But response is a binary or categorical observation that may say little about durability. A high response rate with short-lived responses can produce a very different patient experience from a more modest response rate accompanied by prolonged disease control. PFS incorporates time and progression events, but even PFS is not automatically a valid proxy for OS. Its interpretability depends on assessment frequency, censoring rules, treatment discontinuation, imaging quality, and the extent to which progression reflects the outcomes that matter in the disease.
The problem compounds when you factor in mechanism of action. Correlations between surrogate and definitive endpoints vary widely by tumor type, line of therapy, and drug class. There are no universal quantitative correlation coefficients that hold across the oncology landscape. This means your endpoint validation strategy cannot be template-driven. It must be indication-specific, therapy-specific, and grounded in the actual disease biology your compound targets.
A protocol that borrows an endpoint from a successful trial in another setting may therefore import the wrong assumptions. A surrogate that performs well in first-line disease may be less informative after several prior therapies. A marker that appears useful for cytotoxic treatment may behave differently for immunotherapy, targeted therapy, or a combination regimen. Even within the same tumor type, changes in standard of care can alter the connection between early disease control and survival.
The practical question is not whether PFS or ORR is a “good” endpoint in the abstract. It is whether the endpoint is sufficiently connected to treatment benefit in the exact clinical context in which your product will be evaluated.
The most expensive assumption in oncology protocol design is that hitting a surrogate endpoint means your drug works. A statistically positive surrogate result does not, by itself, establish longer survival or better quality of life.
Regulatory Evolution: The 2025 FDA Draft Guidance on OS Assessment
In August 2025, the FDA released draft guidance titled Approaches to Assessment of Overall Survival in Oncology Clinical Trials. The document recommends that sponsors evaluate overall survival in randomized oncology studies used to support marketing approval. It also recommends that OS be assessed as a pre-specified safety endpoint even when it is not the primary efficacy endpoint. The public comment deadline is October 20, 2025.
The status of the document matters. This is draft guidance, not a binding regulation or an automatic requirement for every randomized clinical study. Its recommendations do, however, provide a clear indication of the agency's expectations for study planning and regulatory dialogue. Sponsors should not describe the guidance as a mandate, but they should take its proposed approach seriously when designing randomized oncology trials intended to support approval.
The guidance points toward a dual-track treatment of OS. Overall survival remains the gold-standard definitive efficacy endpoint, while also serving as an important safety assessment when a trial relies on an alternative primary endpoint. This does not mean that every randomized design must be powered to detect an OS benefit or that every trial must use OS as a primary endpoint. It does mean that sponsors should explain, before study initiation, how OS will be collected, followed, analyzed, and interpreted in a study that may generate an approval-supporting efficacy claim.
For a protocol team, the implication is concrete: in randomized oncology studies used to support marketing approval, OS assessment should be built into the design rather than left to an informal data-collection process. The operational plan may need appropriate statistical justification, reliable survival-status follow-up, a defined analysis strategy, and sufficient observation time to interpret mature data. The exact approach will depend on the disease, expected survival, treatment setting, crossover, subsequent therapies, and the intended regulatory pathway.
The regulatory trajectory has been building toward this. The FDA has historically accepted surrogate endpoints reasonably likely to predict clinical benefit under accelerated approval pathways. That pathway remains available. But the agency is increasingly emphasizing that sponsors should not allow an early surrogate result to replace systematic assessment of survival. The 2023 FDA, AACR, and ASA joint public workshop on overall survival laid intellectual groundwork for this discussion; the 2025 draft guidance develops it into a more explicit recommendation for randomized oncology studies supporting approval.
| Dimension | Earlier Common Practice | Direction Reflected in the 2025 Draft Guidance |
|---|---|---|
| OS in surrogate-primary trials | Sometimes collected without a fully specified role in the protocol | Sponsors are encouraged to prespecify OS assessment in randomized oncology studies used to support marketing approval |
| Accelerated approval pathway | A surrogate “reasonably likely” to predict benefit may support the pathway | The pathway remains available, with parallel attention to OS assessment |
| Statistical planning for OS | OS may be treated as exploratory or insufficiently developed | Sponsors should address the OS analysis and its operating characteristics in the protocol and statistical analysis plan |
| Follow-up duration | Follow-up may narrow after the primary endpoint is reached | The protocol should justify follow-up needed to interpret OS and safety meaningfully |
| Safety signal integration | Survival concerns may be addressed through later surveillance | OS is considered from the trial-design stage when an alternative endpoint drives the primary analysis |
This distinction between recommendation and obligation should not be used as a reason to defer planning. Draft guidance often becomes part of the conversation long before it becomes a final document, and reviewers will still expect sponsors to explain why their design is capable of detecting important survival risks. A study that collects OS inconsistently or stops follow-up without a defensible rationale may create avoidable questions even if it technically falls outside the guidance's intended scope.
The right response is not to redesign every trial around OS as the primary endpoint. It is to make the role of OS explicit. Is OS a key secondary endpoint, a safety endpoint, a component of a hierarchical testing strategy, or an outcome that will be followed descriptively? What events are required for interpretation? How will crossover and post-progression therapy affect the analysis? These decisions should be made before database lock, not reconstructed when a reviewer asks for them.
Validation Frameworks: Moving Beyond Individual-Level Surrogacy
The industry has relied on the Buyse et al. criteria, established in 2000, as a structural framework for surrogate endpoint validation. The framework demands two levels of evidence: individual-level surrogacy, which measures the correlation between the surrogate and the clinical endpoint within individual patients, and trial-level surrogacy, which measures the correlation between the treatment effect on the surrogate and the treatment effect on the clinical endpoint across trials.
Here is the operational reality most sponsors underestimate: individual-level surrogacy alone is insufficient. A strong patient-level correlation between PFS and OS does not guarantee that a treatment improving PFS will also improve OS. Trial-level surrogacy — the harder metric to demonstrate — is what actually helps predict whether an efficacy signal will survive regulatory and HTA scrutiny. It requires a meta-analytic evidence base across multiple trials, and for many drug classes and tumor types, that evidence base simply does not exist yet.
The distinction can be illustrated by a simple but consequential scenario. Suppose patients who remain progression-free generally live longer than patients whose disease progresses early. That observation supports individual-level association. It does not establish that a new treatment producing a PFS improvement will produce a corresponding OS improvement. The treatment could delay radiographic progression without changing the later course of disease, or it could alter the availability and effectiveness of subsequent therapy. The patient-level relationship may remain strong while the trial-level treatment effect relationship remains weak.
This creates a direct planning implication for your clinical development strategy. If you are relying on a surrogate endpoint to support accelerated approval, you need to map — during protocol design, not after — the validation status of that surrogate in your specific indication. The questions are precise:
1. Is there published trial-level surrogacy data for this surrogate, in this tumor type, in this line of therapy? If the answer is no, your regulatory and HTA risk profile is materially higher. You may need to compensate through stronger OS planning, longer follow-up, or a more conservative interpretation of the surrogate result.
2. Does the mechanism of action of your compound introduce a dissociation risk between the surrogate and OS? Immune checkpoint inhibitors, for example, can produce delayed survival curves that diverge from early PFS signals. Similar concerns may arise when treatment effects emerge gradually, when pseudoprogression complicates imaging interpretation, or when post-progression treatment strongly influences survival.
3. What is the maturity timeline for OS data in your trial design? If your protocol truncates follow-up before OS data matures, you are creating an evidence gap that may haunt your HTA submission. A primary PFS readout can be timely while the OS follow-up continues, but that requires funding, retention planning, and a clear data strategy.
4. Have you prespecified the statistical methodology for OS analysis, including handling of crossover and subsequent-line therapies? Post-hoc OS analyses carry diminished regulatory weight and are often difficult to interpret. The protocol should identify the estimand, define censoring rules, describe interim analyses, and explain how treatment switching will be addressed.
5. Does the surrogate measure the benefit that matters in this disease setting? A statistically robust endpoint may still be a poor representation of symptom burden, functional status, or treatment burden. The closer the surrogate is to the patient's experience, the easier it is to build a coherent clinical benefit narrative.
6. Will the trial generate evidence that can travel across decision contexts? Regulators, payers, and clinicians may ask different questions of the same dataset. A protocol designed only for an early regulatory milestone may leave too little evidence for comparative effectiveness, quality-of-life assessment, or economic modeling.
Validation is not a one-time label attached to an endpoint. It is a context-dependent argument supported by disease-specific, treatment-specific, and trial-level evidence. The strength of that argument should determine how much weight the protocol places on the surrogate and how aggressively it invests in definitive outcomes.
You do not validate a surrogate endpoint once and deploy it universally. You validate it for a specific compound, in a specific disease, against a specific standard of care. Anything less is an avoidable design risk.
The HTA Disconnect: Reconciling Accelerated Approval with Reimbursement
This is where the strategic bottleneck actually lives for most pharmaceutical organizations, and it is where endpoint selection decisions carry their most consequential financial impact. Regulatory agencies such as the FDA and health technology assessment bodies operate on different evidentiary standards, and your protocol must account for both simultaneously.
The FDA may accept endpoints reasonably likely to predict clinical benefit for accelerated approval and, in some configurations, regular approval pathways. HTA bodies — whether NICE, G-BA, AIFA, or their equivalents — typically demand more. They may want formal surrogacy validation studies, direct overall survival data, stronger comparative evidence, or patient-reported outcomes that support claims about quality of life and utility. A drug that moves through accelerated approval on PFS data can stall in reimbursement negotiations if the sponsor cannot demonstrate a credible link between that PFS improvement and outcomes payers recognize as clinically meaningful.
The gap is not theoretical. It can appear as delays to market access, restricted formulary placement, uncertainty in managed-entry arrangements, and unfavorable cost-per-QALY calculations. An approval based on a surrogate endpoint may establish regulatory availability without resolving the payer's uncertainty about long-term benefit. The organization may celebrate the authorization while the commercial team is still trying to explain how the early endpoint translates into survival, function, symptoms, and resource use.
This is especially challenging when the confirmatory evidence is expected after approval. A post-approval trial may be able to address uncertainty eventually, but the timing, design, and result are not guaranteed. If the original protocol did not preserve patients for long-term follow-up, collect the right PRO instruments, or capture treatment sequences, the sponsor may have limited ability to close the evidence gap later.
The strategic response requires integrated endpoint planning from phase II forward:
- Primary endpoint for regulatory filing: Select the endpoint that offers the strongest available regulatory position, but acknowledge its validation limitations transparently in the development and regulatory strategy. A faster endpoint is not automatically the strongest endpoint if its interpretation depends on assumptions that have not been demonstrated.
- Co-primary or key secondary endpoint: Prespecify OS with a defensible analysis plan and follow-up strategy where the disease and trial objectives justify it. In randomized oncology studies supporting approval, the 2025 draft guidance makes clear that sponsors should plan explicitly for OS assessment even when OS is not the primary efficacy endpoint. Whether OS is powered for hypothesis testing or assessed with another defined role should be explained rather than left ambiguous.
- Patient-reported outcomes: Only 11% of positive surrogate endpoint trials demonstrated quality-of-life improvement. PROs are therefore not decorative additions to the clinical study report. They can clarify whether disease control is accompanied by preserved function, reduced symptoms, or an acceptable treatment burden. Design the PRO collection strategy into the protocol from day one, including instrument selection, assessment timing, missing-data handling, and the estimand relevant to the treatment experience.
- Health resource utilization: If the value dossier will depend on hospitalization, supportive care, treatment administration, or subsequent therapy assumptions, those variables should be captured prospectively where feasible. Retrospective reconstruction is rarely as clean as a protocol-based collection strategy.
- Post-marketing evidence generation: Plan the confirmatory study architecture before approval, not after. The FDA accelerated approval pathway expects confirmatory evidence, and HTA bodies will often demand additional data to resolve uncertainty. The post-approval plan should be connected to the initial trial's endpoint definitions and patient population rather than treated as a separate evidence universe.
- Subgroup and treatment-sequence analysis: Payers may ask whether the observed effect applies to clinically relevant subgroups and how the intervention performs within actual treatment pathways. Prespecified subgroup logic and transparent documentation of subsequent therapies can make later interpretation more credible.
The organizations that treat endpoint strategy as a regulatory-only decision are the ones that achieve approval without achieving market access. That is not a success metric. It is a failure of operational planning.
Strategic Protocol Design: Balancing Speed and Definitive Clinical Benefit
The central tension in oncology protocol design has not changed: faster development against harder evidence requirements. What has changed is the cost of getting the balance wrong, and the specificity of what regulators and payers now expect.
Your protocol design framework must integrate three imperatives simultaneously. Not sequentially. Simultaneously.
Imperative one: Endpoint architecture that survives dual scrutiny. Build your protocol with a primary endpoint that reflects the strongest achievable regulatory position, while making OS assessment a prespecified component of the evidence plan. It should not be an afterthought, an unstructured exploratory analysis, or a post-marketing commitment that the organization hopes to fund later. For randomized oncology studies used to support marketing approval, the August 2025 FDA draft guidance recommends that OS be assessed even when an alternative endpoint drives the primary efficacy analysis.
This requires more than adding an OS row to the schedule of assessments. The protocol should define survival-status collection, death ascertainment, follow-up responsibilities, analysis timepoints, and the relationship between OS and other endpoints. It should also address crossover effects, subsequent-line therapy confounding, treatment discontinuation, and interim OS analyses with methodological rigor.
The statistical analysis plan must match the role assigned to OS. If OS is a key secondary endpoint, the multiplicity strategy and event assumptions should be clear. If it is primarily a safety assessment, the protocol should still explain how important imbalances will be reviewed and what follow-up is needed to make that review meaningful. If the trial is not designed to establish an OS benefit, the document should not imply that it can do so.
Imperative two: Validation evidence mapped before protocol lock. Before your protocol is finalized, you need a complete surrogacy validation landscape for your specific surrogate, indication, and line of therapy. The review should cover the quality of available trial-level evidence, the consistency of findings across prior regimens, the effect of crossover and subsequent treatment, and the relevance of historical trials to the current standard of care.
If trial-level surrogacy data does not exist, your risk mitigation strategy must compensate. Options may include longer follow-up, larger sample sizes for OS assessment, a co-primary endpoint design that includes a definitive clinical endpoint, or a development sequence that allows earlier randomized evidence to inform a later confirmatory design. None of these choices is free. That is precisely why the decision belongs in protocol strategy rather than in the response to a regulatory information request.
Imperative three: HTA-aligned evidence generation baked into trial design. Do not separate your regulatory dossier from your value dossier. The endpoints, timepoints, and analyses that feed your FDA submission should also feed your HTA submission wherever possible. Patient-reported outcomes, health resource utilization data, and mature OS data are not supplementary when the commercial question is whether the treatment produces meaningful value. They are part of the evidence architecture that determines whether an approved drug generates revenue or gathers shelf space.
The practical design exercise is to work backward from the decisions the evidence must support. A regulator may ask whether the treatment has a favorable benefit-risk profile. A payer may ask whether the benefit is large enough, durable enough, and relevant enough to justify the price. A clinician may ask whether the treatment improves outcomes compared with the regimen already used in practice. Patients may care most about symptoms, function, time without deterioration, and the burden of treatment. One endpoint rarely answers all of these questions.
That does not mean every protocol should become an unmanageable collection of endpoints. An overloaded design can dilute statistical power, increase missing data, burden sites, and make the final narrative less coherent. The objective is disciplined integration: select a primary endpoint for a defensible reason, preserve the ability to interpret OS, measure quality of life in a way that can support clinical claims, and define the analyses before the results create pressure to change the story.
A useful internal review before protocol lock should ask:
- What clinical benefit is the surrogate intended to represent?
- What evidence supports that relationship at the patient and trial levels?
- What features of the mechanism or treatment setting could weaken the relationship?
- How will the protocol detect and interpret an OS imbalance?
- What follow-up is needed for a credible survival assessment?
- How will crossover and subsequent therapies affect the estimand?
- Which outcomes will be required for reimbursement discussions that will occur after regulatory review?
- Which data cannot be recovered reliably if it is not collected prospectively?
These are not administrative questions. They determine whether the trial produces an approval-enabling result, a clinically persuasive result, or merely a statistically positive result that leaves the central question unresolved.
The strategic objective is not to reject surrogate endpoints. That would be impractical and, in many disease settings, would slow access to potentially valuable treatments. The objective is to use them with calibrated confidence. A surrogate may be the right primary endpoint when it has strong validation, the disease is rapidly progressive, and the clinical context supports early decision-making. In another setting, the same surrogate may require more conservative claims, stronger OS follow-up, or a different endpoint hierarchy.
The 2025 FDA draft guidance reinforces that distinction. It does not convert OS assessment into a universal binding requirement for every randomized trial, nor does it eliminate the role of surrogate endpoints. It recommends a more deliberate approach for randomized oncology studies used to support marketing approval: assess OS prospectively, define its role, and avoid treating survival as an optional data stream simply because an earlier endpoint reached significance.
Your next oncology protocol should therefore be designed to deliver more than a statistically significant result on a surrogate marker. It should generate an evidence package that explains what the result means, how confidently it predicts definitive benefit, what happened to survival, how patients experienced treatment, and whether the evidence will remain persuasive beyond the regulatory decision. Anything short of that is not necessarily a failed trial. But it is a fragile development strategy — and the data now makes that fragility visible to everyone reviewing the submission.