
The risk, inconveniently, may remain exactly where it was.
That is the uncomfortable direction of travel in pharmacovigilance. Regulators are no longer interested only in whether a risk minimisation measure was deployed. They want evidence that it reached the right people, changed what those people knew or did, and contributed to safer use in practice. The difference is substantial; it is also where many programmes quietly lose their credibility.
The revised EMA Good Pharmacovigilance Practice Module XVI and its Addendum II took legal effect on August 6, 2024. The revision sharpens the classification of risk minimisation measures, separates messages from tools, and places greater weight on structured evaluation, stakeholder involvement, and the distinction between process indicators and outcome indicators.
This is not a cosmetic rewrite. It is a direct challenge to the familiar industry habit of confusing dissemination with effectiveness.
The shift in GVP Module XVI: classification is no longer a filing exercise
The revised GVP Module XVI categorises risk minimisation measures into two broad components: RMM messages and RMM tools. The tools are then divided into educational or safety advice tools and risk minimisation control tools.
That structure matters because it forces a more disciplined question: what is the measure actually designed to do?
A message may communicate a contraindication, a monitoring requirement, a dosing restriction, or a warning about a specific adverse reaction. A tool may operationalise that message through educational material, a prescriber guide, a patient card, a controlled access process, or another intervention intended to shape clinical behaviour.
The old temptation was to treat all of these as interchangeable pieces of the same regulatory package. They are not. A safety message can be accurate and still fail to influence practice. An educational tool can be distributed widely and still fail to reach the clinicians who prescribe the product. A control measure can be technically present while creating so much friction in the clinical workflow that users bypass it, misunderstand it, or treat it as administrative noise.
The revised framework makes this distinction harder to ignore. It also supports a more realistic view of the safety system: a risk minimisation programme is not a document. It is a chain of actions involving regulators, marketing authorisation holders, healthcare professionals, patients, pharmacists, and sometimes health-system infrastructure.
Break one link and the programme may continue to generate excellent-looking records.
A measure is not effective because it exists, and it is not ineffective because it is imperfect. The question is whether it changes risk-relevant behaviour in the real system.
For marketing authorisation holders, the practical implication is that the evaluation plan should be designed alongside the intervention itself. If the team cannot explain how a measure is expected to influence prescribing, dispensing, monitoring, administration, or patient behaviour, it will struggle to define meaningful indicators later.
That is the point at which many programmes discover they have been measuring visibility instead of safety.
Why past PASS studies struggled to prove effectiveness
The evidence base is not especially flattering. A review of 93 Post-Authorization Safety Studies assessed by the EMA Pharmacovigilance Risk Assessment Committee between 2016 and 2021 found that 39.8% did not reach a conclusion on the effectiveness of the risk minimisation measures under evaluation.
Nearly four in ten studies failed to provide a conclusion. That is not a minor methodological wrinkle; it is a signal that the evaluation model itself often begins too late, asks the wrong question, or relies on evidence that cannot support the conclusion being sought.
The same review found that 62.4% of the studies focused on healthcare professionals’ awareness, knowledge, and behaviour. Those are legitimate areas of inquiry, but they are not interchangeable.
Awareness means a healthcare professional has encountered or recognises the safety information. Knowledge means the professional can recall or understand it. Behaviour means the professional acts in accordance with it. Clinical outcome means the measure contributes to reducing the relevant risk. Each sits at a different level of the causal pathway.
The industry frequently stops at the first or second level because those measures are easier to collect. A survey can ask whether a clinician has seen a brochure. It can test recall of a contraindication. It can assess whether a monitoring recommendation is recognised. These are useful signals, but they do not, on their own, establish that patients received the appropriate monitoring or that the adverse event rate changed.
This is where corporate optimism enters the room wearing a tidy methodology section.
A high awareness score can coexist with poor implementation. A clinician may know that laboratory monitoring is recommended but lack access to testing, time, system prompts, or a clear workflow. A patient may receive a safety card but not understand when to act on it. A prescriber may recall the warning but still make a different decision when faced with a complex patient and a crowded clinic.
The evaluation must therefore preserve the distinction between:
- Reach: whether the intended audience received or encountered the measure.
- Comprehension: whether users understood the message and its relevance.
- Adoption: whether the expected action occurred.
- Implementation quality: whether the measure functioned in the clinical environment as designed.
- Risk-related outcomes: whether the intervention contributed to safer use or reduced occurrence of the targeted harm.
A programme that measures only reach has established distribution. It has not established effectiveness.
The cross-sectional study problem
Cross-sectional designs were the most frequently used approach in the reviewed PASS: 77.4% of studies used them, compared with 29.0% using cohort designs. Cross-sectional studies can be efficient and appropriate for assessing awareness or knowledge at a defined point in time. They are less powerful when the claim concerns sustained behaviour, temporal change, or clinical risk reduction.
The problem is not that cross-sectional research is inherently weak. The problem is methodological overreach. A snapshot cannot easily show whether behaviour changed after the intervention, whether the change persisted, or whether the observed pattern differs from what would have happened without the measure.
A well-designed cross-sectional study can still contribute valuable evidence, especially when it is linked to a clear evaluation question and supported by other data. But it should not be asked to carry the entire burden of proof.
For risk minimisation measures effectiveness evaluation, the study design should follow the causal question—not the availability of a convenient survey panel.
Process indicators and outcome indicators: two sides of the same programme
The revised GVP Module XVI places greater emphasis on structured evaluation using process and outcome indicators. The distinction sounds straightforward; in practice, it is where evaluation plans either become useful or collapse into a list of activity metrics.
Process indicators show whether the system operated
Process indicators examine implementation. They can address questions such as:
- Were the educational materials distributed to the intended healthcare professionals?
- Were materials available in the correct language and format?
- Did the programme reach the settings where the product is prescribed or administered?
- Were prescribers, pharmacists, and patients exposed to the relevant information?
- Did the control measure function without unnecessary interruptions or workarounds?
- Were communications delivered within the required timeframe?
- Did users report practical barriers to applying the measure?
These indicators are not trivial. If a programme never reached the target audience, there is little value in debating its theoretical effectiveness. Distribution gaps, low participation, missing translations, inaccessible formats, and poor integration into clinical workflows can all explain why an intervention failed.
But process data cannot substitute for outcomes. A company can report that thousands of educational materials were distributed while remaining unable to show that anyone changed their practice.
Outcome indicators test whether the measure made a difference
Outcome indicators go further. They assess whether the intended change occurred. Depending on the risk and intervention, this may involve:
- Correct prescribing in the population for whom the risk is most relevant.
- Appropriate patient selection or exclusion.
- Compliance with recommended laboratory or clinical monitoring.
- Correct dosing, duration, or administration procedure.
- Reduction in contraindicated co-medication.
- Reduction in preventable exposure to the identified risk.
- Changes in the frequency or severity of the adverse event of concern.
The correct outcome indicator depends on the risk minimisation objective. There is no universal metric, despite the industry’s occasional fondness for universal templates. A patient card designed to prompt urgent action should not be evaluated using the same logic as a controlled distribution system. A prescriber guide addressing a complex monitoring requirement needs a different assessment from a warning about a simple dosing restriction.
The causal pathway should be explicit:
| Evaluation layer | Core question | Typical evidence |
|---|---|---|
| Reach | Did the measure reach the intended users? | Distribution records, delivery data, participation rates |
| Understanding | Did users understand the safety message? | Knowledge assessments, user testing, structured interviews |
| Behaviour | Did users take the recommended action? | Prescribing audits, monitoring records, dispensing or administration data |
| Implementation | Did the measure work within the clinical workflow? | Process review, barrier analysis, user feedback |
| Outcome | Did the targeted safety risk change? | Safety surveillance, utilisation data, clinical or epidemiological outcomes |
The table is not a regulatory template. It is a defence against category errors.
If a programme reports high reach and improved knowledge, the conclusion should reflect that evidence. It should not quietly upgrade those findings into proof of clinical effectiveness. A survey can tell us that a clinician knows the recommendation. It cannot, by itself, tell us that the recommendation was followed consistently across the relevant patient population.
The most dangerous sentence in an RMM report is often the one that turns awareness into effectiveness without showing the missing steps.
The 12–18 month window: early review is not premature
Guidance indicates that the initial evaluation of a risk minimisation programme’s effectiveness should typically occur within 12 to 18 months after implementation. The timing has a practical logic: it allows enough exposure for the programme to operate while preserving the opportunity to amend it before an ineffective measure becomes institutional furniture.
That window should not be treated as a ceremonial milestone. It is a decision point.
By 12 to 18 months, the evaluation team should be able to determine whether the measure:
1. Reached the intended audience across the relevant healthcare settings.
2. Produced the expected level of awareness and understanding.
3. Was feasible within actual prescribing, dispensing, monitoring, or administration workflows.
4. Generated evidence of the intended behavioural change.
5. Shows any meaningful relationship with the safety outcome it was designed to address.
6. Requires modification, reinforcement, replacement, or escalation.
The final point is the one that tends to disappear into committee language. Evaluation is supposed to enable iteration. If a measure does not work, the answer is not to write a more confident conclusion about the same data. The answer is to improve the measure, change the channel, remove friction, redesign the tool, or reconsider whether the intervention is proportionate to the risk.
Timing must reflect the risk and the mechanism
The 12–18 month expectation is a typical window, not a substitute for clinical judgement. Some risks may emerge quickly and require earlier monitoring. Others may depend on longer exposure, seasonal patterns, cumulative treatment, or rare outcomes that demand a longer observation period.
The evaluation plan should therefore explain:
- When the intervention became operational in each relevant market or setting.
- How long users would reasonably need to encounter and adopt it.
- When the targeted behaviour should become measurable.
- When the safety outcome could plausibly change.
- Which interim signals would trigger an earlier review.
There is a difference between allowing a measure time to work and allowing an ineffective measure to hide behind insufficient follow-up. The former is scientific caution; the latter is governance by delay.
A credible plan should also predefine how inconclusive findings will be handled. The fact that a study does not demonstrate effectiveness does not necessarily prove that the measure failed. But it may show that the design, data source, implementation, or endpoint was inadequate. An inconclusive result should prompt diagnostic analysis—not an automatic press release about reassurance.
Stakeholder integration: patients and healthcare professionals are part of the design
The revised module explicitly emphasises direct involvement of healthcare professionals and patient representatives in the design, user-testing, dissemination, and evaluation of risk minimisation measures.
This is more important than it may sound. Too many safety tools are designed in an echo chamber populated by regulatory specialists, medical writers, and compliance reviewers. The resulting material is accurate, comprehensive, and nearly impossible to use at the point of care. It satisfies internal review while asking the end user to perform a small administrative pilgrimage before making a clinical decision.
Healthcare professionals can identify where a measure collides with workflow. They know whether a monitoring requirement is visible at the point of prescribing, whether the information appears in the systems they actually use, and whether the proposed action is realistic under routine conditions. They can also expose the difference between a recommendation that sounds clear in a conference room and one that survives contact with a busy clinic.
Patients contribute a different form of expertise. They can identify ambiguous language, hidden assumptions, inaccessible formats, and the practical circumstances in which a warning will be ignored or misunderstood. Patient involvement is not a decorative gesture added to demonstrate modernity. It is a method for testing whether the intervention makes sense to the person who may need to act on it.
User-testing should happen before approval, not after disappointment
User-testing is most valuable before the tool becomes fixed. It should examine more than whether participants prefer one colour palette over another. The relevant questions include:
- Can the intended user identify the risk quickly?
- Do they understand what action is required?
- Can they distinguish urgent action from routine advice?
- Do they know when and where to seek help?
- Does the tool fit the decisions users actually make?
- Does the language work for people with different levels of health literacy?
- Does the measure create confusion with other product information?
- Can the required action be completed in the real care pathway?
A polished document that fails these tests is not a successful communication product. It is a well-formatted obstacle.
Stakeholder involvement should continue during effectiveness evaluation. Users can help explain why a measure performed poorly. If awareness is high but behaviour remains unchanged, the barrier may not be knowledge. It may be access, time, competing priorities, unclear responsibility, lack of integration into electronic systems, or a control measure that creates unacceptable operational friction.
Without that context, an evaluation can identify the gap but not the mechanism behind it.
What a stronger RMM evaluation framework looks like
A robust evaluation framework does not begin with the available dataset. It begins with the safety problem.
The team should define the chain from identified risk to desired change, then decide what evidence would demonstrate each meaningful step. That approach is less convenient than reusing a familiar survey instrument, but pharmacovigilance has never been improved by convenience pretending to be methodology.
A practical framework usually includes the following sequence:
1. Define the risk in operational terms.
Specify the harmful exposure, event, population, treatment context, or clinical behaviour the measure is intended to address. Broad descriptions produce broad, untestable endpoints.
2. State the intended behaviour.
Identify what the prescriber, pharmacist, patient, nurse, or other user should do differently. If the expected action cannot be stated clearly, the measure is not ready for evaluation.
3. Separate the indicators.
Map reach, understanding, behaviour, implementation, and outcomes as distinct evidence layers. Do not allow a strong result in one layer to conceal missing evidence in another.
4. Choose data sources that match the question.
Surveys may assess awareness and knowledge. Utilisation data may examine prescribing patterns. Medical records may support monitoring assessments. Spontaneous reports and other safety data may contribute to outcome analysis, although they require careful interpretation.
5. Include a route to explanation.
Quantitative findings should be supported by qualitative investigation where the result is unexpected. Numbers can show that a gap exists; users often explain why.
6. Predefine decision rules.
Establish what findings would lead to continuation, modification, additional communication, redesign, or escalation. Otherwise, the evaluation becomes an observational ritual with no operational consequence.
7. Plan the reassessment.
Risk minimisation is iterative. A measure should have a defined path for review after changes, especially when the first evaluation identifies poor reach, weak comprehension, or limited behavioural adoption.
This is not about building an unnecessarily elaborate study for every routine safety update. The available facts do not support a claim that standard risk minimisation measures always require a full observational PASS. Regulatory expectations remain proportionate to the nature of the risk, the intervention, and the specific request.
The point is more basic: where additional risk minimisation measures require evaluation, the evidence should be capable of answering the question regulators have actually asked.
The regulatory expectation is evidence of performance, not evidence of activity
The revised GVP Module XVI does not eliminate uncertainty. It does, however, make it harder to disguise activity as impact.
Marketing authorisation holders should expect greater scrutiny of how measures are classified, how users were involved, how indicators were selected, and whether the evidence supports the language used in the conclusion. A programme built around distribution logs and an awareness survey may still have a role, but it should not be presented as a complete assessment of risk reduction.
The strategic mistake is to treat evaluation as a report that arrives after implementation. The better approach is to treat it as part of the intervention’s architecture. Design the measure around the behaviour it must change. Test the tool with the people who will use it. Establish the data pathway before the first package is distributed. Review performance within the 12–18 month window. Then modify the programme when the evidence demands it.
That may sound less impressive than declaring alignment across stakeholders and optimising the communication ecosystem. It is also far more useful.
I have watched enough pharmacovigilance programmes pass internal review because every document was present, every stakeholder had attended the meeting, and every action had been recorded. None of that guarantees that a patient was safer. The new expectation is a welcome correction: show how the measure worked, where it failed, and what you changed when reality refused to follow the slide deck.
The actionable takeaway is simple. For every risk minimisation measure, ask what changed in practice—not what was sent, uploaded, approved, or acknowledged. If the evaluation cannot answer that question, the programme has produced optics and documentation; it has not yet produced evidence of effectiveness.