Metric Validity Explained: When a High Safety Score Hides Weak Evidence
A safety metric can be accurate and still be unfit for the decision in front of a leader. Metric validity asks whether the measure captures the exposure, behavior, and evidence that should change action before a serious event.

Key takeaways
- 01Metric validity asks whether a safety measure represents the exposure and decision it is meant to guide.
- 02A score can be statistically correct while remaining weak evidence if it depends on reporting behavior, hides absolute conditions, or measures activity instead of control.
- 03Leaders should test every metric against the decision it changes, the exposure it represents, and the field evidence that can confirm or challenge it.
A safety dashboard can show a strong result while the organization remains unable to see the exposure that matters most. The problem is not always bad arithmetic. It is often a valid calculation attached to the wrong decision.
Metric validity gives leaders a sharper test. It asks whether the measure represents the condition, control, or behavior that should change action before serious harm occurs.
Metric validity is the degree to which a safety measure represents the exposure or control condition it claims to describe and supports the decision it is meant to guide. A valid metric remains useful when reporting behavior, operating conditions, workload, or management attention changes around it.
Definition
A metric is valid when its meaning survives contact with the operation. That requires more than a clean formula because the data may depend on who reports, what supervisors notice, which denominator leaders select, and whether the measure is connected to a real control.
Andreza Araujo has made this distinction central to her safety work. In *Muito Além do Zero*, the warning is practical: an attractive absence-of-accidents result does not prove that serious exposure is under control. The measure must be read alongside the conditions that produce it.
What decision should the metric change?
Start with the decision, not the data source. A board metric may support resource allocation, a plant metric may trigger a control review, and a supervisor metric may require a pause before the next task. If the owner cannot state the action that changes when the value moves, the measure is probably reporting activity rather than guiding risk management.
This is why the leading-indicator calibration routine matters. Calibration is not a cosmetic exercise. It connects the measure to a defined decision and a responsible owner.
Does the metric represent the exposure?
A useful metric has a visible relationship with the exposure. Near-miss counts may reflect reporting trust as much as hazard frequency. Training completion may show attendance without showing competence at the point of work. A low injury rate may describe recorded outcomes while critical controls remain untested.
The Headline Podcast has explored the same issue through the percentile trap described by Dr. Thomas Krause. An organization can rank in the 90th percentile while only 60% of workers trust their supervisor. The ranking is real, but it does not describe the absolute condition leaders need to repair.
Which behavior can distort the number?
Every metric creates attention, and attention changes behavior. When a target rewards a low count, people may report less. When a dashboard rewards completed observations, teams may produce more observations without improving the work. When leaders celebrate a score without asking what workers stopped saying, the measure starts to protect itself.
James Reason's analysis of latent conditions helps here because the visible number can be the final layer of a deeper design problem. A distorted metric is not only a data problem. It can reveal incentives, workload, fear, or weak ownership upstream.
What evidence can confirm the score?
Metric validity improves when the number is paired with field evidence. Compare a reported control result with worksite verification, maintenance history, permit quality, or worker descriptions of the task. The point is not to create another dashboard. It is to test whether the score still means what leaders think it means.
The evacuation-drill analysis shows why a single response-time result can be weak evidence. A faster time may reflect a smaller attendance group, a familiar route, or a staged exercise rather than stronger emergency readiness.
What should a weak metric trigger?
A weak metric should trigger a review of the measure and the operating condition it represents. Replacing the number immediately can hide the reason it failed. Leaders should first identify whether the problem sits in the definition, the data collection, the incentive, the denominator, or the control itself.
This review belongs with the person who owns the decision, not only with the analyst who prepares the report. Across 25+ years of EHS leadership, Andreza Araujo has treated measurement as a management conversation because the value of a metric appears in what leaders change after seeing it.
Recommendation
Before a safety metric reaches the board, write one sentence that names the exposure, the decision, the possible distortion, and the evidence that will challenge the result. If the sentence cannot be written plainly, keep the measure in review rather than treating it as proof.
For Headline Podcast readers, the practical test is simple. Do not ask whether the score is good. Ask what it makes leaders see, what it leaves invisible, and which operating decision changes because of it.
Frequently asked questions
What is metric validity in safety management?
Can a high safety score be misleading?
How should a leader test a safety metric?
About the author
Andreza Araújo
Safety Culture Expert | Senior EHS Executive
Andreza Araújo is a safety culture expert and senior EHS executive with more than 25 years of experience in environment, health and safety. She is a Civil Engineer and Occupational Safety Engineer from Unicamp, holds a Master's degree in Environmental Diplomacy from the University of Geneva, and completed sustainability studies at IMD Switzerland. Andreza has served in Global Head of EHS roles in Fortune 500 environments, leading cultural transformation programs across multinational operations. She has represented Brazil as a speaker at the United Nations in Paris and has spoken at the International Labour Organization in Turin. She is the author of more than 16 books on safety culture in Portuguese, Spanish, English and German. Her work has earned more than 10 EHS awards, including two recognitions from Indra Nooyi, former PepsiCo CEO.
- Civil & Safety Engineer (Unicamp)
- M.A. Environmental Diplomacy (University of Geneva)
- Sustainability Cert (IMD Switzerland)
- People Management & Coaching (Ohio University)
- UN Paris speaker representative for Brazil
- ILO Turin speaker
- LinkedIn Top Voice
- Indra Nooyi PepsiCo CEO recognition (2x)
Documentaries
Watch Andreza's documentaries
Three productions on safety culture, organizational failure and the human lessons behind major disasters.
Podcasts
Listen to Andreza's podcasts
She hosts three shows on safety leadership, EHS and organizational culture, in English and Portuguese.