Safety Indicators and Metrics

Metric Validity Explained: When a High Safety Score Hides Weak Evidence

A safety metric can be accurate and still be unfit for the decision in front of a leader. Metric validity asks whether the measure captures the exposure, behavior, and evidence that should change action before a serious event.

By 3 min read
metrics dashboard representing metric validity explained when a high safety score hides weak evidence — Metric Validity Expla

Key takeaways

  1. 01Metric validity asks whether a safety measure represents the exposure and decision it is meant to guide.
  2. 02A score can be statistically correct while remaining weak evidence if it depends on reporting behavior, hides absolute conditions, or measures activity instead of control.
  3. 03Leaders should test every metric against the decision it changes, the exposure it represents, and the field evidence that can confirm or challenge it.

A safety dashboard can show a strong result while the organization remains unable to see the exposure that matters most. The problem is not always bad arithmetic. It is often a valid calculation attached to the wrong decision.

Metric validity gives leaders a sharper test. It asks whether the measure represents the condition, control, or behavior that should change action before serious harm occurs.

Metric validity is the degree to which a safety measure represents the exposure or control condition it claims to describe and supports the decision it is meant to guide. A valid metric remains useful when reporting behavior, operating conditions, workload, or management attention changes around it.

Definition

A metric is valid when its meaning survives contact with the operation. That requires more than a clean formula because the data may depend on who reports, what supervisors notice, which denominator leaders select, and whether the measure is connected to a real control.

Andreza Araujo has made this distinction central to her safety work. In *Muito Além do Zero*, the warning is practical: an attractive absence-of-accidents result does not prove that serious exposure is under control. The measure must be read alongside the conditions that produce it.

What decision should the metric change?

Start with the decision, not the data source. A board metric may support resource allocation, a plant metric may trigger a control review, and a supervisor metric may require a pause before the next task. If the owner cannot state the action that changes when the value moves, the measure is probably reporting activity rather than guiding risk management.

This is why the leading-indicator calibration routine matters. Calibration is not a cosmetic exercise. It connects the measure to a defined decision and a responsible owner.

Does the metric represent the exposure?

A useful metric has a visible relationship with the exposure. Near-miss counts may reflect reporting trust as much as hazard frequency. Training completion may show attendance without showing competence at the point of work. A low injury rate may describe recorded outcomes while critical controls remain untested.

The Headline Podcast has explored the same issue through the percentile trap described by Dr. Thomas Krause. An organization can rank in the 90th percentile while only 60% of workers trust their supervisor. The ranking is real, but it does not describe the absolute condition leaders need to repair.

Which behavior can distort the number?

Every metric creates attention, and attention changes behavior. When a target rewards a low count, people may report less. When a dashboard rewards completed observations, teams may produce more observations without improving the work. When leaders celebrate a score without asking what workers stopped saying, the measure starts to protect itself.

James Reason's analysis of latent conditions helps here because the visible number can be the final layer of a deeper design problem. A distorted metric is not only a data problem. It can reveal incentives, workload, fear, or weak ownership upstream.

What evidence can confirm the score?

Metric validity improves when the number is paired with field evidence. Compare a reported control result with worksite verification, maintenance history, permit quality, or worker descriptions of the task. The point is not to create another dashboard. It is to test whether the score still means what leaders think it means.

The evacuation-drill analysis shows why a single response-time result can be weak evidence. A faster time may reflect a smaller attendance group, a familiar route, or a staged exercise rather than stronger emergency readiness.

What should a weak metric trigger?

A weak metric should trigger a review of the measure and the operating condition it represents. Replacing the number immediately can hide the reason it failed. Leaders should first identify whether the problem sits in the definition, the data collection, the incentive, the denominator, or the control itself.

This review belongs with the person who owns the decision, not only with the analyst who prepares the report. Across 25+ years of EHS leadership, Andreza Araujo has treated measurement as a management conversation because the value of a metric appears in what leaders change after seeing it.

Recommendation

Before a safety metric reaches the board, write one sentence that names the exposure, the decision, the possible distortion, and the evidence that will challenge the result. If the sentence cannot be written plainly, keep the measure in review rather than treating it as proof.

For Headline Podcast readers, the practical test is simple. Do not ask whether the score is good. Ask what it makes leaders see, what it leaves invisible, and which operating decision changes because of it.

Explore more real safety conversations on Headline Podcast.

Topics safety-indicators-and-metrics metric-validity leading-indicators safety-dashboard safety-leadership

Frequently asked questions

What is metric validity in safety management?
Metric validity is the degree to which a safety measure represents the exposure, control condition, or decision it claims to inform. A valid metric is not merely calculated correctly. It must also be relevant to the risk and useful for changing action.
Can a high safety score be misleading?
Yes. A high score can hide weak evidence when the measure rewards low reporting, uses a denominator that masks absolute exposure, or counts activities without testing whether a critical control works.
How should a leader test a safety metric?
Ask what decision the metric changes, what exposure it represents, which behaviors can distort it, and what field evidence can confirm it. If those answers are unclear, the measure should not carry executive confidence on its own.

About the author

Andreza Araújo

Safety Culture Expert | Senior EHS Executive

Andreza Araújo is a safety culture expert and senior EHS executive with more than 25 years of experience in environment, health and safety. She is a Civil Engineer and Occupational Safety Engineer from Unicamp, holds a Master's degree in Environmental Diplomacy from the University of Geneva, and completed sustainability studies at IMD Switzerland. Andreza has served in Global Head of EHS roles in Fortune 500 environments, leading cultural transformation programs across multinational operations. She has represented Brazil as a speaker at the United Nations in Paris and has spoken at the International Labour Organization in Turin. She is the author of more than 16 books on safety culture in Portuguese, Spanish, English and German. Her work has earned more than 10 EHS awards, including two recognitions from Indra Nooyi, former PepsiCo CEO.

  • Civil & Safety Engineer (Unicamp)
  • M.A. Environmental Diplomacy (University of Geneva)
  • Sustainability Cert (IMD Switzerland)
  • People Management & Coaching (Ohio University)
  • UN Paris speaker representative for Brazil
  • ILO Turin speaker
  • LinkedIn Top Voice
  • Indra Nooyi PepsiCo CEO recognition (2x)

Documentaries

Watch Andreza's documentaries

Three productions on safety culture, organizational failure and the human lessons behind major disasters.

Podcasts

Listen to Andreza's podcasts

She hosts three shows on safety leadership, EHS and organizational culture, in English and Portuguese.

Summarize with AI