Safety Indicators and Metrics

Safety Metrics: 3 Tests That Separate Activity From Control Assurance

A safety metric becomes useful when it helps a leader decide whether exposure is controlled. This article presents three tests that separate activity counts from evidence of control assurance.

By 7 min read updated
metrics dashboard representing safety metrics 3 tests that separate activity from control assurance — Safety Metrics: 3 Tests

Key takeaways

  1. 01A completed safety activity does not prove that a critical control is present, usable, and effective under pressure.
  2. 02Test every important metric for decision relevance, field evidence, and consequence when the control is absent.
  3. 03Separate inspections, training, and observations from evidence that the exposure is controlled.
  4. 04Use metrics to trigger a named management decision rather than reward a cleaner report.
  5. 05A smaller dashboard with verified control evidence is more useful than a large dashboard that only describes effort.

A senior leader opens the monthly safety dashboard and sees a reassuring picture. Training completion is above target, inspections are closed on time, and corrective actions have declined. The next morning, a critical task begins with a barrier that is available in the procedure but missing from the work area.

The dashboard did not necessarily contain false data. It contained data that answered an easier question than the one the operation needed to answer. It showed activity, while the decision concerned control assurance. That distinction matters because a team can become highly efficient at recording prevention work without proving that the exposure is actually controlled.

Across more than 250 cultural transformation projects, Andreza Araujo has emphasized the gap between declared conformity and the conditions people face during real work. Her book Safety Culture: From Theory to Practice is useful here because it treats culture as something visible in decisions, priorities, and operating conditions, not only in statements or completion rates.

Why activity counts are not enough

Activity metrics are attractive because they are easy to define, collect, compare, and display. A manager can report how many toolbox talks occurred or how many actions were closed without opening a difficult discussion about whether the control changed the task.

The problem begins when the proxy becomes the objective. A dashboard can also conceal risk-perception distortions that make familiar exposure look ordinary, so the metric needs to be read alongside field evidence. Once closure, attendance, or inspection volume is treated as proof of prevention, teams learn to protect the number. They may split one meaningful review into several records, close an action with a document update, or repeat a familiar observation because it is easier to report than a deeper control weakness.

This does not make activity metrics useless. It means they need a clear relationship with the exposure they are meant to influence. If that relationship is missing, the dashboard reports effort but cannot support a serious risk decision.

Test 1: Does the metric connect to a defined exposure?

The first test asks whether the metric is attached to a specific exposure, task, population, or control. “Safety observations completed” is not an exposure. It is a volume. “Percentage of critical lifting tasks in which the exclusion zone was established and maintained” is closer to a control question because it identifies the work and the condition that matters.

A useful metric lets the reader answer four questions without guessing. Which harm or high-consequence event is relevant? Where can the exposure occur? Which control is expected to prevent or limit it? Who has authority to change the condition when the control is weak?

When the metric cannot answer those questions, the organization should not present it as evidence of prevention. It may remain an administrative measure, but its label and interpretation need to stay modest.

James Reason’s work on latent failures reinforces this discipline. The visible action at the end of a chain is rarely the whole explanation for an unsafe outcome. A metric that focuses only on what was recorded after the task can miss the design, authorization, maintenance, staffing, and supervision conditions that allowed the exposure to remain.

Test 2: Is there field evidence behind the number?

A metric becomes more credible when its result can be checked against the work. This does not require every data point to come from a long audit. It requires a defined sample, a clear verification method, and enough contact with the field to reveal when the record and the task diverge.

Suppose a dashboard reports that 96 percent of permit reviews were completed before work started. The number says little about control strength unless the review also checks whether the permit matched the task, whether isolations were confirmed, whether the people doing the work understood the boundaries, and whether the permit remained valid when conditions changed.

Field evidence can include a direct observation, a worker explanation, a supervisor decision record, a photograph tied to the control, a maintenance status, or a test of whether the barrier remains usable during the difficult part of the task. The method should fit the control. A signature is not equivalent to a functional test.

In The Illusion of Compliance, Andreza Araujo examines how documents can create the appearance of control while the operating condition remains unchanged. If the evidence is only another record generated by the same process, the measurement may be measuring paperwork consistency rather than protection.

Test 3: Does the metric trigger a decision when the control is weak?

A metric earns its place on an executive dashboard when a threshold, trend, or exception leads to a defined decision. Without that link, the number may create discussion without creating action.

For each important measure, write the response before publishing the result. If critical controls are not verified, who can pause the task? If the same weakness appears across sites, who can fund a design change? If a contractor cannot demonstrate the required barrier, who can change the scope or authorization? If the metric worsens after a production change, which leader must review the decision?

This step exposes a common weakness in reporting systems. The organization may have a metric owner who prepares the chart, but no decision owner who controls the condition. In that case, the dashboard can identify a problem without giving anyone the authority to solve it.

Andreza’s Antifragile Leadership provides a useful leadership lens for this problem. A strong leader does not treat bad information as a failure of loyalty. The leader uses it to improve the quality of the decision system, especially when the first response to an unfavorable signal would otherwise be to explain it away.

What activity metrics still do well

Activity measures can reveal whether a management system is functioning. Training attendance can show whether access was provided. Inspection completion can show whether a review process is operating. Corrective-action aging can show whether decisions are being delayed. These are legitimate questions, provided the answer is not stretched beyond the evidence.

The right approach is to pair activity with the control result. Report the number of inspections, then show the proportion that found a meaningful control weakness and the time taken to correct it. Report training completion, then test whether workers can recognize the exposure and apply the critical control in the task. Report closed actions, then verify whether the field condition changed.

This pairing keeps activity visible without allowing it to impersonate assurance. A high completion rate can coexist with a weak control result, which is precisely when the system needs attention.

How to build a decision-linked metric

Start with the decision, not the available data. Ask what a supervisor, plant manager, or board member needs to decide before the exposure becomes an event. Then define the control condition that would support that decision and the evidence that can confirm it.

A practical metric specification can use five fields. Name the exposure, state the required control, define the evidence, identify the decision owner, and set the response when the result is outside the acceptable condition. The formula is secondary. A precise formula cannot rescue a weak control definition.

For example, a confined-space readiness measure might test whether the entry authorization, atmospheric testing, rescue arrangement, communication method, and trained roles are ready for the actual entry. The result should not be a generic readiness percentage if one missing rescue element makes the entry unacceptable. The metric must preserve the decision logic of the control.

That is where many dashboards lose meaning. They average conditions that should not be averaged, turning one decisive failure into a tolerable decimal. A leader should know which missing element requires a stop, not only the overall score.

How leaders should read an improving trend

An improving trend deserves curiosity before celebration. Ask whether the work changed, whether the exposure changed, whether the control became easier to use, and whether the evidence comes from different shifts, teams, and locations.

If the people who collect the data are judged mainly on completion, they may have an incentive to reduce the visibility of difficult conditions. This is not a character judgment. It is a predictable effect of a measurement system that rewards appearance more than evidence.

Select several reported successes and compare the record with the task. Speak with the people who execute the control. Review whether the condition holds during changeover, maintenance, night work, contractor activity, or production pressure. If the improvement disappears in those moments, the metric has identified administrative progress rather than reliable protection.

What to remove from the executive dashboard

Remove measures that have no decision owner, no defined exposure, no field verification, or no response when the result worsens. A metric can be popular and still be operationally empty.

Remove duplicate counts that describe the same activity in different language. Remove averages that hide a single unacceptable control failure. Remove targets that encourage reporting volume while discouraging difficult escalation. Replace them with measures that show where exposure exists, which barrier is weak, and what leadership decision is pending.

The goal is not a dashboard that looks alarming. The goal is a dashboard that makes the next responsible decision harder to avoid. That standard connects the number to the work, the control, and the person who can change the condition.

Control assurance is a management practice, not a reporting style

Safety metrics become valuable when they help an organization see whether protection exists before harm occurs. Ask whether the metric connects to a defined exposure, whether field evidence supports the number, and whether the result triggers a named decision when the control is weak.

If any answer is no, treat the measure as activity data rather than proof of assurance. Leaders who make that distinction protect the integrity of the dashboard and the people who rely on it. They also move the conversation toward whether the work is safer because the organization changed something that matters.

Listen to Headline Podcast conversations about safety leadership, evidence, and better risk decisions.

Topics safety-indicators-and-metrics control-assurance leading-indicators critical-controls safety-dashboard ehs-leadership risk-decisions headline-podcast

Frequently asked questions

What is the difference between a safety activity metric and a control assurance metric?
An activity metric counts what the organization did, such as inspections completed or people trained. A control assurance metric tests whether a specific protection exists, works in the real task, has an owner, and is restored when it fails. Activity can support assurance, but it cannot replace evidence that the exposure is controlled.
Why do safety dashboards often look better than field conditions?
Dashboards often collect data that is easy to count, while field conditions require observation, verification, and a decision about whether the control is strong enough. When a reporting process rewards closure more than evidence, the dashboard can show high completion even though the exposure, usability, or ownership problem remains.
How many safety metrics should an executive dashboard contain?
There is no universal number. The dashboard should contain the smallest set that supports the decisions leaders must make about serious exposure, control ownership, resources, and unresolved weakness.
What should leaders ask when a leading indicator improves?
Leaders should ask what changed in the work, which control became stronger, what evidence proves the change, and whether the improvement appears across shifts and locations. If the only change is a higher completion rate, the indicator may be measuring reporting discipline rather than risk reduction.
Can lagging indicators be removed from safety reporting?
Lagging indicators can still provide context about harm that occurred, but they should not carry the full burden of proving prevention. Leaders need exposure and control evidence because low injury counts may coexist with weak barriers or limited reporting.

About the author

Andreza Araújo

Safety Culture Expert | Senior EHS Executive

Andreza Araújo is a safety culture expert and senior EHS executive with more than 25 years of experience in environment, health and safety. She is a Civil Engineer and Occupational Safety Engineer from Unicamp, holds a Master's degree in Environmental Diplomacy from the University of Geneva, and completed sustainability studies at IMD Switzerland. Andreza has served in Global Head of EHS roles in Fortune 500 environments, leading cultural transformation programs across multinational operations. She has represented Brazil as a speaker at the United Nations in Paris and has spoken at the International Labour Organization in Turin. She is the author of more than 16 books on safety culture in Portuguese, Spanish, English and German. Her work has earned more than 10 EHS awards, including two recognitions from Indra Nooyi, former PepsiCo CEO.

  • Civil & Safety Engineer (Unicamp)
  • M.A. Environmental Diplomacy (University of Geneva)
  • Sustainability Cert (IMD Switzerland)
  • People Management & Coaching (Ohio University)
  • UN Paris speaker representative for Brazil
  • ILO Turin speaker
  • LinkedIn Top Voice
  • Indra Nooyi PepsiCo CEO recognition (2x)

Documentaries

Watch Andreza's documentaries

Three productions on safety culture, organizational failure and the human lessons behind major disasters.

Podcasts

Listen to Andreza's podcasts

She hosts three shows on safety leadership, EHS and organizational culture, in English and Portuguese.

Summarize with AI