How to Test a Safety Policy at the Point of Work in 21 Days
A safety policy matters only when it changes a real decision under pressure. This F2 guide gives EHS managers and supervisors a 21-day field test for triggers, authority, control evidence, escalation, and restart conditions.

Key takeaways
- 01Test one policy against one observable operational decision instead of reviewing the whole management system.
- 02Define the trigger, first decision owner, backup owner, control evidence, escalation path, and restart condition.
- 03Observe the policy during changed conditions and production pressure, not only during routine work.
- 04Compare how workers, supervisors, and leaders describe the same decision across shifts.
- 05Use the evidence to make one control or authority change, then repeat the test before declaring the policy reliable.
A safety policy can be signed by the chief executive, displayed in every entrance, and still fail when a supervisor must choose between an approved control and a pressured shortcut. The useful question is not whether people can repeat the policy. It is whether the policy changes a real decision when work conditions move away from the plan.
This 21-day field test gives an EHS manager and frontline supervisors a practical way to find that answer. It focuses on observable decisions, evidence, ownership, and the response that follows a concern.
What you need before starting
Choose one policy that is supposed to influence high-consequence work, such as stop-work authority, permit control, contractor management, or energy isolation. Do not test the entire management system at once. A narrow scope makes the result easier to verify, while a real task keeps the exercise connected to operational pressure.
Nominate one EHS owner, one operations owner, and two supervisors from different shifts. Include workers who perform or support the selected task, because a policy that only works in the office has not reached the point of work.
Step 1: Select one decision the policy must change
Write the decision in operational language. For example, a stop-work policy may require a supervisor to pause a task when a critical control is missing or cannot be verified. A useful decision is specific enough that two observers can agree on whether it happened.
Verify the choice by asking three people from different roles to describe the decision without showing them your wording. If their answers describe different triggers, the policy is not ready for testing. The common error is choosing a broad value statement instead of an observable decision.
Step 2: Define the trigger that should start the decision
List the condition that should cause the policy to become active. The trigger may be a changed work boundary, a missing permit condition, an unavailable competent person, or a control whose status is uncertain. Write what a worker can see, hear, or confirm rather than using a phrase such as unsafe situation.
Verify the trigger during a short field walk with a supervisor and a worker. If the answer depends on personal courage or a vague sense of concern, clarify the condition before continuing. The common error is making the trigger so general that it can be applied after the exposure has already occurred.
Step 3: Map the first decision owner and backup owner
Assign the first decision to the role closest to the exposure, provided that role has the authority and competence to control the task. Name a backup owner for nights, weekends, contractor interfaces, and periods when the primary supervisor is absent.
Verify the map with a short night-shift scenario. Record who receives the call, who can pause the work, and who can authorize a restart. The common error is naming EHS as the only owner even though EHS is not present when the decision must be made.
Step 4: Translate the policy into a field question
Create one question that a supervisor can ask before work starts and when conditions change. A strong question asks what has changed, which control is protecting the task, and what evidence confirms that the control is available. It should open a decision, not invite a ceremonial yes.
Verify the question during a pre-task conversation. If it produces only a group affirmation, replace it with a prompt that requires a named control, an owner, and a visible confirmation. The common error is turning the policy into a script that people recite without examining the work.
Step 5: Test the policy against a changed condition
Select a legitimate change that the team could encounter, such as a revised access route, a delayed isolation, a weather shift, or a contractor substitution. Do not create an unsafe exposure for the exercise. Present the changed condition in a controlled discussion or use a real change that has already been safely managed.
Verify whether the supervisor can explain what must stop, what must be reassessed, and who owns the next decision. Ask the worker what evidence would allow the task to restart. The common error is testing only the original plan, because policies are easiest to follow when no uncertainty or production pressure is present.
Step 6: Observe the response without coaching it
For one shift, observe how the policy appears in normal work. Note the trigger, the first response, the people consulted, the control checked, and the point at which work continued or stopped.
Verify the record against the work itself rather than against a completed form. James Reason's work on latent failures is useful here because a visible deviation may reveal a design problem in authority, planning, or control ownership. The common error is treating an imperfect response as an individual attitude problem before checking what the system made difficult.
Step 7: Compare policy language with supervisor language
At the end of the shift, ask the supervisor to explain the decision in their own words. Compare that explanation with the written policy. Look for missing triggers, hidden approval steps, different meanings of pause, and assumptions about who carries the risk while work is delayed.
Verify the comparison with a worker from the same task. The policy is more credible when the worker and supervisor can describe the same decision path. The common error is editing the policy immediately, before separating a wording problem from an authority or resource problem.
Step 8: Test the response after a concern is raised
A policy is not complete when it tells people to speak up. It also has to define what happens after the concern is raised. Record who acknowledges it, who investigates the exposure, who decides the next action, and how the outcome returns to the person who raised the concern.
Verify the response by tracing one concern from first report to closure. Ask whether the person received a decision, whether the control changed, and whether the same issue appeared on another shift. The common error is measuring reporting volume while ignoring whether the response makes future reporting rational.
Step 9: Check whether production pressure changes the rule
Review a decision made near a deadline, during a staffing shortage, or while equipment availability was limited. The purpose is not to blame production. It is to see whether the policy remains usable when the conditions that create risk also create pressure to continue.
Verify what happened to the control, the schedule, and the person who raised the issue. A policy has reached the point of work when leaders can explain which decision changed because of safety evidence, even when the change creates cost or delay. The common error is declaring the policy effective because it works during routine operations.
Step 10: Convert findings into one control change
Choose one change that removes friction from the correct decision. It might be a clearer authority statement, a revised permit field, a backup contact, a better control verification point, or a required management response after a stop. Keep the change narrow enough to implement and visible enough to test.
Verify the change with the same task and roles used in the first observation. If the result depends on one unusually committed supervisor, the design is still fragile. The common error is issuing more communication when the evidence points to a missing resource, unclear role, or weak control.
Step 11: Repeat the test on another shift
Run the decision test with a different supervisor, shift, contractor interface, or operating condition. Variation matters because a policy that works only with the author present is not a reliable operational control. Keep the test question and evidence standard constant while allowing the work context to differ.
Verify whether the second team reaches the same decision and whether the same owner receives the escalation. If the result changes by shift, document the local condition that explains the difference. The common error is forcing identical behavior without fixing different access, staffing, equipment, or leadership conditions.
Step 12: Publish the decision rule and review date
At the end of the 21-day test, publish the short decision rule, named owners, restart evidence, and first review date. Explain what changed because of the test and what remains unresolved. A concise rule is more useful than a new policy document that adds pages but leaves the decision path uncertain.
Verify adoption by asking three roles to explain the rule without opening the document. Schedule a review after the next meaningful change, incident, or high-risk task. The common error is closing the project after publication, even though the policy becomes trustworthy only when later decisions continue to follow it.
How to judge the result
Classify the result as ready, conditional, or not ready. Ready means the trigger, authority, control evidence, escalation path, and restart condition are understood across roles and shifts. Conditional means one supporting condition still needs an owner. Not ready means people cannot agree on when the policy applies or who can act.
Andreza Araujo's Safety Culture: From Theory to Practice treats culture as something visible in decisions, habits, and leadership responses, not as a slogan detached from work. The question is not whether the organization has a good policy. The question is whether the policy changes behavior when the operation becomes inconvenient.
For leaders who want to deepen the diagnosis, Andreza Araujo's safety culture books and guides provide a useful reference for connecting policy, leadership, and field evidence.
Frequently asked questions
What is a safety policy field test? It is a short, controlled review of whether a written safety policy changes an observable operational decision when work conditions, information, or pressure change.
Who should lead the test? An EHS owner should coordinate the evidence, while an operations owner should own the practical decision and resources required to make the policy usable.
How long should the test take? Twenty-one days is long enough to observe more than one shift and response cycle while keeping the scope small enough for a focused improvement.
Should the test measure reporting volume? Reporting volume can provide context, but stronger evidence is whether concerns receive a clear response and whether the decision or control changes when evidence requires it.
What should happen if shifts use the policy differently? Compare work conditions, authority, staffing, equipment, and response paths before rewriting the document. Different outcomes often reveal a design gap rather than a motivation problem.
A policy earns credibility when a worker can trigger the right decision, a supervisor can act without searching for permission, and a leader can show what changed after the concern was raised. Use the 21-day test to make that chain visible before the next high-consequence task depends on it.
Frequently asked questions
What is a safety policy field test?
Who should lead the test?
How long should the test take?
Should the test measure reporting volume?
What should happen if shifts use the policy differently?
About the author
Andreza Araújo
Safety Culture Expert | Senior EHS Executive
Andreza Araújo is a safety culture expert and senior EHS executive with more than 25 years of experience in environment, health and safety. She is a Civil Engineer and Occupational Safety Engineer from Unicamp, holds a Master's degree in Environmental Diplomacy from the University of Geneva, and completed sustainability studies at IMD Switzerland. Andreza has served in Global Head of EHS roles in Fortune 500 environments, leading cultural transformation programs across multinational operations. She has represented Brazil as a speaker at the United Nations in Paris and has spoken at the International Labour Organization in Turin. She is the author of more than 16 books on safety culture in Portuguese, Spanish, English and German. Her work has earned more than 10 EHS awards, including two recognitions from Indra Nooyi, former PepsiCo CEO.
- Civil & Safety Engineer (Unicamp)
- M.A. Environmental Diplomacy (University of Geneva)
- Sustainability Cert (IMD Switzerland)
- People Management & Coaching (Ohio University)
- UN Paris speaker representative for Brazil
- ILO Turin speaker
- LinkedIn Top Voice
- Indra Nooyi PepsiCo CEO recognition (2x)
Documentaries
Watch Andreza's documentaries
Three productions on safety culture, organizational failure and the human lessons behind major disasters.
Podcasts
Listen to Andreza's podcasts
She hosts three shows on safety leadership, EHS and organizational culture, in English and Portuguese.