Safety Culture Evidence: Survey vs Field Observation vs Decision Audit, Which Signal Should Leaders Trust?
Safety culture surveys, field observations, and decision audits reveal different parts of organizational reality. This comparison helps leaders choose the evidence source that fits the decision, then test it against the conditions and choices shaping work.

Key takeaways
- 01A survey reveals how people experience priorities, pressure, and voice, but it does not prove that controls work in the field.
- 02Field observation shows work-as-performed and the conditions that make a procedure reliable or fragile.
- 03A decision audit traces authority, exceptions, escalation, and ownership when safety becomes inconvenient.
- 04Disagreement between the three sources is useful evidence because it reveals where stated culture fails to reach operating decisions.
- 05Leaders should choose the lead method by the decision they need to improve, then test its conclusion with the other two methods.
A safety survey can report strong trust while a supervisor quietly cancels a control check to protect output. A field observation can show correct behavior while the same operation approves temporary changes without a clear owner. A decision audit can expose that contradiction, yet it can also miss the daily conditions that shape how people work. Leaders who treat one source as the whole truth are not measuring safety culture. They are measuring the part of culture that their chosen method can see.
Survey data, field observation, and decision audit each answer a different question. A survey reveals how people perceive priorities, pressure, and voice. Field observation shows how work is performed under real conditions. A decision audit examines whether choices, exceptions, and escalations preserve the controls leaders say matter. The strongest diagnosis does not average these sources into one score. It compares their disagreements and asks what operating condition produced them.
Across more than 250 cultural transformation projects, Andreza Araujo has treated culture as an operating system rather than a communication campaign. That perspective changes the executive question from “What is our culture score?” to “Which decisions become easier or harder because of the culture we have?” The answer determines which evidence should lead the next intervention.
Evaluation criteria for safety culture evidence
The three methods should be evaluated against the decision they need to support. If a board is deciding whether the organization can absorb rapid growth without weakening controls, the evidence must show more than employee sentiment. If a supervisor is trying to understand why a procedure is bypassed, a survey score alone will not reveal the friction in the task. If the organization is preparing for a major change, the evidence must show how authority, escalation, and exceptions work when the plan no longer matches the field.
Five criteria make the comparison practical. First, consider proximity to the work, because evidence gathered far from the task can hide local pressure. Second, examine the time horizon, since a survey may describe a broad climate while an observation captures one shift. Third, test behavioral specificity, which shows whether the method identifies a decision or only a general impression. Fourth, assess response bias, including fear, optimism, and the desire to present the operation well. Fifth, ask whether the evidence points to an owner who can change the condition.
| Criterion | Survey | Field observation | Decision audit |
|---|---|---|---|
| Primary question | How do people experience priorities, pressure, and voice? | What happens during real work under current conditions? | How do leaders and teams make, escalate, and review safety-critical choices? |
| Best strength | Reach across locations and roles | Visibility into work-as-performed and local workarounds | Traceability from risk recognition to ownership and action |
| Typical blind spot | General sentiment without a visible operating example | Behavior altered by observation or limited to one moment | Paper decisions that are not connected to the worker experience |
| Useful output | Patterns by role, site, shift, or issue | Specific friction, exposure, and control conditions | Decision pathway, exception pattern, and governance gap |
The most important criterion is not statistical precision. It is decision usefulness. A method deserves priority when it can change the next decision, identify the person who can change it, and produce evidence that the change reached the point of work.
Survey data reveals the experienced culture
A well-designed survey is valuable because it gives people a route to describe conditions that leaders cannot continuously observe. It can show whether employees believe production pressure overrides safety, whether supervisors invite bad news, and whether people understand what happens after they raise a concern. When the sample covers different roles and shifts, the pattern can reveal that the same policy is experienced differently by maintenance, operations, contractors, and frontline supervisors.
The survey becomes weak when leaders mistake agreement for proof. Employees may agree that safety matters because the statement is socially expected, while their answers about workload, stop-work confidence, or response to bad news reveal a different operating reality. A high score on values is therefore less useful than a specific question about what happens when a target and a control compete.
Survey design also determines whether the results can guide action. Questions should distinguish personal belief from observed management behavior. “My manager cares about safety” is broad. “My manager has delayed work when the required control was unavailable” points toward a decision that can be verified. The second question is harder to answer, yet it creates a bridge between perception and evidence.
Leaders should use survey data to select where to look next, not to close the investigation. If maintenance reports low confidence in isolation planning while production reports strong safety commitment, the disagreement is not noise. It may indicate that the culture is supportive in conversation but unreliable at the boundary between planning and execution.
Field observation reveals the operating culture
Field observation is the closest of the three methods to the conditions in which exposure is created or controlled. It can show whether a pre-job discussion changes the work sequence, whether a supervisor notices a missing barrier, and whether the crew can pause without waiting for permission from a distant manager. It also reveals the small adaptations that formal systems often ignore, including shortcuts that have become normal because the designed process is too slow or incomplete.
Observation is not a search for individual fault. James Reason’s work on latent conditions remains useful here because repeated behavior often points to a system that makes the unsafe choice easier, faster, or more rewarded than the safe one. A worker who crosses a temporary boundary may be responding to a layout problem, an unavailable tool, a schedule constraint, or a permit that does not match the task. The observation should capture the condition around the action.
The method has limits. A planned walk can produce performance for the visitor rather than the normal work pattern. A single observation can also overrepresent an unusual shift. To reduce that distortion, leaders should observe the task at different points in its cycle, ask what changes when the supervisor leaves, and compare the work with the decision and control requirements established before the task began.
Field observation should end with a testable statement. “The team needs to be more engaged” is not a finding. “The crew cannot verify the temporary isolation boundary because the permit identifies the equipment but not the isolation owner” is a finding that engineering, operations, and the permit issuer can act on.
Decision audits reveal the governance culture
A decision audit follows a safety-relevant choice through its evidence, authority, exception, and review. It asks what information was available, who had the right to decide, which control was accepted, and how the organization responded when conditions changed. This method is especially useful for temporary changes, contractor interfaces, overdue actions, staffing decisions, and work that continues after a control is found unavailable.
The audit is different from checking whether a form was completed. A completed form can show that a process exists while hiding the decision that mattered. The audit should reconstruct the sequence. When a task continued without the planned safeguard, did the supervisor understand the exposure, did the team know who could stop the work, and did the manager receive a clear escalation before accepting the remaining risk?
Decision audits are powerful because they expose ownership. They often show that a problem has been described by several functions but owned by none. A risk register may name an action, a meeting may repeat the concern, and a dashboard may show the item as open, yet nobody has authority to change the design, schedule, staffing, or budget that keeps the exposure in place.
The method can become sterile when it examines only approved decisions. Include exceptions, reversals, and decisions made under pressure. Those moments reveal whether the organization treats a control as a genuine boundary or as a preference that can be traded away without a visible conversation.
Why the three signals disagree
Disagreement is often the most useful result. A favorable survey with weak field practice may indicate that leaders communicate the right values but have not removed the friction that shapes behavior. Strong field practice with poor survey trust may indicate competent crews working inside a climate where people expect concerns to be ignored. A clean decision audit with weak observations may mean that governance is documented while the design of work still defeats execution.
These contradictions should be treated as hypotheses, not as proof of deception. Each method has a different sampling frame and a different relationship with the people providing the evidence. The executive task is to connect the sources around one concrete issue. If the issue is stop-work authority, compare survey confidence, observed pause behavior, and the last three decisions in which work was delayed or allowed to continue.
A useful review asks three questions. What do people say? What do they do when the task changes? What do leaders decide when the control becomes inconvenient? The gap between those answers identifies the part of culture that deserves attention.
How to combine evidence without averaging it
Do not create a composite culture score before interpreting the sources. Averages can hide the exact disagreement that carries the operational risk. Instead, select one theme, such as contractor control, bad-news escalation, or temporary change, and build a small evidence chain around it.
Start with the survey to identify where confidence, pressure, or voice differs by role and location. Follow the pattern into the field and observe the task that produces the concern. Then trace a recent decision involving the same control, including who accepted the exception and what evidence was reviewed. The sequence turns a broad perception into an operating diagnosis.
Each source should produce a different artifact. The survey produces a segmented perception pattern. The observation produces a condition tied to a task. The decision audit produces an ownership and escalation map. When those artifacts point to the same friction, the intervention has a stronger basis than any method could provide alone.
Use the same chain after the intervention. A new procedure is not proof of cultural change, and a better survey score is not proof that the control works. Look for the changed decision, the changed field condition, and the changed experience of the people who must carry the control.
Decision matrix for leaders and EHS teams
| Situation | Lead with | Why it fits | Follow with |
|---|---|---|---|
| Leaders need a broad view across sites or roles | Survey | It reveals where trust, pressure, and voice differ across the organization. | Field observation in the highest-contrast group |
| A procedure is frequently bypassed | Field observation | It exposes the task friction and local condition that make the designed method unreliable. | Decision audit of the procedure owner and exception path |
| Controls are repeatedly accepted late or informally | Decision audit | It traces authority, escalation, and the conditions that permit risk to remain open. | Survey questions about confidence to challenge the decision |
| People report low trust after a serious event | Survey | It provides a broad view of voice, fairness, and confidence in follow-up. | Field observation and audit of the response decisions |
| A major change is moving from design to execution | Decision audit | It tests whether ownership and exceptions are clear before exposure reaches the field. | Observation during the first operating cycle |
The matrix does not prescribe one permanent sequence. It identifies the source most likely to answer the immediate question, then names the evidence needed to test the conclusion. The order should change with the problem. A broad climate concern may begin with a survey, while an active exposure should begin in the field rather than wait for a questionnaire.
Recommendation by context
For a board or executive team, lead with a survey only when the question concerns patterns across the organization. Pair it with a decision audit before funding a culture program, because leaders need to know whether the reported climate is connected to authority, resources, and escalation. The executive output should be a small number of operating decisions, not a larger dashboard.
For an EHS manager diagnosing a recurring deviation, lead with field observation. Watch the work under normal conditions, ask what makes the intended control difficult, and identify the decision that created the workaround. Use the survey to test whether the same friction is experienced elsewhere, then audit the ownership of the corrective action.
For a site preparing for a major operational change, lead with a decision audit. Confirm who can approve exceptions, what evidence is required before the change proceeds, and how the organization will respond when the new process does not perform as designed. Observe the first operating cycle and survey the affected workforce after they have encountered the real conditions.
For a safety culture transformation, use all three methods, but keep them tied to one or two critical operating themes. Andreza Araujo’s experience across 25+ years in multinational EHS leadership supports a simple discipline. Culture work becomes credible when leaders connect what they say, what the work requires, and what their decisions reward. The evidence should make that connection visible.
Survey data, field observation, and decision audits are not rival measures of the same thing. They are lenses on experienced culture, operating culture, and governance culture. Leaders should trust the signal that best fits the decision in front of them, then challenge it with the other two. When the sources disagree, the gap is not a reporting inconvenience. It is the evidence that tells you where culture is failing to reach the work.
Headline Podcast extends these questions through conversations about safety leadership, operational control, and the conditions that shape behavior. Explore the Headline Podcast for evidence-led perspectives from the international safety community.
Frequently asked questions
What is the best way to measure safety culture?
Are safety culture surveys reliable?
What does field observation reveal that a survey cannot?
What is a safety decision audit?
Should leaders combine survey, observation, and audit results into one score?
About the author
Andreza Araújo
Safety Culture Expert | Senior EHS Executive
Andreza Araújo is a safety culture expert and senior EHS executive with more than 25 years of experience in environment, health and safety. She is a Civil Engineer and Occupational Safety Engineer from Unicamp, holds a Master's degree in Environmental Diplomacy from the University of Geneva, and completed sustainability studies at IMD Switzerland. Andreza has served in Global Head of EHS roles in Fortune 500 environments, leading cultural transformation programs across multinational operations. She has represented Brazil as a speaker at the United Nations in Paris and has spoken at the International Labour Organization in Turin. She is the author of more than 16 books on safety culture in Portuguese, Spanish, English and German. Her work has earned more than 10 EHS awards, including two recognitions from Indra Nooyi, former PepsiCo CEO.
- Civil & Safety Engineer (Unicamp)
- M.A. Environmental Diplomacy (University of Geneva)
- Sustainability Cert (IMD Switzerland)
- People Management & Coaching (Ohio University)
- UN Paris speaker representative for Brazil
- ILO Turin speaker
- LinkedIn Top Voice
- Indra Nooyi PepsiCo CEO recognition (2x)
Documentaries
Watch Andreza's documentaries
Three productions on safety culture, organizational failure and the human lessons behind major disasters.
Podcasts
Listen to Andreza's podcasts
She hosts three shows on safety leadership, EHS and organizational culture, in English and Portuguese.