< Home
Software quality depends on what you choose to observe
Instrumentation does not just reveal quality. It defines what teams can see.
By Brian Ochan | 12/08/2026
Every couple of weeks or so, I am on call at work. Being on call is stressful enough by itself, especially the quiet anxiety that something might go wrong and you will have to make sense of it quickly. While on call a couple of weeks ago, I spent more time than usual moving between AWS dashboards, logs and Sentry, looking for reassurance that everything was fine. Nothing appeared alarming. Error rates were low, services were healthy and the graphs stayed mostly within their expected ranges.
But our application is used globally, serving users across different markets and handling millions of requests. That made me wonder what fine actually meant. If a call came in saying something was down, where would I begin? How would I decide which signals mattered and work backwards towards the problem?
Every dashboard begins with an act of omission. A complex system produces more information than any team could reasonably collect, display or understand. We choose a handful of signals, give them names and thresholds, then arrange them into something we can monitor at a glance.
That question is not new to me. For my master’s thesis at Lund University, I studied dashboards and the decisions people make from them. I was interested in whether the way information was visualised could influence how accurately people interpreted it and the conclusions they reached. What has stayed with me most is a simpler question: how much of reality can a dashboard ever show? Years later, I keep returning to that question in software engineering.
What a dashboard leaves out
One definition we used in my thesis described a dashboard as a visual display of the most important information, consolidated on a single screen so it can be monitored at a glance. The phrase most important is doing a lot of work here. Important to whom? For which decision? Over what period? Under which conditions?
Someone has already answered those questions before the dashboard reaches us. They decided which signals to collect, which ones to aggregate, where to put the thresholds and what to leave out. That omission is not necessarily a flaw. It is the point. A dashboard that showed everything would help us see very little.
In software, we usually start with metrics, logs and traces. OpenTelemetry calls these observability signals 1: evidence that helps us understand what a system is doing. But the system only emits what we instrument it to emit. A dashboard is therefore not the system made visible. It is a model of the system built from the evidence we chose to collect. This essentially means models can be accurate and still be incomplete.
The green dashboard problem
Consider a planning application that lets someone configure a product and save their work. The dashboard says the save endpoint is available. Response time is well below the agreed threshold. 99.9 per cent of requests succeed. There are no exceptions and the database is healthy. From an operational point of view, the feature works, but that tells us nothing about the people who never reach the endpoint. Perhaps the save action is difficult to discover. Perhaps users are unsure what will be saved. It could also be that the request succeeds but they cannot find their work when they return.
The service can be healthy while the task it exists to support is failing. This is close to what information-systems research calls cognitive fit 2: a representation is most useful when it fits the task someone is trying to perform. The same dashboard can answer one question very well and another very badly.
Did the save service execute successfully?
That is an operational question
Did the user successfully save what they intended to create?
That is an entirely different question
The data has not become incorrect. The question has changed. OpenTelemetry makes a similar distinction when discussing service-level indicators 3, noting that useful indicators should represent behaviour from the user’s perspective, not only the machinery underneath it. Availability matters, but a service being available does not guarantee the outcome a user expected. This is where system health and software quality begin to separate. The request succeeded. The process stayed alive. The logs contained nothing alarming. Every recorded fact may be correct. The dashboard did not lie. We asked it a bigger question than the data could answer.
Instrumentation is a design decision
Suppose we instrument that save flow with three events:
It looks comprehensive enough. We can calculate a success rate and alert when failures rise. But what does save_succeeded actually mean? That the server returned 200? That the data was persisted? That it can be retrieved after a refresh? That the user understood the save had completed? That what we stored matched what they thought they had created? Clearly, each definition measures something different and the denominator matters too. If our success rate includes only attempted saves, we exclude anyone who wanted to save but could not find the action. If we aggregate everything into one global number, we may hide a problem affecting one browser, device or market. Even a "green" threshold is a decision. Green does not mean “good”. It means a measurement landed on the acceptable side of a boundary somebody chose.
That is why I increasingly think of instrumentation as a design decision. Event names, dimensions, sampling, time windows and thresholds encode assumptions about how we expect the system to behave. An event can look like an objective fact. In practice, it is often a hypothesis expressed in code.
The person reading the dashboard matters
My thesis did not produce the simple result we initially suspected. We ran a double-blind randomised field experiment with 87 business practitioners to explore whether visual similarities in dashboard data could contribute to decision bias. We did not find a clear relationship between the visual features we tested and the specific bias we were looking for. What was more interesting was how interpretation differed between people. Participants with higher interpretation accuracy aligned more closely with the correct base rates, although the study was exploratory and the result was not strong enough for a broad claim.
That nuance has stayed with me because dashboards are not inherently clarifying or misleading. Their usefulness depends on the data, the design, the question and the person reading them. Presenting correct data does not guarantee correct understanding. The underlying data may have one source. Meaning rarely does.
Different signals answer different questions
The answer is not to distrust dashboards per se. I rely on them every time I am on call. Nor is the answer to instrument everything. More data can create its own form of blindness. I believe a more useful approach is to recognise that different signals answer different questions:
Did the software execute as expected?
Operational telemetry can help
Did people complete the flow?
Behavioural instrumentation can help
Why did they behave that way?
Qualitative evidence can help
None is a substitute for the others. An exception may tell us a request failed. Funnel data may tell us people abandoned the task before making the request. A support conversation or usability session may tell us they did not trust what would happen if they completed it. Taken together, those signals get us closer to the system we think we are observing. The questions I now find useful are not only What should we measure?
- What decision will this measurement support?
- What could still be going wrong while this indicator remains green?
- What evidence would make us question the story the dashboard is telling?
These are not arguments against measurement. They are arguments for being more deliberate about what our measurements can and cannot tell us.
Quality is a judgement made from evidence
Software quality is not a hidden property waiting for the right dashboard to reveal it. Even formal quality models resist collapsing it into one number 4. Reliability, performance, usability, accessibility and maintainability describe different parts of the same product. Improving one does not automatically improve the others. What we call quality is ultimately a judgement made from incomplete evidence. So I believe software quality depends on what you choose to observe. So does what you choose to improve. A healthy dashboard is valuable evidence, not a verdict.