Data Quality Under the Lens is a new Real World Data Science column. Each edition explores real-world moments where data quality shaped outcomes, sometimes driving failure, sometimes preventing it. From near misses to hard lessons learned, we look at what happens when data is up to the task… or falls short.
If you spot a real world problem and think data quality could lie at the heart of the story, send it in to the RWDS mailbox and our Data Quality Detectives will analyse whether the Silent Drift, Proxy Trap, Spreadsheet Cascade, Governance Vacuum or Metric Mirage is responsible.
The Case of the Month
By the early 1990s, the US Department of Defense (DoD) had a problem that will sound familiar to anyone who has worked with a large organisation’s legacy systems: it had too many systems doing essentially the same thing.
The DoD had inherited a collection of systems developed independently by different parts of the organisation. As this was the early days of computing, groups based in different locations had created a multitude of good, homegrown IT systems that met local needs. These were built on a variety of technology stacks with many types of hardware and software with systems often “growing” organically.
Early in my data management career, I was part of a team brought in to understand these systems and help the DoD determine how they could be consolidated. By this point, 37 systems were involved in paying DoD personnel. They all worked, but we needed to decide what should happen to enable the organisation to manage them as a whole.
For example, a Pentagon manager might want to know how many employees worked at a particular location. This sounds like a straightforward question, but the answer depended on what each system meant by ‘employee’. A significant proportion of the DoD workforce had more than one job within the organisation. Different payroll systems handled these arrangements differently: one system might count people by counting paychecks, while another might treat the primary and secondary jobs differently. The resulting numbers could therefore be different without either system necessarily containing an obvious error, and the systems could be answering different questions while appearing to answer the same one.
This became particularly important when the DoD needed to decide which of its existing systems should provide the basis for a more consolidated approach. The 37 systems had similar purposes and, when their processes were compared, they looked remarkably similar. Process models showed the same basic pattern: personnel information went in, it was processed, and pay was produced.
Process analysis therefore did not provide a useful way of distinguishing between the systems. We needed to look somewhere else: at the data.

What Actually Happened?
Our team was asked to reverse-engineer selected legacy systems. Rather than starting with the documentation describing what a system was supposed to do, we examined the systems themselves to recover the business rules, requirements and data structures embedded within them. This revealed something that the process models had hidden: although the systems appeared to perform the same broad function, they did not necessarily represent the underlying business in the same way.
The distinction became particularly clear when we tested the systems against specific requirements. One memorable example involved a hypothetical but deliberately highly specific employee: a one-legged engineer working in waist-deep water, underneath rotating helicopter blades, on overtime. The point of this unusual example was not that the DoD had large numbers of employees fitting that description. It was that a requirement this specific could expose differences between systems. If a system’s data structures could represent the necessary combination of circumstances, that provided evidence that it could support the requirement. If they could not, the system could be ruled out.
This gave the DoD a way of making a decision based on evidence about the systems rather than on competing claims about which system was best. It also helped address the organisational politics surrounding the consolidation. Each system had people and resources associated with it, and there were understandable concerns about what would happen if a particular system was replaced. A transparent comparison based on agreed requirements provided a more objective basis for deciding which system could meet the organisation’s needs.
Disaster or Near-Miss?
The fact that the DoD could not always answer basic questions about its workforce consistently shows that the fragmented systems were already causing real problems. However, the greater risk arose when it needed to choose a system to take forward. Selecting one based on popularity, existing use, or apparently similar processes could have embedded its limitations into an enterprise-wide system.
By reverse-engineering the data models, we were able to identify differences that the process analysis had missed and test the systems against the DoD’s actual requirements. This meant the consolidation decision could be based on evidence rather than assumption and the underlying weaknesses were exposed before they were carried into the next generation of systems.
Why This Matters Now
Technology has advanced, but this problem hasn’t disappeared. In June 2026, the US Office of Personnel Management announced a $395 million contract with Oracle to modernise federal human-resources systems. The programme aims to consolidate more than 100 separate human-resources management systems across the US federal government into a single platform. The stated case for consolidation includes duplicated systems and functions, inconsistent application of policies and the cost of maintaining a fragmented HR environment.
The scale is very different from the DoD example of the early 1990s, but the underlying question remains relevant: what knowledge about the organisation is embedded in the systems being replaced?
Replacing many systems with one does not automatically create better data. A new system still needs to represent the concepts, rules and exceptions that matter to the organisation. If those things are poorly understood during the transition, consolidation can simply move existing problems into a new system. The DoD experience suggests that understanding legacy data should be part of the modernisation process, rather than something left until after a new system has been selected.
The Data Quality Pattern
This case illustrates a classic Governance Vacuum: systems were developed to meet local needs, but there was insufficient enterprise-level agreement about how the organisation’s data should be defined and represented.
The Practitioner Takeaways
We can take a lot from this case!
Define the question before counting the data. A seemingly simple question such as “How many employees do we have?” may have more than one valid interpretation. Agreeing what is being counted is a prerequisite for producing a useful number.
Don’t assume that systems doing the same job represent the same things. Similar processes do not necessarily mean compatible data. Look at the definitions, relationships and business rules embedded in the data.
Understand legacy systems before replacing them. Old systems may contain years of accumulated organisational knowledge, including requirements and exceptions that are not documented elsewhere. Reverse engineering can help uncover this knowledge before it is lost.
A single system is not automatically a single source of truth. Consolidation can reduce duplication, but data quality still depends on whether the new system represents the organisation consistently and appropriately.
Data models can be valuable evidence when organisations are making decisions about systems. Looking at the data structures helped reveal differences that were not apparent from looking at the processes alone.
The one-legged engineer is more than an entertaining anecdote. The example worked because it forced the team to ask a precise question about what the system needed to be able to represent. Sometimes an unusually specific requirement is exactly what reveals a hidden difference between systems.
- About the author:
- Peter Aitken, PHD is an acknowledged Data Management authority, an Associate Professor at Virginia Commonwealth University, President of DAMA International, and has founded several companies that have helped more than 200 organizations leverage data–specific savings have been measured at more than $1.5B USD. His latest is Anything Awesome LLC.
Copyright and licence : © 2026 Peter Aiken
This article is licensed under a Creative Commons Attribution 4.0 (CC BY 4.0) International licence.
How to cite :
Peter Aiken 2026. “Data Quality Under the Lens: the DOD’s One-Legged Engineer Problem”, 2026. URL