insight
Operational vs analytical data: why one store fails at both
Software architect and consultant. Works with business-led product development, distributed systems, operational AI, and production software delivery.
Answer first
The warehouse answers questions about last month. The operational platform feeds systems acting right now. Different latency, different failure modes, different definition of correct.
An organization with a working data warehouse and a new operational use case will usually try to serve the second from the first. The data is already there, the pipelines already run, and building something parallel looks like duplication. Six months later the workflow is slow, the nightly load window has become a business constraint, and an analyst query has taken production down twice.
The two workloads have different requirements almost everywhere it matters.
Correctness means different things
For analysis, correct means complete. A report that omits yesterday's late-arriving records is wrong, and waiting for completeness is the right trade. Batch windows exist for this reason.
For an operational decision, correct means current. A case worker deciding now needs the state as it is, and a value that was accurate at three this morning may be actively misleading. Completeness that arrives after the decision has no value.
These pull in opposite directions. A platform tuned for one serves the other badly, and the compromise usually satisfies neither.
Failure behaves differently
A failed analytical load is an incident with a working day to fix it. The dashboard is stale, someone notices in the morning, the job is rerun. Nobody is blocked.
A failed operational feed stops work. Cases cannot be processed, or worse, they are processed against stale data and the error surfaces weeks later in something a customer notices. The platform needs the availability characteristics of the system it feeds, which is a different operational posture and a different on-call expectation.
- Analytical: recovery measured in hours, retry acceptable, degradation visible and tolerable
- Operational: recovery measured in minutes, retries must be safe to repeat, degradation has to be explicit rather than silently serving old values
Where the shared foundation still helps
None of this argues for two disconnected estates. Definitions, lineage, and quality rules should be shared, because an operational system and a report that disagree about what a customer is will produce an argument nobody can settle.
Share the semantic layer and the governance. Separate the serving path. One place defines what a customer is, and two paths deliver it under different guarantees.
One definition, two serving paths. The mistake is either two definitions, or one path asked to satisfy both sets of guarantees.
Signs you are using the wrong one
- A business process waits for a nightly load, and people have adapted their working hours to it
- An analyst query has degraded a production workflow, and the response was a policy about when to run reports
- The same entity has two different values depending on which system is asked, and reconciling them is somebody's routine job
- Operational users have built spreadsheets to work around freshness, which is the clearest signal and the easiest to miss
What we recommend
Decide the serving path from the decision the data supports rather than from where the data currently lives. If a person or a system acts on it during the working day, it needs an operational path with the availability and latency that implies.
Build the first one around a single workflow that genuinely matters. That proves the pipeline, the ownership model, and the quality rules against something real, and gives every later flow a pattern. Attempting a general operational platform before any specific workflow depends on it produces infrastructure that nobody trusts enough to use.
- What is the difference between an operational data platform and a data warehouse?
- A warehouse is optimised for analysing what has already happened, where correct means complete. An operational platform feeds systems acting during business operations, where correct means current. The two definitions pull in opposite directions, and a platform tuned for one serves the other poorly.
- Can we serve operational use cases from our existing warehouse?
- Sometimes, for tolerant use cases. It fails when a business process starts waiting on a load window, or when analytical queries begin affecting production. Both are signals that the workflow needs an operational serving path rather than a faster batch.
- Does that mean building two separate data platforms?
- No. Share definitions, lineage, and quality rules, because an operational system and a report that disagree about the same entity create arguments nobody can settle. Separate the serving path and its guarantees. One definition, two paths.
- Where should an operational data platform start?
- With one workflow that genuinely matters. That proves the pipeline, ownership, and quality rules against something real and gives later flows a pattern to follow. General platforms built before any specific workflow depends on them tend not to earn trust.
Newsletter
Occasional notes on building systems that hold up
A short email when we publish something worth your time. Architecture, integration, and operational AI in regulated organizations. No cadence promises, no forwarding your address.