Data engineering
Pipelines, lineage and metric definitions your regulator and your finance director can both agree with.
- Change data capture
- Lineage
- dbt
- Data contracts
- Regulatory reporting
- Retention
- Code
- DAT
- Class
- B · 6 to 11 months
- Engagement
- TYPICAL 5-10 MONTHS · FIXED-PRICE DISCOVERY
- Stages
- 05
- Deliverables
- 06
- Sections
- 06
Overview
Reporting disagreements are rarely a maths problem. They happen because the same field is computed three ways in three places and nobody can trace which one the board pack used. We build pipelines with lineage from source column to published figure, define each metric once, and keep the definitions reviewable by the people who own them. Cambridge Mutual’s regulatory reporting run went from nine days of manual reconciliation to an overnight batch with a signed lineage trail behind every number.
Benefits
05 pointsOne definition per metric, versioned and reviewed by its business owner, so “active customer” means the same thing in the board pack and the regulatory return.
Column-level lineage from source system to published figure. Any number in a report can be traced back in minutes instead of reconstructed over a fortnight.
Tests on the data, not only the code: freshness, volume, referential and distribution checks that fail the pipeline before a wrong figure reaches a report.
Retention and minimisation enforced in the pipeline. Personal data is masked, tokenised or dropped at ingest according to its classification, rather than cleaned up afterwards.
Reruns are deterministic. Replaying a pipeline for a past date reproduces the figures as they stood, which is what makes a restatement defensible.
Workflow
05 stagesSource and definition audit
We trace the numbers people actually argue about back to their sources and record every place each metric is currently computed. This is the step where the disagreements get named and owned.
Ingest and contracts
Change data capture or batch extraction from source systems, with a schema contract at each boundary so an upstream change is caught at ingest rather than in a report three weeks later.
Modelling layer
Transformations as version-controlled, tested SQL. One definition per metric, reviewed by its owner, with the reasoning recorded next to the code instead of in a decommissioned wiki.
Serving and access
Published models exposed to BI tools, downstream systems and regulatory returns, with row- and column-level access mapped to your existing data classification.
Operate
Freshness and quality alerting routed to the team that owns the data, plus a documented backfill and restatement procedure for when a source system corrects history.
Deliverables
06 items- Metric definition catalogue: one owner, one definition and one implementation per published figure.
- Ingest pipelines with schema contracts, using change data capture where the source supports it.
- Version-controlled transformation models with freshness, volume, referential and distribution tests.
- Column-level lineage graph from source system through to each published figure.
- Access model mapped to data classification, with masking and tokenisation applied at ingest.
- Backfill and restatement runbook, including how to reproduce a prior reporting period exactly.
Questions
04 entriesWhichever your volume and team shape justify, and for most organisations we work with a well-modelled warehouse is the honest answer. Architectural labels tend to arrive before the problems they solve. We recommend the smallest thing that meets your reporting, retention and access requirements, and tell you what would have to change for the larger option to pay off.
Yes, and the classification drives the design. Orrery Health’s clinical pipelines mask direct identifiers at ingest, pseudonymise at rest, and restrict re-identification to a single named and audited role. We build to the classification you already hold; if you do not have one, producing it is the first week of work.
Common, especially with mainframe and packaged ERP sources. Depending on the platform we use log-based change data capture, batch extracts against a replica, or a nightly file drop. All three are workable. The constraint changes achievable latency, not feasibility, and we agree the freshness target per source before committing to it.
Your team, and we build for that from the start. Transformations are SQL in your repository, tested in your CI, running on infrastructure you own. We pair with named engineers throughout, and the last six weeks of an engagement are deliberately run by them with us in support.
Start a project
02 locationsEnterprise systems consultancy
- Manchester, United Kingdom
- Oslo, Norway