Reliability engineering
Service level objectives derived from what customers actually notice, and the operational practice to hold them.
- SLOs
- Error budgets
- Incident response
- Failure rehearsal
- On-call
- Capacity
- Code
- REL
- Class
- B · 6 to 11 months
- Engagement
- TYPICAL 3-6 MONTHS · QUARTERLY REVIEW OPTIONAL
- Stages
- 05
- Deliverables
- 06
- Sections
- 06
Overview
An uptime percentage nobody derived from customer impact is theatre. We set objectives from real user journeys, instrument them honestly, and build the operational habits that keep them met: on-call people can sustain, blameless review with tracked actions, and capacity planned against measured headroom. Heliostat went from eleven severity-one incidents a year to two, and cut mean time to recovery from 4 hours 40 minutes to 38 minutes, without adding headcount.
Benefits
05 pointsObjectives tied to journeys customers notice, such as “payment confirmed within 3 seconds”, rather than a server-uptime figure that stays green through an outage.
Error budgets give a defensible answer to “ship or stabilise”. When the budget is spent, reliability work takes priority automatically instead of being argued for each quarter.
Incidents get shorter. Instrumentation and runbooks are built around the failure modes review actually surfaces, so responders stop starting from zero at 3am.
On-call that people will do for years: rotations sized against measured page volume, every alert actionable, and a written rule for what is allowed to wake someone.
Capacity planned from measured headroom and a growth model, so peak events are rehearsed rather than survived. Heliostat absorbed a 6.2x seasonal peak with no degradation.
Workflow
05 stagesJourneys and objectives
We map the handful of journeys that matter commercially and define an objective for each, with the business owner agreeing the target and — more importantly — what happens when it is missed.
Instrumentation
Objectives measured from the customer’s side of the system, using real user timings and synthetic probes rather than host metrics alone. Anything we cannot measure honestly does not become an objective.
Incident practice
Severity definitions, a single command channel, and blameless review with tracked actions. We facilitate the first several reviews with you, then hand the facilitation over.
Failure rehearsal
Controlled failure injection against the dependencies most likely to break — a slow database, a dead region, an expired certificate — in a scheduled window, with every finding feeding back into the runbooks.
Capacity and cadence
Headroom measured under load, a growth model with named assumptions, and a monthly reliability review where the error budget decides what gets built next.
Deliverables
06 items- Service level objectives per customer journey, with agreed targets and named owners.
- Instrumentation and dashboards measured from the customer’s perspective, in your own observability stack.
- Alert catalogue where every alert is actionable, routed and has a runbook — and the ones that are not are deleted.
- Incident response process: severity matrix, roles, communication templates and escalation path.
- Failure rehearsal reports with the remediation each exercise produced.
- Capacity model with measured headroom, growth assumptions and the trigger points for scaling.
Questions
04 entriesAlmost certainly not. Each additional nine roughly multiplies cost, and above about 99.95% the binding constraint is usually your dependencies rather than your own system. We work back from what customers notice and what your commercial commitments require. For most of our clients that lands between 99.9% and 99.95%, and the money saved goes into recovery speed instead.
Then do not run one until the objectives require it. A sustainable rotation needs at least six people, so with a smaller team the real options are a follow-the-sun arrangement with a partner, an out-of-hours scope covering severity-one only, or accepting a lower overnight objective and saying so honestly. We model those options against your measured page volume rather than assuming the full rotation.
In the form we run it, yes. Injection is scheduled, scoped to a named dependency, announced in advance and has an abort control. It is a fire drill, not a surprise. We start in pre-production and only move a given failure mode into production once the team has handled it there first. Kestrel Energy notified their regulator ahead of the first production exercise and had no objection.
No. We build the capability and hand it over, and we do not sell managed operations. Teams that outsource on-call permanently lose the feedback loop between the code they write and how it behaves at 3am, and reliability degrades quietly from there. We will sit alongside your rotation as a second pair of hands for the first two to three months.
Start a project
02 locationsEnterprise systems consultancy
- Manchester, United Kingdom
- Oslo, Norway