Skip to content

Vanta Logistics

Rebuilding dispatch for 6.2 million parcels a day

Vanta Logistics lost four hours of dispatch during peak week 2022 and 340,000 parcels missed their promised day. We rebuilt the scheduler as an event-driven service that has since carried 6.2 million parcels on its busiest day without a dispatch outage.

Client
Vanta Logistics
Sector
Logistics
Run
14 months · Mar 2023 – Apr 2024
Filed
2024-06-03
Capability
Platform modernisation · Systems integration · Reliability engineering
Stack
Go · Apache Kafka · PostgreSQL 16 · gRPC · Redis · Kubernetes · Google Cloud · Terraform · Grafana

Impact

  • 99.98%

    Dispatch availability in peak week, up from 97.1%

  • 40 s

    Reroute after a depot outage, down from 17 minutes

  • 6.2M

    Parcels routed on the busiest single day

Challenge

Logistics
14 months · Mar 2023 – Apr 2024

The scheduler was a single-writer monolith against one Oracle instance: it could not be scaled out, and it could not be tested, because no environment other than production had realistic parcel volume. A deploy took ninety minutes and a rollback took longer, so the team froze changes for the eight weeks around peak — which meant every fix found in October waited until January. When the scheduler stalled on 12 December 2022 it took nineteen minutes to notice and three and a half hours to recover, because recovery meant replaying a batch nobody had replayed before.

Solution

We extracted routing, capacity and depot assignment into separate Go services communicating over Kafka, so a depot outage now reroutes rather than blocks. Before any of it went live it ran in shadow against a mirrored copy of production traffic for eleven weeks, and we compared its routing decisions against the monolith’s parcel by parcel until the divergences were explainable. Failure was rehearsed rather than assumed: monthly game days kill a depot, a broker and a database primary in turn, and recovery runbooks are the artefacts those exercises produce. Deploys went from ninety minutes to six, which removed the argument for the peak-season freeze — the team now ships around forty times a week, including through December.

Plates

  • Service topology diagram showing routing, capacity and depot assignment as three separate Go services exchanging events over Kafka, with parcel intake on the left and depot handheld devices on the right.
    Plate 01One monolith became three services with independent failure domains.
  • Shadow-run comparison view showing routing decisions from the legacy monolith and the replacement service side by side for the same parcels, with a divergence count of 41 out of 2.4 million and each divergence categorised.
    Plate 02Eleven weeks of shadow traffic before a single real parcel moved.
  • Game day timeline recording a deliberately failed depot during a rehearsal, showing detection at eleven seconds, automatic reroute at forty seconds, and the runbook steps the on-call engineer followed.
    Plate 03Recovery times come from rehearsals, not from estimates.

Testimony

Our dispatch system had one test when they arrived. Arcwright wrote 1,400 more before they changed a single line of it. Six months later a release takes twenty minutes instead of nine hours, and nobody stays late for it.

Petra Lindqvist · VP Engineering, Vanta Logistics5 out of 5