Olist 03: 11 Days, One Strike, and a Quietly Recalibrated Promise
A strike, a KPI that looked untouched, and the goalpost that had quietly moved to make it look that way.
Series, Part 3 | A plain-language read on the Olist Brazilian e-commerce root-cause report (full notebook on Kaggle, link at the end)
When I was an Account Manager at Amazon, I saw this play out more than once. Something breaks in fulfillment, an SLA is about to slip, and the instruction that comes down from leadership is almost never “push through and hold the line.” It’s “reset what the customer expects.” The promise itself, it turns out, is a variable you can move.
That’s not some shady trick. It’s a tool operations teams reach for constantly. Looking at Olist’s data with that same eye, I recognized the same playbook almost immediately.
1. A nationwide strike, and a chart that shows nothing
In May 2018, Brazilian truckers launched a nationwide strike that lasted about 11 days, and road freight essentially stopped. A disruption at that scale should leave an obvious dent in the delivery data of any e-commerce platform that depends on road transport.
But look at the most obvious chart, average delivery days by month, and the strike weeks barely register. The line keeps trending downward like nothing happened. A country-wide logistics shutdown, and the data looks untouched.
Inside the strike window (shaded red), the delivery-time curve shows no visible disruption, as if the shock never happened.
That’s the first thing worth questioning. Was it really untouched, or were we just looking in the wrong place?
2. Zoom in to the week, and the damage shows up
Monthly aggregation has a built-in problem. An 11-day disruption gets spread across a 30-day bucket and smoothed out almost entirely. And this particular chart tracks delivery month, meaning a package in transit can span two calendar months, which smears the impact even further.
Switch the lens and look directly at what happened during those two strike weeks, at the customer’s actual buying behavior. The damage shows up immediately.
Orders dropped hard. During the two strike weeks, weekly order volume fell to 954 and 1,083, down roughly a third from the usual range of 1,450 to 1,700.
Actual delivery time stretched out. Median delivery time went from 8 days the week before the strike to 12 and then 10 days during it.
At weekly resolution, the late-delivery rate first spikes to 27.6% and 24.8% two weeks before the strike (driven by Mother's Day demand), then drops to 4.0% and 0.4% during the strike itself, among the lowest in the entire window.
3. The real insight: the ruler got moved
Stop here and look at only one number, and the conclusion you’d reach is completely wrong.
That number is the late delivery rate. During the two strike weeks, it sat at 4.0% and 0.4%, among the lowest in the entire observation period.
Taken at face value, that reads as “fulfillment performance actually improved during the strike,” which makes no sense on its own. Here’s what actually happened. The promised delivery window itself got stretched wide open. The typical median promise sat around 20 days. During the strike weeks, it jumped to 36.
Loosen the ruler and hitting the target gets a lot easier. This isn’t the delivery system absorbing a shock and holding steady. It’s the platform seeing the shock coming and moving its own standard out of the way first, measuring a harder environment with a looser ruler.
This playbook is deeply familiar from the Amazon AM world. When a fulfillment system faces a systemic hit, what gets communicated down the chain is never “hold the line.” It’s “reset the customer’s expectations.” All the data can show is the trace this move leaves behind. What actually got discussed in whatever room the decision was made, the data doesn’t say. But the trace it leaves is more than enough to recognize the play.
4. A fork in the road that’s easy to get wrong
Before going further, one detail needs to get ruled out, or two unrelated events end up getting merged into one story.
In the two weeks before the strike, the late delivery rate had already spiked to 27.6% and 24.8%. At a glance, it’s tempting to read that as the strike hitting early.
But checking order volume for those same two weeks points to a different cause: Mother’s Day. Around May 13th is one of the biggest shopping peaks of the year, and order volume in those two weeks hit a new high for the entire observation window. Demand surged on its own, with no connection to the strike.
Get this fork wrong, and a purely seasonal spike ends up getting charged to the strike’s account. Keeping the two apart is what makes the next judgment, how big the strike’s actual impact was, hold up.
5. Two weeks later, everything resets
About two weeks after the strike ended, order volume, actual delivery time, and the promised window all returned to normal.
The report’s phrasing for this, locked in deliberately, sums it up well. Visible, brief, absorbed. The disruption was real and visible, but short-lived, and the system worked through it. An episode of this size, 2,037 orders, 2.1% of the entire observation window, could never be the explanation behind a structural problem as large as 97% of customers never buying twice. It’s an operational footnote that got handled cleanly, not the main plot.
6. The final word: a stable KPI can mean two very different things
What we saw. The strike really did stretch out actual delivery times, but the platform moved its promise standard at the same time, which made the surface-level “late delivery rate” metric look completely untouched.
What this means from an operator’s seat. Anyone who reads operational data should hold onto this. A stable KPI can mean the underlying reality is genuinely stable, or it can mean the ruler measuring it got moved. Judging whether a metric is healthy means never looking at the number alone. It means checking whether the baseline behind it changed too.
A question worth carrying forward. Next time you see a report showing “98% SLA compliance,” it might be worth asking one more question. Is that 98% because things actually got better, or because the ruler got looser first?
The full report, including all code, charts, and cited sources, is published on Kaggle:[Kaggle]
Next up: the part of the report most likely to get skipped over, and the part that demanded the most discipline. Why “delivery is a secondary lever” is a conclusion that got calculated, not one that got guessed.
Photo by Maayan Nemanov on Unsplash
Inside the strike window (shaded red), the delivery-time curve shows no visible disruption, as if the shock never happened.
At weekly resolution, the late-delivery rate first spikes to 27.6% and 24.8% two weeks before the strike (driven by Mother's Day demand), then drops to 4.0% and 0.4% during the strike itself, among the lowest in the entire window.