Olist 04: The Number Is Not the Answer. The Question After It Is

This is the last post in the series: what separates running the numbers from actually deciding what they mean for the business.

Series, Part 4 (Final) | A plain-language read on the Olist Brazilian e-commerce root-cause report (full notebook on Kaggle, link at the end)

A lot of e-commerce analysis posts I’ve come across online stop at the same point. A coefficient turns out significant, the p-value comes in under 0.05, there’s a chart, and a sentence that reads something like “delivery experience significantly affects retention.” Then it’s over.

This report has a coefficient just like that too. But for me, the moment it got calculated wasn’t the moment a conclusion showed up. It was the moment a question showed up. A simple question, simple enough to skip right past. This number, so what?

This post isn’t about another finding in the report. The first three posts already covered plenty of those. This one is about what actually goes through my head after a number gets calculated.


1. A number, before it gets translated, isn’t really anything

Under the delivery-anchored definition, the effect of “was the order late” comes out to −0.57 percentage points, and it’s statistically significant.

Honestly, seeing that number the first time didn’t produce any reaction at all. −0.57 doesn’t sting, it doesn’t even feel like something that happened in the real world. It reads more like a symbol sitting in a row of a regression output table, with a few asterisks next to it.

Stop the analysis right there, write “delivery delay has a significant negative effect on repeat purchase,” and technically nothing about that sentence is wrong. Logically it holds up fine. But it doesn’t answer any question an operator actually cares about. It just translates a mathematical fact into a sentence that sounds more sophisticated. Nothing has actually moved forward.

The real first step isn’t accepting the number. It’s forcing yourself to translate it into something you can actually picture. If every late order on this platform suddenly arrived on time overnight, how much would the repeat rate go up? Run the math, and the answer lands somewhere between four and seven hundredths of a percentage point.

That’s the moment the uneasy feeling shows up for the first time. Because the next question comes automatically. Is a number this small worth spending resources on? If I were running operations, holding a limited budget and a limited team, would I actually put this on the list of things to fix this quarter?

Statistics can’t answer that question. All statistics can tell you is that an effect is real, not noise. It can’t tell you whether it matters. And for anyone actually making decisions, whether it matters is the only question worth spending time on.

2. A number without a reference point doesn’t mean anything

Looking at 0.04 to 0.07 percentage points on its own, there’s no way to tell if that’s big or small. Every isolated number is like this. Strip away the reference point, and it’s not really anything.

So the next step isn’t calculating something else. It’s stopping and asking a question first. What should this number get measured against?

This isn’t a technical question, it’s a judgment call. Pick the wrong reference point, and a real but small effect can get dressed up as an important finding. Pick the right one, and a real but small effect gets honestly placed back where it belongs. The reference point the report settled on is a rule of thumb that’s been tested repeatedly across the e-commerce industry. An effect that can genuinely carry retention usually needs to reach double digits, somewhere around 15%, before it counts as a lever a business can actually lean on.

Put 0.04 to 0.07 next to 15%, and the answer becomes obvious immediately. This isn’t “a slightly smaller effect.” This is a gap of two orders of magnitude, the kind of gap where pouring in every available resource still wouldn’t close the distance.

There was, honestly, a bit of reluctance in writing this conclusion down. The coefficient is significant, the p-value looks clean, and if you wanted to, you could lean entirely on the word “significant” and dress this up as a legitimate finding. But something written that way is written for statistical software, not for a person trying to make a decision. A significant but negligible effect, framed as though it were worth investing in, misleads anyone reading this report to plan a budget around it, even when every individual number in it is true.

Calculating an effect size was never about landing on a precise number. It was about answering a real question for a real person: should this get done, or not. That’s what this whole step was actually doing.

3. When one path doesn’t work out, an operator doesn’t stop there

If the analysis ended at “delivery isn’t the lever,” this report would have failed at its job. Because anyone genuinely responsible for a business, hearing “this path doesn’t work,” has the same next thought every time: so what does?

That’s exactly what the report does next. Delivery experience doesn’t clear the bar, so the search moves to a lever that might actually clear it. There’s a group of customers in the data, 14.7% of the total, holding 28.8% of total transaction value, who’ve recently started to churn. This group surfaced on its own from the data. It wasn’t a hunch that went looking for evidence afterward.

But finding this group is only half the work. The other half is answering a different question. Is winning them back actually worth it? That question sounds simple, and it’s exactly the step that gets skipped most often. A lot of analysis stops at “we found a group of high-value customers at risk,” adds a line like “recommend strengthening retention efforts for this segment,” and calls it a conclusion. It sounds like one. It doesn’t answer anything a budget decision could actually be built on.

So the report goes further and runs the numbers. How much it costs to win back one customer in this group. Roughly what odds that attempt has of working. How much revenue one recovered customer contributes over a full year. Put those three numbers together, and that’s what actually answers “is this worth it,” instead of stopping at “this group matters,” which is true but empty on its own.

Both paths, delivery experience and customer win-back, produced numbers. Only one of them produced numbers worth building a decision on. Not because one dataset was cleaner than the other, but because only one path asked “what does this mean for the business” all the way through, instead of stopping at “is this coefficient significant.”

4. This is the actual dividing line

A data operator’s job ends at calculating a number. An analyst’s job ends at answering the question an operator actually cares about.

What sits between those two isn’t technique, isn’t tooling, isn’t even knowledge. What sits between them is where you place yourself. Standing in the calculator’s seat, the job is producing a correct number. Standing in the decision maker’s seat, the job is answering whether something should get done, whether it’s worth doing, where it ranks against everything else competing for the same budget. Same data, same code, even the same model, and depending on which seat you’re sitting in, what comes out looks completely different.

This is also why this report, from day one, was written with conclusions up front and evidence following behind. Because anyone who’s actually going to use this report to make a decision, in the first second of opening it, isn’t asking “how was your model specified.” They’re asking “so what should I do.” Methodology and evidence matter, but their place is supporting that answer once it’s been given, not delaying it.


A significant coefficient only proves one thing: that this is real.

Whether it’s worth doing, whether it matters, where it ranks, that always needs another sentence to answer. And that sentence never comes from a calculation. It comes from someone who understands the business, willing to sit with the number for one more minute after the math is done.


The full report, including all code, charts, and cited sources, is published on Kaggle: [link placeholder]

This is the last post in the series. If you’ve followed it from the first post to here, thank you for sticking with this report all the way through. It started with a growth curve hiding a contradiction, passed through a timeline that nearly got read backward and a KPI ruler that got quietly moved, and lands here, on the plainest sentence in the whole thing: once the math is done, there’s still work left to do.