Every Estimate Is a Coastline
I hate software estimation.
Well, I hate what happens after software estimation. The estimating itself is mostly harmless. Six people sit in a room, stare at a paragraph somebody typed into a planning document, talk around the large holes in what we know, and eventually produce a number because a number is what the spreadsheet accepts.
Then the number gets sent up the rungs, loses every caveat we attached to it, and comes back down as a date. We (the engineers) curse ourselves for not padding it by 50% more.
Someone always says estimates aren’t commitments. Everyone nods. A roadmap gets built around the estimate anyway and, through some organizational alchemy I don’t fully understand despite participating in it for years, it becomes a commitment.
I’ve done my part. I’ve turned guesses into suspiciously specific numbers of engineering weeks because “somewhere between a month and several months… depends what we find” looked bad in the cell. I’ve also gotten annoyed when somebody else’s project took longer than they said it would. We all know estimation is inexact, but I’m not convinced any of us actually behaves as though it is.
At a previous company, we had a “contact state table” and a “contact state history table”.
Before I go any further, this was a legacy PHP monolith approaching 20 years old. Setting up new Kafka topics was discouraged until sometime in the early 2020s. Using MySQL to track the state of billions of contacts was not my dream architecture. It was what we had.
The first table stored where a contact was in each workflow. Waiting for a delay. Sitting at a condition. Ready for an email. The history table was its audit trail. Support, reporting, and segmentation used it to see how the contact got there.
Eventually the current table held more than a billion states. Its history passed 150 billion rows and 10 terabytes across our sharded databases. Every day, we upserted roughly 100 million state changes into one table and then did it again in the other. Only opens and clicks grew faster. Yes, we stored those in MySQL too. Please remember the monolith.
The databases were starting to wobble. The fleet cost a genuinely upsetting amount of money every month.
The job was to save the DBs from doom (and maybe dent the bill while we were at it).
Then a new CDP project brought in Debezium and, apparently, we had change data capture now. Before that, BI tables came from an enormous nightly Airflow job that essentially ran SELECT * against all of our tables. Lol. We could stop writing the history table and keep the audit trail in BigQuery instead.
I was told I was not allowed to plug directly into the CDC stream. Data Engineering pulled the changes into BigQuery and compacted them. We needed the raw history, so we quietly read the uncompacted source. For some reason, it lagged by as much as eight hours. Fine for some reports. Less fine when support was trying to explain why a contact had not received an email.
We considered partitioning the history by date, keeping recent data in MySQL, and dropping old partitions after BigQuery caught up. The in-house migration engine did not support partitions. I added support, but its new container image would not build, so we could not release it. I no longer remember why. Figures.
Meanwhile, support kept asking for history and the databases kept writing both copies.
From far away, this was one box and one arrow. State changes go to BigQuery. History table goes away. Nobody mentioned the failed approaches, the red tape, or the trifles of doing any of this inside a 20-year-old PHP monolith.
Donald Rumsfeld (please keep reading) was unfortunately onto something with the known knowns, known unknowns, and unknown unknowns. Even a broken clock is right twice a day.
At the start, you know roughly what the project should do and which systems it touches. You know some of the questions too. Who reads the history? How fresh does it need to be? Discovery answers a few of them. But the CDC stream you’re not allowed to touch, the eight-hour lag, the container image that won’t build… none of that was visible in the sentence “move the audit history to BigQuery.”
If complete knowledge were required before estimating, we could provide beautifully accurate estimates immediately after finishing the work.
Which brings me, somehow, to the coastline of Britain.
How long is the coast of Britain?
It sounds like a boring question with a boring answer. But it’s nuanced! Somebody measured it. Then several other people measured it. AND THEY GOT DIFFERENT NUMBERS.
Lewis Fry Richardson found the same problem while studying whether shared borders made countries more likely to fight. Spain reported that its border with Portugal was 987 kilometers. Portugal said the same border was 1,214 kilometers.
His work appeared posthumously in 1961 as The Problem of Contiguity: An Appendix to Statistics of Deadly Quarrels. It attracted almost no attention. Then Benoit Mandelbrot came across it, interpreted Richardson’s slopes in terms of dimension, and gave the problem its famous title in 1967: How Long Is the Coast of Britain? Statistical Self-Similarity and Fractional Dimension. Classic Mandelbrot.
Imagine walking a divider along a map of a coastline. Take long strides and you step past bays and inlets. The measurement comes out fairly short. Make the stride smaller and the divider follows some of the bays it skipped. Smaller again and it catches smaller bends, rocks, and whatever else the map can resolve. The measured coast gets longer.
If you keep going, eventually you run into tides, grains of sand, and atoms. There isn’t one useful answer at every possible scale. The length means very little unless you know how it was measured.
Here’s the same idea with a ruler you can drag:
Drag the ruler toward fine and watch the measured length grow.
The history-table project was measured with a long ruler. So are most projects, unless you’re cleaning up spaghetti code everybody already understands (lucky you, honestly). “Put the audit history in BigQuery and delete the table” was one stride. Looked manageable.
Then the ruler got shorter. Who still reads the history? How stale can it be? Can we use the CDC stream directly? Wait, we have a CDC stream now?! What keeps recent history available while BigQuery catches up? Can the migration tool partition the table? Who owns each part once we’re done?
Leave the uncertainty out of the estimate and it still shows up later, usually with interest and during a status meeting.
This doesn’t make every early estimate secretly good. Some estimates are lazy. Some are politically convenient. Sometimes an engineer hears a date hidden inside the question and reverse-engineers an answer that won’t make the meeting uncomfortable.
I have certainly never done that.
Richardson also noticed that coastlines don’t all react the same way when the ruler shrinks. A fairly smooth one settles down. A jagged one keeps revealing more bends, so its measured length grows much faster.
Three of the borders Richardson measured, drawn at one scale from modern outlines (Natural Earth) and walked with the same divider. Shorten the ruler and Britain nearly triples, while South Africa barely moves.
Some work stays roughly the same size as you inspect it. A small change in a familiar system, with settled requirements and no new dependencies, might actually be as boring as it looks.
Other work gets bigger fast, like a new integration, a whole new product, or some COBOL written in the 60s and 70s undergirding our entire financial system (or so I’ve heard). The rate depends on the team, the system, the quality bar, and what everyone already knows. The same feature can be smooth for the team that built the platform and an absolute fjord situation for the team meeting it for the first time. How long the coast is depends a lot on who’s holding the ruler.
Eventually, a project stops. A coastline doesn’t. Once the work is over, we know how long that version took. We still don’t know how long it had to take.
Refusing to estimate is generally unwise. Somebody else will make the guess, and now you have an estimate you don’t believe and no say in it. Hedging has limits too. I once tried that with a short-tempered director who needed a number lower than the one I gave. He screamed at me in a meeting. This did not improve the estimate.
So give the number. Just send the ruler up with it: how closely you’ve looked, and what’s most likely to make it grow. When it grows anyway, reopen the date, the approach, or the scope instead of quietly eating the difference.
We deleted the history table in 2024, later than anyone wanted and with some rough edges. It was supposed to be one box and one arrow. I mean, I guess it still is if you stand far enough away 😉