Big Data Is a Storage Bill, Not a Strategy

I’ve now sat in four meetings this year where a senior person said some version of “we need a big data strategy.”

In none of those meetings did anyone say what question they wanted answered.

That gap is the whole subject of this post.

What Actually Changed

Let me be fair to the hype first, because something real did happen.

Storing and processing very large volumes of data used to require expensive specialised hardware. Hadoop changed that. It spreads data across a cluster of ordinary machines and moves the computation to wherever the data already sits, rather than dragging terabytes across a network to a central processor. It tolerates machines failing, because with enough cheap machines, some of them are always failing.

Cloudera has been commercialising this for three years. Hortonworks spun out of Yahoo in June. McKinsey published a substantial report in May arguing this is a genuine productivity frontier, and I think they’re broadly right.

So the capability is real. My argument is about what companies are doing with it.

The Failure Mode

Here’s the pattern I keep seeing.

A company decides data is strategic. It provisions a cluster. It starts retaining everything, because storage is now cheap and nobody wants to be the person who deleted the thing that turned out to matter. Log files, clickstreams, sensor readings, transaction details, all of it, indefinitely.

Eighteen months later there are hundreds of terabytes, a meaningful line item on the infrastructure budget, two engineers who keep the cluster alive, and not one decision in the business that is being made differently.

“You didn’t build a data capability. You built a warehouse, filled it, and hired security. Nobody has been inside to look for anything.” — Sameer Gupta

The clue is that the project’s success metrics are all about volume. How much data we’re collecting. How many sources we’ve connected. How fast we can query it. Those are inputs. Nobody is measuring outputs, because nobody defined one.

The Question Comes First

The discipline I’d push, and it’s unglamorous, is to run this in the opposite direction.

Start with a decision somebody makes badly. Not a dataset. A decision. Which customers get a retention call. How much stock to hold in a given depot. Which transactions get manually reviewed. Which machine gets serviced next.

Then ask three things:

  • What would we need to know to make that decision better? Be specific. “More data” is not an answer.
  • Do we have it, or could we get it? Sometimes the answer is that the relevant data was never captured, and no amount of cluster solves that.
  • If we knew it, would we actually act differently? This is the question that kills most projects, and it should. If the answer is no, because the process is contractually fixed or the decision is really made politically, then stop. You’ve saved a year.

I’ve watched teams get to that third question and discover the honest answer was no. That’s not a failure. That’s the cheapest possible outcome, arrived at before the spending.

The Constraint Nobody Budgets For

Here’s the part I’d put in front of a CFO.

The cost curve on storage has collapsed. The cost curve on people who can ask a good question of a large dataset and correctly interpret the answer has not moved at all. If anything it’s gone up, because demand has risen and supply hasn’t.

  • A cluster is a purchase order. You get it in weeks.
  • Someone who can formulate a hypothesis, test it properly, and recognise when a correlation is an artefact of how the data was collected, is a hire, and a difficult one.

Which means the binding constraint on your data strategy is not infrastructure. It’s a small number of scarce people, and buying more storage does nothing whatsoever to relieve it.

“The cheap part of this got cheaper and the expensive part didn’t move. Everyone is optimising the half that was already solved.” — Sameer Gupta

Three Things I’d Do Instead

  • Fund one question, properly. One decision, one small team, one measurable outcome, ninety days. If it works you’ll have a case study to fund the next one. If it doesn’t you’ll have learned something for the price of a quarter.
  • Retain deliberately, not reflexively. Every dataset you keep is a cost, a security exposure, and eventually a legal liability. “We might need it” is not a retention policy.
  • Hire the interpreter before the infrastructure. One person who knows what to ask will get more from a laptop and a well-chosen sample than a team with a cluster and no hypothesis.

Final Thoughts

There’s going to be a lot of money spent on this over the next few years, and a good deal of it will be spent on storing things nobody looks at.

The companies that do well won’t be the ones with the largest clusters. They’ll be the ones who picked three decisions that actually matter, got the data required for exactly those, and changed how they operate as a result.

Data isn’t the new oil. Oil is worth something in a barrel. Data is worth something only after somebody asks it a question.