Every Company Wants a Data Scientist. Few Know Why.

Four years ago the title didn’t exist. It was coined, more or less, by DJ Patil at LinkedIn and Jeff Hammerbacher at Facebook, who needed a word for people who were doing something that wasn’t quite statistics, wasn’t quite engineering, and wasn’t quite business analysis.

Now it’s on hiring plans everywhere, including at companies that could not tell you what the person would do on their first Monday.

I’ve watched a few of these hires land and I want to write down what goes wrong, because it is almost never about the candidate.

What the Job Actually Requires

Drew Conway drew a Venn diagram a while back that has stuck with me because it’s honest about the awkwardness of the role. Three circles:

  • Statistics. Knowing that a difference can look real and not be, and how to tell.
  • Programming. Being able to get the data yourself, clean it, and reshape it without waiting three weeks for someone in IT.
  • Domain knowledge. Understanding the business well enough to know which questions are worth asking and which answers are obviously wrong.

The overlap of all three is the job. The overlap of the first two without the third is, in Conway’s framing, the danger zone: technically flawless analysis of something nobody needed to know.

Almost nobody has all three when they arrive. The third one, in particular, cannot be hired for. It has to be acquired inside your building, which takes months, which means your data scientist will be unproductive for a while and you need to have decided in advance that this is acceptable.

“You are not hiring a finished product. You are hiring two thirds of one and agreeing to supply the rest, and most companies never budget for that.” — Sameer Gupta

Where These Roles Die

Four failure modes, in rough order of how often I’ve seen them.

They cannot get to the data. This is the big one and it is embarrassingly common. You hire an expensive specialist and then they spend their first quarter raising tickets to get read access to systems, because the person who owns the warehouse has a change control process designed for a different era. I’ve seen six-figure hires spend four months in a queue.

There is no route to production. They build something that works, in a notebook, on their laptop. Then it needs to run every night against live data, with monitoring and a rollback plan, and there is no team whose job that is. So the model lives in the notebook forever and gets run manually when someone remembers. Nothing changes in the business.

They get turned into a reporting function. Within six months they’re producing weekly dashboards, because dashboards are what stakeholders know how to ask for. This is a waste of the salary and everyone senses it. The person leaves within the year.

Nobody senior owns the outcome. The analysis says the current approach is wrong. The person whose approach it is disagrees. Absent an executive who wants the truth more than they want to be right, the analysis loses. It always loses.


What I’d Do Before Hiring

If I were setting one of these roles up, in this order:

  • Name the decision. Not a department, not a data source. One decision, made regularly, that you believe is currently made badly. If you can’t name it, you’re hiring on fashion and should stop.
  • Sort out data access first. Before the offer letter. Have the credentials, the environment, and the permissions ready on day one. This alone puts you ahead of most companies.
  • Decide who ships it. Somebody has to take a working model and make it run reliably. If that’s the same person, say so and hire for it. If it’s an engineering team, get their commitment before, not after.
  • Give them a sponsor with authority. Someone senior enough to act on an unwelcome finding, and personally invested in the answer rather than in the initiative.
  • Set a ninety-day question, not a mandate. “Improve our use of data” is unfalsifiable. “Tell us whether our churn model is better than the current rules” can be answered, and answering it builds the credibility for the next thing.

On Where They Sit

I don’t think there’s a universal answer, but the two common structures fail in predictable directions.

A central team builds better technical practice and worse business intuition. They’re further from the decisions, so their work is more rigorous and less used.

Embedded in a business unit is the opposite. Closer to real problems, more likely to be acted on, and more likely to reinvent things badly in isolation because nobody is comparing notes.

For most companies starting out I’d embed them, because the failure mode of irrelevance is worse than the failure mode of duplication. You can consolidate later. You cannot retrofit relevance.

Final Thoughts

The demand for this role is real and it isn’t going away. The techniques are getting more accessible, which if anything increases the value of the person who knows which technique to apply and whether the result means anything.

But the hire only works if the organisation around it works. Access, a production path, a sponsor, and a specific question. Get those four things right and a good person will pay for themselves quickly.

Get them wrong and you’ll conclude that data science doesn’t work for your business, when what actually happened is that you never let it start.