Data & AI

Data Scientist mock interview

Expect experiment design and causal reasoning over model trivia. The recurring trap is a question where the data cannot support a confident answer — saying so, and stating what you would need, is the mark.

Data loops test two things that pull against each other: rigour with methods, and the judgement to know when a rough answer is the right one. Candidates who are strong on one are often visibly weak on the other.

Expect questions where the data is deliberately insufficient. Stating the assumption you are making, and what would change your answer, is most of the mark.

The rounds you can practise

Each round is run by the interviewer built for it, with its own scoring axes — not one generic interviewer asked to change subject.

Metrics, experiments, causal reasoning, and decisions under uncertainty.

Scored on Analytical communication · Data & inference rigor · Decision quality

Metrics, experiments, causal reasoning, and decisions under uncertainty.

Scored on Analytical communication · Data & inference rigor · Decision quality

Hands-on problem solving, code quality, debugging, and engineering judgment.

Scored on Problem-solving communication · Engineering correctness · Adaptability & craft

System design, trade-offs, and war stories from production.

Scored on Technical communication · Engineering depth · Ownership & collaboration

Metrics, experiments, causal reasoning, and decisions under uncertainty.

Scored on Analytical communication · Data & inference rigor · Decision quality

Data Scientist interview questions

Six you can expect, in the register interviewers actually use. Answer them out loud before you read the next section — reading a question and answering one are different skills, and only the second is marked.

  1. Design an experiment to test whether this feature increases retention.
  2. The metric moved but the experiment was underpowered. What do you tell the team?
  3. How do you tell correlation from causation with only observational data?
  4. Walk me through a model you built that did not get used.
  5. How would you detect that a metric moved for a boring reason?
  6. What would you need to answer this question properly, that you do not have?

The same question, answered badly and well

The gap between these two is most of your score, and it is easier to see than to be told.

The metric moved but the experiment was underpowered. What do you tell the team?

What loses marks

Reporting the lift. It is the answer stakeholders want and it is the reason this question is asked at all.

What scores

Say plainly that the result does not support a confident conclusion, give the range it is consistent with, and state what would settle it — more time, a bigger sample, a different unit of randomisation. The willingness to say 'we cannot know yet' is the mark.

Why Data Scientist candidates get cut

The post-mortem nobody sends you. These are specific to this loop rather than general interview advice.

Model trivia in place of experimental judgement. Most of this loop is causal reasoning, not algorithms.
Producing a confident number from data that cannot support one.
No sense of what the business would do differently depending on the answer.

Start with a general round

For Data Scientist, a general round runs with Marcus metrics, experiments, causal reasoning, and decisions under uncertainty. It is the fastest way to find out which round you actually need to work on. The first session is free.

Practise a Data Scientist interview →

Question guides

Before you practise, it is worth reading how the common rounds are marked: tell me about yourself, behavioral questions and STAR, and system design.

Similar roles

All roles