Data & AI

Machine Learning Engineer mock interview

The loop straddles research and production. Be ready for both the modelling conversation and the deployment one: training/serving skew, drift, and how you would know the model got worse.

Data loops test two things that pull against each other: rigour with methods, and the judgement to know when a rough answer is the right one. Candidates who are strong on one are often visibly weak on the other.

Expect questions where the data is deliberately insufficient. Stating the assumption you are making, and what would change your answer, is most of the mark.

The rounds you can practise

Each round is run by the interviewer built for it, with its own scoring axes — not one generic interviewer asked to change subject.

Metrics, experiments, causal reasoning, and decisions under uncertainty.

Scored on Analytical communication · Data & inference rigor · Decision quality

Metrics, experiments, causal reasoning, and decisions under uncertainty.

Scored on Analytical communication · Data & inference rigor · Decision quality

Hands-on problem solving, code quality, debugging, and engineering judgment.

Scored on Problem-solving communication · Engineering correctness · Adaptability & craft

System design, trade-offs, and war stories from production.

Scored on Technical communication · Engineering depth · Ownership & collaboration

Metrics, experiments, causal reasoning, and decisions under uncertainty.

Scored on Analytical communication · Data & inference rigor · Decision quality

Machine Learning Engineer interview questions

Six you can expect, in the register interviewers actually use. Answer them out loud before you read the next section — reading a question and answering one are different skills, and only the second is marked.

  1. How would you know this model got worse in production?
  2. Walk me through training/serving skew you have actually hit.
  3. How do you decide a model is good enough to ship?
  4. What do you monitor once it is live?
  5. How do you handle a feature that is available at training time and not at inference?
  6. When is the right answer not to use a model?

The same question, answered badly and well

The gap between these two is most of your score, and it is easier to see than to be told.

How would you know this model got worse in production?

What loses marks

"We'd monitor accuracy." In most production systems the label arrives days later or never, so this answer describes something you cannot actually do.

What scores

Proxy signals first — input distribution drift, prediction distribution shift, downstream behaviour — then delayed label evaluation where labels do arrive, and a named threshold with an action attached. The absence of ground truth is the whole problem.

Why Machine Learning Engineer candidates get cut

The post-mortem nobody sends you. These are specific to this loop rather than general interview advice.

Strong on modelling, empty on deployment. The loop straddles both and candidates almost always prepare one.
No answer for the gap between offline metrics and the business outcome.
Treating a model as finished at ship rather than as something that decays.

Start with a general round

For Machine Learning Engineer, a general round runs with Marcus metrics, experiments, causal reasoning, and decisions under uncertainty. It is the fastest way to find out which round you actually need to work on. The first session is free.

Practise a Machine Learning Engineer interview →

Question guides

Before you practise, it is worth reading how the common rounds are marked: tell me about yourself, behavioral questions and STAR, and system design.

Similar roles

All roles