Data & AI

AI / LLM Engineer mock interview

A newer loop with less settled conventions. Expect evaluation to dominate — how you measure whether a non-deterministic system is working — alongside retrieval design, prompt strategy and cost per request.

Also advertised as AI Engineer, LLM Engineer, GenAI Engineer — the loop is the same, and so is the practice below.

Data loops test two things that pull against each other: rigour with methods, and the judgement to know when a rough answer is the right one. Candidates who are strong on one are often visibly weak on the other.

Expect questions where the data is deliberately insufficient. Stating the assumption you are making, and what would change your answer, is most of the mark.

The rounds you can practise

Each round is run by the interviewer built for it, with its own scoring axes — not one generic interviewer asked to change subject.

Metrics, experiments, causal reasoning, and decisions under uncertainty.

Scored on Analytical communication · Data & inference rigor · Decision quality

Metrics, experiments, causal reasoning, and decisions under uncertainty.

Scored on Analytical communication · Data & inference rigor · Decision quality

Hands-on problem solving, code quality, debugging, and engineering judgment.

Scored on Problem-solving communication · Engineering correctness · Adaptability & craft

System design, trade-offs, and war stories from production.

Scored on Technical communication · Engineering depth · Ownership & collaboration

Metrics, experiments, causal reasoning, and decisions under uncertainty.

Scored on Analytical communication · Data & inference rigor · Decision quality

AI / LLM Engineer interview questions

Six you can expect, in the register interviewers actually use. Answer them out loud before you read the next section — reading a question and answering one are different skills, and only the second is marked.

  1. How do you evaluate a system whose output is different every time?
  2. Walk me through your retrieval design and where it fails.
  3. How do you keep cost per request under control as usage grows?
  4. What is your fallback when the model returns something unusable?
  5. How do you stop a prompt change from silently regressing something else?
  6. When would you fine-tune rather than change the prompt or the context?

The same question, answered badly and well

The gap between these two is most of your score, and it is easier to see than to be told.

How do you evaluate a system whose output is different every time?

What loses marks

"We check the outputs look good." Manual spot-checking does not scale past a handful of cases and cannot catch a regression.

What scores

A held-out set with graded cases, a scoring method you can defend — deterministic checks where possible, model-graded where not, human-labelled on a sample — run on every prompt or model change. Evaluation is the discipline of this role and the loop is mostly about it.

Why AI / LLM Engineer candidates get cut

The post-mortem nobody sends you. These are specific to this loop rather than general interview advice.

No evaluation story. It is the single most common reason candidates fail this loop.
Prompt engineering as the entire answer, with nothing on retrieval quality, latency or cost.
No plan for the model being confidently wrong in front of a user.

Start with a general round

For AI / LLM Engineer, a general round runs with Marcus metrics, experiments, causal reasoning, and decisions under uncertainty. It is the fastest way to find out which round you actually need to work on. The first session is free.

Practise a AI / LLM Engineer interview →

Question guides

Before you practise, it is worth reading how the common rounds are marked: tell me about yourself, behavioral questions and STAR, and system design.

Similar roles

All roles