Engineering

DevOps / SRE mock interview

This loop is dominated by incidents. Expect to walk through a real outage — detection, mitigation, root cause, and what you changed so it could not recur — with the interviewer probing every step you skipped.

Also advertised as DevOps Engineer, Site Reliability Engineer — the loop is the same, and so is the practice below.

Engineering loops are the most structured of any discipline, and the most predictable. Expect a recruiter screen, one or two technical rounds, a system design conversation at mid-level and above, and a behavioral round that is usually underestimated.

The technical rounds are rarely won on the answer. They are won on narration — whether the interviewer can follow your reasoning while you work, and whether you say the trade-off out loud instead of silently picking one.

The rounds you can practise

Each round is run by the interviewer built for it, with its own scoring axes — not one generic interviewer asked to change subject.

System design, trade-offs, and war stories from production.

Scored on Technical communication · Engineering depth · Ownership & collaboration

System design, trade-offs, and war stories from production.

Scored on Technical communication · Engineering depth · Ownership & collaboration

Hands-on problem solving, code quality, debugging, and engineering judgment.

Scored on Problem-solving communication · Engineering correctness · Adaptability & craft

System design, trade-offs, and war stories from production.

Scored on Technical communication · Engineering depth · Ownership & collaboration

Technical Deep-Dive

Theo · Principal Engineer

Hands-on problem solving, code quality, debugging, and engineering judgment.

Scored on Problem-solving communication · Engineering correctness · Adaptability & craft

DevOps / SRE interview questions

Six you can expect, in the register interviewers actually use. Answer them out loud before you read the next section — reading a question and answering one are different skills, and only the second is marked.

  1. Walk me through your worst incident, start to finish.
  2. How did you know it was happening before a customer told you?
  3. What is the difference between an alert that pages and one that does not?
  4. How would you roll out a risky change to a service with no maintenance window?
  5. What does your ideal post-mortem contain?
  6. How do you decide an SLO, and what happens when you burn the budget?

The same question, answered badly and well

The gap between these two is most of your score, and it is easier to see than to be told.

Walk me through your worst incident, start to finish.

What loses marks

The root cause and the fix, told as a technical anecdote. It answers a smaller question than the one asked.

What scores

Detection, mitigation, root cause, prevention — in that order, with the times attached and the part where you mitigated before you understood. Interviewers push hardest on the step candidates skip, which is almost always detection.

Why DevOps / SRE candidates get cut

The post-mortem nobody sends you. These are specific to this loop rather than general interview advice.

Conflating mitigation with fixing. Restoring service and finding the cause are different jobs and doing them in the wrong order costs users.
Post-mortems with a person in the root cause line — blameless is not a nicety, it is whether your incidents produce information.
Alerting philosophy that amounts to alert on everything, which is indistinguishable from alerting on nothing.

Start with a general round

For DevOps / SRE, a general round runs with Kai system design, trade-offs, and war stories from production. It is the fastest way to find out which round you actually need to work on. The first session is free.

Practise a DevOps / SRE interview →

Question guides

Before you practise, it is worth reading how the common rounds are marked: tell me about yourself, behavioral questions and STAR, and system design.

Similar roles

All roles