What this role’s resume is actually judged on — the language postings use, what a reviewer looks for in the first ten seconds, and the specific way AI / LLM Engineer resumes go wrong.
A first pass takes seconds and is looking for three things. If they are not near the top, the rest of the page rarely gets read.
Most resumes fail one bullet at a time. The difference is almost always specificity — a number, a constraint, or a consequence.
Built AI features using GPT-4 and LangChain.
Built the eval harness for our support assistant — 400 graded cases run on every prompt change — which caught a 12-point regression before release and made prompt edits safe.
Terms that recur in real postings for this role. They belong in the bullet that proves them, not in a list at the bottom — our own checker weights requirement coverage at 40% and raw keyword matching at 20%, and most serious systems make a similar trade.
A keyword you cannot defend in an interview is worse than a missing one. How the format checks work.
This is the fastest way to be filtered out for these roles. Anyone can call an API; the scarce skill is knowing whether the output got better or worse. If your resume has prompts but no evals, it reads as a hobbyist.
A newer loop with less settled conventions. Expect evaluation to dominate — how you measure whether a non-deterministic system is working — alongside retrieval design, prompt strategy and cost per request.
AUC where the business metric should be
Listing tools instead of questions answered
No mention of what happens when a pipeline breaks
Research projects presented as production work
Reading as an analyst who happens to write dbt
Listing technologies you have touched once
Generic resume advice only goes so far — what matters is whether this resume covers this posting. Paste both and get requirement-by-requirement coverage, the keywords you are missing, and the format problems a parser will hit. The first check is free.