Hiring12 August 202611 min read

How candidates cheat in remote interviews

Almost everything written on this subject is published by companies that sell detection. This is not, and it reaches a different conclusion.

Niraj Patil, Founder

I build AI voice interviewers at MockLive. We sell an interview format, not a detection product — which is why I can afford to write this.

You have probably already had this interview. The candidate is articulate. The answers arrive complete — a definition, two or three considerations, a trade-off, a conclusion. Nothing is wrong with any of it.

Then you ask how they measured the improvement they just described, and the shape of the conversation changes. The number is approximate. The timeframe is vague. The person who owned the dashboard is “someone on the platform team”. Thirty seconds ago they were fluent on distributed caching; now they cannot tell you what the latency was before their change.

That gap is the tell. Not hesitation, not nerves, not a long pause — those are ordinary. The tell is fluency without anchoring: strong general knowledge sitting next to collapsed personal specifics. A language model can produce an excellent answer about caching strategy. It cannot tell you what your candidate’s p99 was in March, because it was not there.

Is that proof they cheated?

No, and this is the first place hiring teams go wrong. An anxious candidate goes vague under pressure. A candidate who genuinely never measured the thing goes vague because there is nothing to recall. Someone three years removed from a project has forgotten the number honestly.

The pattern is a reason to ask another question. It is not a verdict, and treating it as one is how good candidates get rejected for being nervous. Everything useful in this article follows from taking that distinction seriously.

How AI interview cheating tools actually work

These products are openly marketed. Some of them advertise being invisible to screen sharing and undetectable to the interviewer, in those words, on their own homepages. There is no point being coy about a category the reader has already searched for.

The mechanism is the same across the category, and it is worth understanding at exactly one level of detail: the tool captures the question — by transcribing audio or reading the screen — sends it to a language model, and surfaces a suggested answer somewhere a screen share will not pick up. That is the whole architecture. It is not sophisticated, and the sophistication is not where the problem lives.

The problem lives in a dependency. Every tool in this category needs a question that sits still. Long enough to be captured, sent, answered, and read. Ideally written down. Ideally predictable. Ideally with no one watching the candidate’s face while they read.

Look at what that describes. A take-home. An asynchronous one-way video screen where the question appears on screen and the candidate records a reply. A timed coding assessment where the prompt sits in a browser tab for sixty minutes. These formats are not incidentally vulnerable. They are close to the ideal operating conditions for the thing you are worried about.

Now look at what a live conversation does to that dependency. The question is spoken. The follow-up depends on what the candidate said forty seconds ago. There is no static prompt to capture, and the window between question and expected answer is a few seconds during which a human being is watching. The tool has not been beaten — it has been given nothing to work with.

Why detection is a losing game

I want to be fair here, because this is the section where it would be easy not to be. Proctoring and detection vendors ship real engineering. The people building those products are not cynics, the problem they are attacking is genuinely hard, and some of what they catch is real. This is a critique of an approach, not of the companies taking it.

The update cycle is asymmetric

Assistance tools are consumer software with a fast release cadence. When a detection method starts working, the countermeasure is a software update that ships in days. A detection vendor discovers the new behaviour, builds a signal for it, tests it, ships it in a quarterly release, and waits for customers to upgrade.

Nobody in that loop is incompetent. The loop is just structurally unfavourable, in the same way that signature-based antivirus is structurally behind whoever is writing the malware. You can run that race. You will not finish it.

False positives are the expensive failure

This is the part buyers consistently underestimate, and it is arithmetic rather than opinion. Suppose five percent of your candidates use assistance, and suppose you have a detector that is ninety percent accurate in both directions — considerably better than anything on the market claims to be.

Illustrative, per 1,000 candidates. Assumes a 5% base rate and 90% accuracy in both directions.
GroupCountFlagged
Actually used assistance5045
Did not95095
Total flagged140, of which 95 are innocent

Two thirds of your flags are wrong. Not because the detector is bad — it is better than real ones — but because the honest group is nineteen times larger, so even a small error rate against it swamps the true positives. Lower the base rate and it gets worse.

The cost of those ninety-five is not abstract. It is a rejection letter to someone who did nothing, an accusation you cannot substantiate if they ask, and — in jurisdictions with automated decision-making rules — a documentation problem. Every one of them tells people about it.

Heavy monitoring costs you the candidates you wanted

Greenhouse surveyed 2,950 job seekers in May 2026. Among the 1,200 U.S. respondents, 38% had already walked away from a hiring process because it included an AI interview, and another 12% said they would roughly half who have left or would.

Be precise about what that measures. It is about AI interviews rather than proctoring specifically, and self-reported intent overstates behaviour. It does not prove that proctoring costs you half your pipeline. What it does establish is that candidates now leave processes over the screening format itself, at a rate large enough to show up in your funnel — and that the people most able to walk are the ones with other offers.

That is the trade nobody puts in the business case. Every candidate lost to friction is a real cost, paid today, to prevent a hypothetical one.

And rigour is already the expensive part

Outsourced human technical interviews are the highest-integrity option available, and they are priced accordingly — in the hundreds of dollars per completed interview, typically with an annual volume commitment. I am deliberately not quoting a figure for a named vendor: none of them publish public pricing, the third-party estimates disagree by more than twofold, and putting an unverifiable number about a competitor into an article about honesty would be a strange way to make this argument.

The structural point survives without the number. Today you can have rigour by paying a lot of money per interview, or by spending senior engineering time you would rather spend elsewhere. Most teams, reasonably, pick the asynchronous option — and then discover it is the format most exposed to the thing this article is about.

Can you detect AI use in interviews?

Not to a standard you would be comfortable defending to the candidate. Detection infers assistance from indirect signals — where someone is looking, how long they paused, whether focus left the window, whether the prose reads as machine-generated. Every one of those has an innocent explanation. People look away when they think. People with dual monitors look away constantly. Non-native speakers write more formally.

You can raise the cost of casual cheating, and that is genuinely worth doing. What you cannot get is a verdict. If a product implies otherwise, ask what it does when it is wrong, and who tells the candidate.

What actually reduces the problem

None of this requires buying anything, including from us. It works because it attacks the dependency rather than the tool.

Make the first round live

Synchronous beats asynchronous by a wide margin, for the reason set out above: there is no static question to capture. If you keep an asynchronous stage, treat it as a filter for obvious no-hires rather than as evidence of ability.

Follow up on the candidate’s own specific claims

This is the single most effective technique available and it costs nothing. When a candidate says they cut latency by forty percent, ask what it was before. Ask how they knew. Ask what they tried first that did not work. Ask who disagreed.

No assistant can answer for a project it was not part of. The candidate who did the work has more detail than they can fit in the answer; the one who did not runs out immediately. You are not catching a tool — you are asking about something only a participant knows.

Ask for the reasoning, not the answer

“What would you do?” has a good generic answer. “Why did you rule out the other option?” does not. Our system design guide goes into why the trade-off is where the signal is — and that holds whether or not anyone is using assistance, which is the point. Questions that resist assistance are usually just better questions.

Score against a rubric

Two interviewers cutting candidates for different reasons is a bigger measurement problem than cheating is, and it is much more common. A written rubric also gives you something defensible if a decision is ever challenged. The behavioural questions guide sets out what those rounds are actually assessing.

Record the round

Not to surveil anyone — to make decisions reviewable. When someone asks why a candidate was cut, a recording and a timestamped rubric line is an answer. Memory three weeks later is not. Tell candidates you are recording, and why.

Tell candidates what is allowed

Write it down and send it before the interview. If notes are fine and a live model is not, say exactly that. Ambiguity produces more cheating than policy does, because candidates who would happily follow a clear rule end up guessing at an unstated one — and a meaningful number guess in their own favour.

Where we fit

This is the product section. It is one section, near the end, and you can skip it without losing the argument.

We build MockLive Hiring: an AI voice interviewer that runs your first round live, asks follow-up questions based on what the candidate actually said, and returns video, a full transcript, and scoring where every point traces to a rubric line and a timestamp. It exists because the first round is the most repetitive and least consistent part of hiring, not because of cheating.

We make no claim to detect cheating, and we are not going to add one. We think products promising a reliable verdict are overpromising, for the arithmetic reasons above. What a live, conversational format does is remove the static question that assistance tools depend on. That is a structural property, not a detection feature, and it is the only honest claim available here.

The question worth asking instead

Almost every conversation about this starts with “how do we catch them”. That is the wrong end. The better question is why your first round is the kind of thing that can be faked at all.

A screen built around a real conversation, with follow-ups that depend on the previous answer, does not need a cheating verdict — there was never a static question sitting there waiting to be fed to something. You are not in an arms race, because you never entered one.

That is a format decision, and you can make it this quarter without buying anything. Most of what is on this page costs nothing but the willingness to run the first round as a conversation.

Common questions

Can you detect AI use in interviews?

Not reliably, and anyone selling certainty is overpromising. Detection tools infer assistance from indirect signals — eye movement, timing, window focus — and every one of those signals has an innocent explanation. You can raise the cost of cheating, which is worth doing. You cannot get a verdict you would be comfortable defending to the candidate.

What are AI interview cheating tools?

Desktop applications that listen to an interview or read the screen, send the question to a language model, and surface a suggested answer where a screen share will not capture it. They are openly marketed and some advertise being undetectable. They depend on the question sitting still long enough to be captured and answered.

How do you prevent cheating in technical interviews?

Change the format rather than adding surveillance. Live and synchronous beats asynchronous, and following up on the candidate's own specific claims is the single most effective technique available — it costs nothing and no assistant can answer for a project it was not part of.

Does interview proctoring work?

It raises the cost of casual cheating and it catches some people. The problem is the trade: the tools it defends against iterate far faster than proctoring vendors ship, false positives carry real legal and human cost, and heavy monitoring measurably drives candidates away. Whether that trade is worth it depends on your volume and your role mix.

What are the alternatives to interview proctoring?

A live conversation with unscripted follow-ups, a structured rubric so two interviewers cut candidates for the same reasons, and a recording so decisions can be reviewed afterwards. These reduce the problem by removing the static question that assistance tools need, rather than by trying to catch the tool in the act.

Should we tell candidates what AI use is allowed?

Yes, explicitly and in writing before the interview. Ambiguity produces more cheating than policy does, because candidates who would follow a clear rule end up guessing at an unstated one. It also gives you a defensible position if you later need to act.

Sources are linked inline. The Greenhouse figures are quoted from the primary release rather than secondary coverage, because at least one widely-read summary inverted the second number. The false-positive table is an illustration with its assumptions stated, not a measurement of any specific product. Last updated 12 August 2026.