How to Assess Developers in the AI Era
Test Domain Skill and AI Skill Separately
Summary
To hire well when everyone uses AI, measure a developer's unaided skill and their skilled use of AI separately, so a passing score tells you what the candidate can do and not what the AI did for them.
Imagine that two candidates take your coding test and both are allowed to use AI the way they would on the job. Both submit working solutions and score 90%.
One is a strong engineer who used AI to move faster and checked everything it produced. The other prompted their way into code they don't understand and couldn't debug if it broke.
If your test rated them identically, you have no way of knowing which one deserves a job offer. The problem here isn't cheating, and more proctoring won't fix it.
Key points
- A test that allows AI everywhere measures the wrong thing: when the same task blends unaided skill and AI usage, a strong engineer and a lucky prompter can get the same score, and the test result gives you no way to tell them apart.
- The fix is separation, and the unaided half must be enforced: measure what the candidate can do without AI in a locked-down test, so that score is the candidate's own. Then measure how well the candidate directs and corrects AI, and score that separately.
- Vendors mostly pick one of two camps: block AI harder, or lean all the way in and grade one blended result. Splitting the two scores clean is the third path, and the only path that tells you who did the work.
On this page
- Why can't a coding test tell a strong engineer from a lucky prompter?
- Should you just let candidates use AI on a coding test?
- How should you assess developers who use AI?
- If the job uses AI, why measure unaided skill at all?
- How do you test unaided skill and AI skill separately?
- How do other assessment vendors handle AI on the test?
- AI-era assessment FAQ
Why can't a coding test tell a strong engineer from a lucky prompter?
When candidates can use AI, a passing score doesn't tell you who did the work. An engineer who checks AI's output and a candidate who blindly accepts it can submit the same correct answer, because AI is most convincingly wrong exactly where the candidate is too weak to notice.
Think about the last time you used AI on something you know well. The AI produced something that had a mistake. You glanced at the output, spotted what was off, corrected it in seconds, and moved on. The AI just needed a small nudge. Now imagine using AI on something you don't know well. The AI produces an answer with the same confidence and you have no way to tell if that answer is right or if there is something incorrect in it.
That asymmetry is the trap:
- Within your skill level, you catch AI's mistakes quickly. You know what is wrong and you fix it fast enough for AI to be a reliable partner that just needs some supervision.
- Above your skill level, AI is just as confident, but you can't see the errors. When the output happens to be correct, you got lucky, but luck doesn't carry over to the next unfamiliar problem.
Both situations look identical to someone from the outside, but only the first one is competence. On the test result you can't tell which candidate got lucky and which candidate earned the score.
You can't proctor that flaw away. AI's failures are invisible in the one place they matter most, above the candidate's actual skill level. No amount of watching candidates work will surface a mistake they can't see themselves.
Should you just let candidates use AI on a coding test?
Letting candidates use AI on every question feels modern, but it destroys the signal you run the test for. One AI-allowed task measures the candidate and the AI at once, and afterward you cannot tell which one earned the score.
The popular answer is to lean in. Everyone uses AI on the job now, the reasoning goes, so let candidates use it on the test too.
When one task lets the candidate use AI freely, it measures two different things at once: what the candidate can do, and what the AI did for them. Blind trust in a correct answer and genuine verification of that answer look the same in the result, and you can't tell which one you actually measured.
Blindly accepting AI's output, or accepting it for the wrong reasons, should count against a candidate. But you only capture that failure if the test is built to make it visible. If you give the candidate an AI assistant and grade only the total test score, then you've built a test that hides that failure.
How should you assess developers who use AI?
Measure two things separately: what the candidate can do unaided, and how well the candidate directs and corrects AI. If you use only one test, keep the two sections separate inside the test, so that a passing unaided score proves real ability and you compare the AI score against that baseline.
If measuring both skills in the same task can't separate judgment from luck, stop trying to. Measure each skill independently, on its own terms.
Unaided skill: can this person solve the problem with their own understanding, without AI? That score is your baseline. Without a baseline, you have nothing to compare AI-assisted work against.
AI skills: can this person direct, question, and correct AI when working with it? That skill is real, and is a growing part of the job.
Measured separately, each score tells you something you can trust. Your overall read of the candidate is then built from two clean signals instead of one blur.
If the job uses AI, why measure unaided skill at all?
Because "the job uses AI" and "you can skip measuring unaided skill" are different claims, and the second doesn't follow from the first. Unaided skill is the baseline that gives the AI score its meaning. Without that baseline, a candidate who got a right answer by luck reads the same as a candidate who actually knew the answer.
Most developers now work alongside AI for most of the day. That shift is the reason you still need a clean read on what a developer can do without AI. Take that baseline away and you're back to the two candidates who both scored 90%, with no way to tell them apart. The unaided task works best as a job-relevant work-sample that resembles the real job role rather than a contrived puzzle.
How do you test unaided skill and AI skill separately?
The unaided half has to be enforced. A locked-down test keeps disallowed AI assistants from running on the candidate's machine, so the baseline score is real. The AI half gets its own hands-on tasks, from prompt engineering to working with AI agents, scored apart from the unaided questions.
An honor system gives you a guess where you need a baseline. TestDome enforces the unaided half with three layers of AI proctoring, included with every pack, with no annual contract or per-seat lock-in. Safe Exam Browser locks the candidate's machine and prevents disallowed tools from running on it. The blocked tools include invisible overlay assistants and hidden windows that only the candidate can see. AI-analysed screen recording watches how the candidate works within the tools they are allowed to use, and flags the use of tools such as ChatGPT or DevTools. AI-analysed webcam proctoring looks for outside help: another person, a second device, or a phone off to the side. It runs on a single camera, so it catches a visible device or an off-screen glance but not a device held fully out of frame.
The AI-skills test covers the work developers actually do with AI. TestDome offers a dedicated library of AI-skills tests, separate from its AI-free coding tests. The tests span AI literacy, prompt engineering, and working with AI agents. For developers who build AI into their own products, the library adds LLM engineering, machine learning, and deep learning. The tasks are hands-on work samples. There's no multiple-choice trivia about AI.
You can build both halves for your own roles with a free TestDome trial, or start from a ready-made test in the library.
How do other assessment vendors handle AI on the test?
Most vendors pick one of two camps: block AI harder with more lockdown, or lean in and grade AI-collaboration inside one combined task. TestDome takes a third path and measures unaided skill and AI skill apart, so neither score contaminates the other.
In the lean-in camp, CodeSignal markets AI-assisted and agentic assessments built to measure "how candidates and employees work with AI in role-relevant scenarios." In that format the candidate works through a task with an AI assistant present, which raises an integrity question of its own: if the candidate is meant to complete the task alongside AI, it is hard to determine what the candidate can do without AI. TestGorilla's one-way, AI-led video interviews are a different thing. There the AI sits on the employer's side, conducting the interview, and the integrity question is about proctoring, because the candidate records answers on their own time outside any controlled environment. Both camps are reasonable responses, and most serious vendors now assess AI skill in some form.
Keeping the two scores clean instead of blending them is what lets you answer the question you started with. How each vendor handles the split is covered in the head-to-head guides: TestDome vs. Codility, vs. HackerRank, and vs. TestGorilla. Of the two candidates who both scored 90%, which one are you actually about to hire?
AI-era assessment FAQ
Should candidates use AI during a coding assessment? The answer depends on the question type. Block AI on the questions that measure unaided skill, and allow it on the questions that measure AI skills, where the task calls for working with it.
Can one test measure both coding skill and AI skill? Yes, when the two skills are scored separately inside the test. Run the unaided questions under lockdown, so AI assistants can't run on the candidate's machine. The AI-skills questions build the AI interaction into the task itself, so both halves can run under the same proctoring and be scored separately.
Is any coding test "AI-proof"? No format is AI-proof, and the AI-cheating guide linked below covers what to do about that. A work-sample test is valuable because it resembles the real job role and predicts work performance. That value does not depend on resisting AI.
Does measuring two skills make the test longer? Not much. The candidate still takes one test, with the two skills covered by separate questions inside it. Keep the test short and job-relevant to protect completion rates and candidate experience.
If you're fighting active AI cheating right now rather than redesigning how you measure skill, start with how to handle AI cheating in a coding assessment.
Sources and review notes
We last reviewed this page against the public URLs below. If procurement or legal needs proof, save screenshots from those pages on the date you rely on them.
References
- Schmidt, F. L., & Hunter, J. E. (1998). The validity and utility of selection methods in personnel psychology: Practical and theoretical implications of 85 years of research findings. Psychological Bulletin, 124(2). https://doi.org/10.1037/0033-2909.124.2.262
- Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11). https://doi.org/10.1037/apl0000994
How we checked this page: every figure above was read from the vendor's own public page on the date shown. Where a detail is uncertain or still rolling out, we say so in the text.