AI-Native Skills Assessment: What to Test and How - IT Staffing - Second Talent
Skip to content

AI-Native Skills Assessment: What to Test and How

Almost every developer uses AI tools now, so the AI-native claim checks nothing. Five dimensions, a 90-minute observed session, and the questions to put to any provider making the claim.

Matt Li By Matt Li 9 min read

TL;DR: Almost every developer uses AI tools now, so a provider claiming AI-native engineers is claiming nothing checkable. The scarce skill is verification: 66 percent of developers name almost-right output as their top frustration and 45 percent say debugging generated code takes longer. Assess across five dimensions in a live 90-minute session, lead with verification, and ask any provider for the written rubric plus what fails it. Free PDF template.

Every provider in this market now advertises AI-native engineers. The phrase has no agreed definition, no accepted test behind it, and adoption figures high enough that it cannot distinguish anyone.

What changed here. An earlier version built its argument on a Gartner AI Skills Framework and a Robert Half hiring quality report, including a claim that only 25 to 30 percent of engineers listing AI tooling demonstrate fluency under observation. Neither source exists and the page carried no external links at all. The version below uses the 2025 Stack Overflow Developer Survey, which measures this subject directly across 33,662 responses, and presents the five dimensions as our own framework, which is what they always were.

Why the claim cannot be verified as stated

The survey settles the adoption question. 84 percent of respondents use or plan to use AI tools, up from 76 percent the year before, and 50.6 percent of professional developers use them daily.

Why the AI-native claim is unverifiable: 84 percent of developers use or plan to use AI tools, so the resume line no longer signals anything

That number does the work here. When adoption approaches universal, asking whether an engineer uses AI tools returns yes from nearly everyone and filters nobody. The resume line has become a default rather than a signal, and screening on it changes almost no outcomes.

Artefacts do not settle it either. A portfolio can be produced with the same tools it is supposed to demonstrate, and what a finished artefact cannot show is why each choice was made. That is the part worth assessing and the part that has to be observed rather than submitted.

The more useful finding is where developers say the difficulty actually sits. 66 percent name “AI solutions that are almost right, but not quite” as their biggest frustration, and 45.2 percent report that debugging AI-generated code is more time-consuming than expected. On trust, only 3.1 percent highly trust the accuracy of what these tools produce while 45.7 percent actively distrust it against 32.7 percent who trust it, and developers with ten or more years of experience are the most sceptical of all.

Read together, those figures describe a market where using the tools is table stakes and judging their output is the work. An assessment that tests the first thing measures something everyone has.

The five dimensions

These are our framework rather than an industry standard, ordered by how much each one predicts. Verification leads because the survey data points there.

Five dimensions of AI-native engineering skill: verification of output, iteration, eval writing, tool fluency and cost and latency reasoning

Verification of output is what an engineer does with an answer that looks right. Whether they check before building on it, what check they choose, and whether they can say what a wrong version would look like. If you assess one dimension, assess this one.

Iteration is what happens when the first answer fails. A weak engineer re-runs the same prompt or abandons the tool. A strong one names the failure mode, changes one thing, and can explain why the output moved, which is the difference between iterating and shuffling.

Eval writing is whether they can test a feature backed by a model at all. Reference inputs and expected outputs, deterministic checks where the answer is checkable, a judged rubric where it is not, and some way of noticing drift. It is the dimension most predictive of production-grade work and the one most engineers have never done.

Tool fluency across more than one tool matters less than it used to but has not stopped mattering. What is worth watching is deliberate switching and knowing each tool’s failure mode, rather than the number of logos.

Cost and latency reasoning separates demos from systems. Which model for which job, what a call costs at volume, where a smaller model is sufficient, and where latency becomes user-visible.

Level 1 versus level 4 criteria across the five AI-native assessment dimensions

Score each dimension from 1 to 4. The endpoints above are easier to recognise than the middle, and once you have seen both ends on a few candidates the middle places itself.

Running the session

Ninety minutes, live, screen-shared, on a problem the candidate has not seen. Everything in the structure below is chosen to be difficult to prepare for.

A 90-minute observed AI-native skills session, from warm-up on the candidate own setup through to cost and latency trade-offs

Start on their own setup, in their own editor with their own configuration. What is already configured tells you more than any question about configuration would, and it settles the tool-fluency dimension in the first ten minutes without asking about it.

The middle section is a small build in a domain they have not worked in, with narration. Unfamiliarity is doing the work: in a familiar domain a candidate can carry the task on prior knowledge, and you learn nothing about how they operate when the tool is the only thing helping.

The highest-signal part is the segment where the obvious generated answer is plausible and wrong. Steer toward a task with that property and then watch. Some candidates check reflexively, some check when prompted, and some build three steps on top of it. That single observation predicts more than the rest of the session combined, and it is the exact failure mode 66 percent of developers say frustrates them most.

Close on evals and then on cost. Both are conversations rather than exercises, and both are hard to bluff because the follow-up question is always “for this problem specifically”.

Making preparation not work

The goal is not catching anyone out. It is ensuring the session measures what it claims to, which means designing around the things a prepared candidate can produce in advance.

What a prepared candidate can fake in an AI skills assessment versus what has to be produced live

A rehearsed narration, a familiar take-home, and fluent terminology are all obtainable without the underlying skill. Handling a wrong answer in an unfamiliar domain, in real time, with someone watching, is not.

Two practical notes. Do not use a well-known problem, because solutions to well-known problems are exactly what the tools have memorised. And do not treat a take-home as evidence of anything beyond willingness to do it, given that the survey found 75.3 percent of developers say the main reason they would ask another person for help is when they do not trust the AI’s answer. That instinct is the skill, and a take-home cannot see it.

What to ask a provider

Most buyers will not run this session themselves. The alternative is testing whether the provider runs something like it.

Four questions to ask a staffing provider claiming AI-native engineers, including sending the written rubric and describing what fails it

Ask for the written rubric assessors actually use, with the levels defined. Then ask them to describe what a candidate does that fails it. The second question is the one that matters, because a marketing page can describe success and only a real rubric describes failure.

Ask how assessors are calibrated against each other, and what happens when two reviewers disagree on the same session. Without an answer, the score means whatever that reviewer meant that day.

Then ask what share of candidates fail at this stage. If nearly everyone passes, the stage is decorative. Our checklist on evaluating IT staffing companies covers the funnel questions this sits inside, and the 15 questions to ask any provider covers the rest of the conversation.

Where a reference is available, ask them what they observed rather than what they were told. Our guide to reference checks that work covers how to get an answer with an example attached.

Where this fits in the wider assessment

AI-native skill is one dimension of hiring an engineer, not a replacement for the rest of it. The same session should still tell you whether they can read an unfamiliar codebase, reason about a system, and explain a decision to someone who disagrees.

It also does not change the commercial questions. Whichever legal entity employs the engineer still determines where classification exposure sits, and the IRS common-law test weighs behavioural control regardless of how the work gets done. Our guide to worker classification in cross-border IT staffing covers that side.

For a cost baseline, the US Bureau of Labor Statistics puts the median wage for software developers at $135,980 as of May 2025, salary before employer taxes and benefits. There is no published premium for AI-native skill specifically, and any provider quoting one is quoting themselves.

The dimension most assessments skip

One thing worth testing that rarely appears on anyone list: what the candidate does with data. Working with these tools means deciding, repeatedly, what goes into a prompt, and an engineer who pastes production records into a hosted model has created a problem that no amount of code quality offsets.

Ask directly during the session. What would you not paste in here, and why. A strong answer distinguishes between a local model, a zero-retention endpoint and a consumer chat window without being prompted. Where personal data is involved, GDPR Article 28 makes the processing terms prescribed rather than negotiable, and that constraint reaches the engineer at the keyboard rather than stopping at the contract.

The same instinct shows up in how they think about failure. The NIST AI Risk Management Framework organises this around governing, mapping, measuring and managing risk, and an engineer who has thought about where a model-backed feature can fail in production will describe something recognisably similar without using the vocabulary.

Common questions

Is ninety minutes really necessary?

For a senior seat where this skill is load-bearing, yes. The verification segment does not work if it is rushed, because the candidate has to get far enough into a task for a wrong answer to matter. Sixty minutes works if you drop the cost conversation.

Can this be assessed asynchronously?

Partly. A recorded session where the candidate narrates is second best and still useful. A written submission is not, since the whole framework rests on watching decisions get made.

What if the candidate uses a tool we do not use?

Let them. You are assessing how they work, not standardising their setup, and an engineer forced into an unfamiliar editor will underperform for reasons that tell you nothing.

Does this apply to non-engineering roles?

The dimensions transfer with the exercises changed. Verification, iteration and cost reasoning apply to anyone using these tools in production work. Eval writing is specific to building features backed by a model.

How often should the rubric be revised?

Roughly every six months, and more often on the tool-fluency dimension. Adoption moved eight points in a single survey cycle, so a rubric written two years ago is testing a different market.

Takeaways

  • 84 percent of developers use or plan to use these tools. Tool use differentiates nobody.
  • Verification is the scarce skill. 66 percent name almost-right output as their top frustration.
  • Observe live on an unfamiliar problem. Artefacts and take-homes cannot show reasoning.
  • Plant one plausible wrong answer. It is the highest-signal fifteen minutes available.
  • Ask any provider for the rubric and for what fails it. The second half is the test.

Where to go next

Download the assessment template (PDF). It is a printable interviewer sheet with the session structure, example tasks, the five-dimension rubric with level criteria, and a scorecard. No sign-up required.

Second Talent assesses engineers on these dimensions and will send you the rubric, including what fails it, before you talk to anyone. Tell us what you are hiring for.

Hire senior engineers on the Second Talent platform

Browse, shortlist, and hire pre-vetted AI-Native Talent across Asia, all in one platform. Free to start, $0 upfront.

Try for Free

Written by

Matt Li is a tech-driven entrepreneur with deep expertise in global talent strategy, digital experience optimization, e-commerce, and Web3 innovation. He is the Co-Founder of Second Talent, a US-based company that connects businesses with top-tier tech professionals worldwide. Since launching the company in 2024, Matt has led its growth by leveraging technology to streamline remote hiring and scale distributed teams. With a background spanning product, operations, and innovation, Matt brings a cross-disciplinary perspective to the evolving digital economy. His work sits at the intersection of global talent, emerging technology, and scalable digital transformation.

More posts by Matt Li →
WhatsApp