The demand for engineers who can put AI systems into production has outpaced supply for several years, and the gap is not closing. That has produced a hiring market where almost every candidate can discuss transformers, retrieval architectures and vector databases fluently, and where fluency has consequently stopped carrying information.
This creates a specific problem for hiring managers. The screen that worked when the field was small — test whether they understand the technology — now passes almost everyone. Teams then discover the difference between candidates six months into an engagement, at considerable cost.
The screen most teams run
A typical loop tests three things: familiarity with model architectures and current tooling, ability to write reasonable code under observation, and general problem-solving. These are all worth knowing and none of them distinguishes a candidate who has operated an AI system in production from one who has built impressive demonstrations.
The reason is that production AI work is dominated by a different class of problem. Very little of it is modelling. Most of it is evaluation design, data quality, failure handling, cost management and the judgement to recognise when a statistical approach is the wrong tool. A candidate can be genuinely strong on architecture and have never once had to explain to a business owner why the system was confidently wrong last Tuesday.
Three signals that actually predict
Across a lot of technical hiring, three things have consistently distinguished people who ship from people who demonstrate.
They talk about evaluation before they talk about models
Ask how someone would approach a new AI problem. Weaker candidates start with architecture — which model, which framework, which retrieval strategy. Stronger candidates start by asking how success will be measured and who decides. That instinct is nearly impossible to fake and it correlates with production experience more reliably than any other single signal, because anyone who has shipped has been burned by an unmeasurable objective.
They have specific, unflattering failure stories
Ask about something that went wrong. The useful answer is specific, technically detailed, and includes the candidate's own misjudgement. Answers that stay abstract, or where every failure was caused by someone else's decision, indicate either limited exposure or limited reflection. Both matter.
The best answers are usually mundane rather than dramatic — a schema change that silently degraded retrieval for weeks, an evaluation set that was accidentally in the training data, a cost curve nobody modelled. Mundane failures are the texture of real operational experience.
They will argue against the work
Describe a use case with a real flaw in it and see whether they push back. Candidates who have owned production systems have learned that shipping the wrong thing is worse than shipping nothing, and they will say so to a hiring manager. Candidates optimising for agreeableness will find a way to make your idea work. In a delivery context that is a liability, because they will do the same thing to a client.
The most valuable thing a delivery engineer can tell a client is that the thing they asked for is a bad idea. Someone who will not do it in an interview will not do it in month four.
What credentials do and do not tell you
Certifications and course completions in this field have become close to uninformative for senior roles. They confirm exposure to material, which is table stakes, and they are frequently a substitute for the operational experience that actually matters.
Two credential signals do carry weight. Sustained open-source contribution to infrastructure — not example notebooks, but tooling other people depend on — indicates someone who has dealt with other people's edge cases. And depth in an adjacent discipline, particularly data engineering, distributed systems or quality engineering, is often a stronger predictor than an AI-specific credential, because those disciplines already teach the operational habits that AI work requires.
The most consistently undervalued background is quality engineering. Engineers from a strong QA or SRE tradition arrive already thinking about failure modes, test design and observability, which is most of what an AI system needs and most of what teams from a pure research background have to learn on the job.
Designing a loop that detects the right thing
Three changes usually surface the signals above without lengthening the process.
- Replace the architecture question with an evaluation exercise. Give a realistic use case and ask the candidate to design how they would measure it, including what they would refuse to ship without. This is more diagnostic than any system-design whiteboard.
- Make one interview a debugging session on a genuinely broken system. Not a puzzle — a plausible mess with a boring cause. How someone narrows the search space is far more informative than whether they find it.
- Have someone senior present a flawed plan and score the pushback. Explicitly evaluate whether the candidate identified the flaw and whether they raised it clearly. Tell them afterwards that this was the test.
The uncomfortable implication
Loops built this way reject candidates who look excellent on paper and pass some who look unremarkable. That is difficult to defend internally, especially when a hiring manager is under pressure to fill a role and a well-credentialled candidate is available.
It is still the right trade. The cost of a wrong senior hire in a delivery organisation is not the salary — it is the engagement that goes sideways, the client trust that does not come back, and the twelve months before anyone admits the problem. Measured against that, an extra two weeks of search is cheap.
Why this sits on our site. We staff the pods we run. Every engineer we place is someone we would put on our own most difficult engagement, and this is the screen we apply. If you are building an internal loop, take any of it that is useful.
Keep reading
More from the team
Why AI pilots stall: it is the operating model, not the model
The pilot worked. It was demoed to the board. Eighteen months later it is still a pilot. That failure has a consistent shape, and it is almost never technical.
What an AI-native global capability centre actually requires
Most GCC business cases are still built on arbitrage. That case is weakening, and the centres being set up on it now will struggle within three years.
Ready when you are
Want to argue with any of this?
We would rather have the disagreement than the polite nod. Bring your situation and we will tell you where this framing does not apply.