How to Vet an AI/ML Engineer

September 8, 2026

(2 min read)

Vetting a senior AI or ML engineer is harder than it looks, because the standard process produces false positives by design. The CV reads beautifully, the technical screen goes well, the references are warm, and six weeks into the engagement you discover the candidate has only ever worked at notebook scale: impressive models, none of which survived contact with production traffic, drift, or a compliance review. This guide covers why generic vetting misses that, what a structured assessment of an AI/ML engineer actually has to test, and the reference questions that surface what interviews cannot.

Why generic technical screens fail for this role

An algorithm-and-coding screen tests what AI/ML work has in common with software engineering, which is the part that rarely sinks the hire. What sinks it is specific to the discipline. Model performance is claimed on the candidate’s terms: a leaderboard score on a clean dataset with a fixed metric proves skill under conditions production never offers. The gap between a working prototype and an operated system is wider in ML than anywhere else in software, because the system degrades silently as the world drifts away from the training data. And the cost of a wrong hire lands late: for most roles a mis-hire is visible in weeks, while a notebook-scale ML engineer produces plausible artefacts, experiments, decks, promising prototypes, for one or two quarters before the absence of anything in production becomes undeniable.

So the vetting question is never “can this person build a model”. It is “has this person operated one, and can they carry accountability for what it does”.

What to assess: four categories, not one screen

Mahala scores every AI/ML candidate against four categories, with 75 out of 100 as the pass mark. The categories translate directly to an in-house process, so they are worth stealing even if you never send us a brief.

Technical acumen. Discipline-specific depth, assessed on the candidate’s actual stack rather than puzzle questions: model architecture choices and their trade-offs, evaluation design (what metric, why, and what it fails to capture), feature and data handling including leakage, and serving: latency, batching, and what the model costs to run per thousand predictions. A senior answers these with numbers from systems they have operated.

Demonstrated impact. Evidence that shipped systems produced named outcomes: what ran in production, for how long, with the candidate accountable for what part, and what changed in the business numbers. Claims here are verified against references, not taken from the CV; the phrase “contributed to” is where inflated ML CVs hide.

Consulting aptitude. Whether the candidate can scope an ambiguous request, defend a trade-off in front of stakeholders, explain a model’s behaviour to a risk officer, and say “the data cannot support that” to someone senior. In enterprise and regulated settings this category carries as much predictive weight as the technical one, and it is the one generic processes never test.

Professional growth. Whether the candidate’s practice has kept pace with a field that reinvents itself every two years: what they have adopted, what they have deliberately not adopted and why. The “why not” answer separates practitioners with judgement from enthusiasts with a feed.

The failure the four-category structure is designed to catch is the single-spike candidate: brilliant in one category, absent in another, carried through a one-dimensional process by the spike.

Reference calls that surface signal

References for ML roles are usually run as character checks, which wastes them. Three questions turn a reference call into evidence.

“Tell me about a model of theirs that failed in production.” Every real practitioner has one. If the reference cannot name one, either the person never operated at production scale or the reference only saw the demos. Follow with what the candidate did in the incident: the answer describes their operational discipline better than any interview.

“How did they handle disagreement with the product team?” ML work generates a specific conflict: the model says one thing, the roadmap wants another. You are listening for whether the candidate communicated uncertainty honestly or told the room what it wanted to hear.

“What governance or compliance constraint blocked them, and what did they do?” In regulated environments this is the daily texture of the job. A candidate who has never met the constraint will meet it for the first time on your budget.

Why six out of seven do not clear the bar

At Mahala, one candidate in seven clears the four categories. The distribution of failures is instructive for anyone vetting in-house. The largest group fails on demonstrated impact: a portfolio of proofs of concept, hackathon wins, and coursework, with nothing that ran unattended in production. The next group is technically strong and fails on consulting aptitude: they cannot scope a vague request or explain a decision without jargon, which in an enterprise team converts to stalled work and frustrated stakeholders. A smaller group fails the reference cross-check: the systems were real, the candidate’s role in them was smaller than the CV said. Almost nobody fails on raw technical knowledge alone, which is precisely why a process that only tests technical knowledge approves almost everybody.

Running this in-house, honestly costed

Nothing above requires an agency. It requires three things most teams are short of. A senior ML practitioner to run the technical category, because scoring an answer about evaluation design takes someone who has designed evaluations. Written criteria fixed before the first interview, so the polished candidate cannot move the goalposts mid-process. And the discipline to verify impact claims with references who worked downstream of the system, at roughly an hour per candidate of senior time, multiplied by every candidate who reaches that stage.

That arithmetic is the actual case for a specialist network: the vetting is done before you see anyone, and what arrives is two to three blind CVs of candidates who already cleared the bar, with the evidence documented. The full framework is on our vetting page, and the specialists it produces are at hire AI engineers.

Frequently asked questions

Do certifications count for anything in AI/ML vetting?

As evidence of baseline knowledge, mildly. As evidence of production capability, no. A cloud ML certification proves the candidate can pass the vendor’s exam; it says nothing about how they behave when a deployed model degrades on a Friday. Weight certifications below any single verified production story.

Take-home assignment or live technical interview?

A timeboxed take-home on a realistically messy dataset, reviewed in a follow-up conversation, beats live coding for this role: you are hiring judgement exercised over hours, not composure under observation. The follow-up conversation is where the vetting happens; anyone can outsource a take-home, few can defend one they did not do.

How long should a proper vetting process take per candidate?

With written criteria and prepared references, roughly four to six hours of senior time: technical assessment, consulting-fit conversation, and reference calls. Teams that find that unaffordable per candidate have discovered why shortlists that arrive pre-vetted exist.

Start from candidates who already cleared it

If there is an AI/ML seat open next to a live roadmap, request a vetted shortlist of AI/ML engineers and interview two to three candidates whose production claims have already been checked, within 72 hours of the brief. Building the process yourself and want a second pair of eyes on it? Book a call; we will walk you through what our assessment tests and why.

Vetting done before you say hello.

Want candidates who already cleared all four categories? Request a vetted shortlist and meet two or three vetted AI/ML Engineers in 72 hours.

Featured Articles

Data engineer vs data scientist

2 min read

Data engineer vs data scientist: which do you need?

Data engineer and data scientist sound similar and are constantly confused, including by the recruiters hiring for them. But they are fundamentally different roles, and hiring the wrong one for your problem means it gets solved slowly, expensively, or not at all.
What Is MLOps? A Practical Guide for Enterprises

2 min read

What is MLOps? A practical guide for enterprises

MLOps is the discipline of running machine-learning models in production reliably. If DevOps is what keeps your software alive, MLOps is what keeps your models alive.
Hiring remote specialists in the UAE

2 min read

Hiring remote specialists in the UAE: the efficient default, not the compromise

Hiring senior AI talent in the UAE and across the Gulf is a different problem from hiring in Europe or North America. The ambition is national policy, the procurement is rigorous, and the local pool of specialists.

Request a Vetted Shortlist

About You
About the role
About the engagement

Book a Call