How Do You Hire Someone Whose Work You Cannot Check?

You cannot evaluate the code, so stop trying. Judge the three things that do not require technical skill: how the candidate scopes a vague problem, how they describe failure, and whether their past work matches your volume. Borrow technical judgment for the rest, deliberately and from someone with no stake in the outcome.

The uncomfortable part of this situation is usually stated wrongly. The problem is not that you lack technical knowledge. It is that you are about to make a decision where the person being assessed knows more about the subject than everyone else in the room, and both of you know it.

That asymmetry does not go away by reading about transformers the weekend before. It goes away by changing what you assess.

Most hiring advice for AI roles is written for people who can read a take-home submission and form a view. If you can, use it: the AI engineer skills checklist and structured technical screens exist for that reader. This article is for the founder, operations lead, or agency owner who cannot, and who is going to hire anyway because the work is not going to wait.

Before going further, one prior question deserves an answer, because it disqualifies a large share of the people who ask this one. An AI hire only makes sense when a defined process is failing at a volume a person cannot absorb. If the problem is a goal rather than a process, the hire will produce something plausible and it will be judged against expectations nobody wrote down. That case is argued in full in whether an AI engineer is worth the cost at all. Read it first if you are not certain.

What Can You Judge Without Technical Skill?

Three things, and they are more predictive than most technical screens.

How they narrow the problem. Describe your situation in the vague terms you would naturally use, and say nothing more. A strong candidate will start asking questions: what does the current process look like, how many items a day, who does it now, what does a wrong answer cost you. A weak one will start describing a solution. This is not a personality observation. Narrowing scope under ambiguity is the single most load-bearing skill in applied AI work, and it is fully visible to a non-technical listener.

How they describe something that failed. Ask what went wrong on their last project. You are listening for specificity and ownership, not for the technical content. "The model underperformed" is a non-answer. "We built a classifier for support tickets, it worked in testing, and it collapsed in production because the training data came from a period when the ticket categories were different" is an answer. You do not need to understand classifiers to hear the difference between those two sentences.

Whether their past work ran at your volume. Systems behave differently at ten items a day and ten thousand. Ask for the numbers around their previous work rather than the work itself: how many items it handled, how quality was measured, what the error rate was, who used it afterwards, and whether it is still running. Someone who was genuinely close to production has these numbers. Someone who built a demo does not, and the gap shows immediately.

None of the three requires you to assess a line of code. All three are hard to fake, because they require having done the thing rather than having read about it.

What Do You Have to Borrow Judgment For?

Two things, and the borrowing has to be arranged rather than improvised.

The first is whether the technical claims are true. The second is whether the proposed approach is reasonable for your problem, or whether it is the approach the candidate happens to know.

You need a technical person for both. The important part is who that person is, because the default choices are mostly bad ones.

Disqualified: anyone selling you the candidate. A vendor's technical screen is real information but it is not independent, and it should not be the only screen.

Disqualified: anyone selling you a platform. A solutions engineer at an AI tooling company will tell you the truth as they see it, and they see it through their product.

Disqualified: anyone whose reputation depends on this hire working. That includes the person who first suggested you needed AI.

What works is a technical person with no position in the outcome. In practice that is a friendly CTO at a company that does not compete with you, an independent engineer paid for two hours, or a technical advisor already on your board. Two hours is genuinely enough for this. You are not asking them to run a hiring process. You are asking them to sit in one conversation and tell you afterwards whether anything the candidate said was wrong.

Brief them narrowly. "Tell me if any technical claim was inaccurate, and whether the approach they described is a normal way to solve this or an unusual one." That question is answerable in a short call and it is the entire value of the borrowed judgment.

What you are assessing Can you judge it yourself? How
Scoping under ambiguity Yes Describe the problem vaguely and count how many questions come back before a solution does
Honesty about failure Yes Ask what went wrong last time; listen for specifics and ownership, not for the technical content
Volume relevance Yes Ask for items handled, error rate, how quality was measured, and whether the system still runs
Communication with non-specialists Yes You are the test. If you cannot follow the explanation, that is data about them, not about you
Whether technical claims are true No Two hours from an independent engineer with no stake in the hire
Whether the approach fits the problem No Same reviewer, asked whether this is a normal solution or an unusual one
Who should NOT use F5 Hiring Solutions Companies needing a W-2 US employee, on-site presence, an engagement under six months, or a self-serve platform for browsing profiles. F5 Hiring Solutions places full-time professionals from India and the Philippines through a concierge process

What Should You Actually Ask?

Three questions, asked in this order, and then stop.

"Before you started, what would you need from me?" The answer separates people who have shipped from people who have studied. Anyone who has done this work will name data access, a decision about who owns the process being changed, and a definition of what correct output looks like. Anyone who says they would just get started has not been through the part where none of that existed.

"What would you refuse to automate here, and why?" Good practitioners narrow scope, and they can explain the boundary. This question also tells you whether the candidate will tell you something you do not want to hear, which matters more over twelve months than any technical skill.

"What went wrong on your last project?" Covered above, but it belongs in the sequence. Ask it last, once the conversation is comfortable enough to get a real answer.

Notice that none of the three has a correct answer you need to know in advance. You are scoring specificity, not content. That is what makes them usable by a non-technical interviewer.

Resist the urge to ask a question you found online whose answer you cannot evaluate. It produces a confident-sounding response you have no way to check, which is worse than not asking. If you want the full technical version to hand to your borrowed reviewer, what to look for when hiring an AI engineer covers it.

How Should a Trial Work?

Structure it as one process, one person, one quarter, with the baseline recorded before anyone starts. The mechanics, including the four numbers to record in week zero and what each phase should produce, are set out in whether an AI engineer is worth the cost at all and there is no reason to restate them here.

The point specific to your situation is different. A trial is how a non-technical buyer converts an unanswerable question into an answerable one. You cannot assess whether someone is a good AI engineer. You can assess whether a number moved in ninety days. The whole purpose of the baseline is to move the decision onto ground where your judgment is as good as anyone's.

Which is also why the baseline has to be recorded before the person starts. Afterwards, you get an argument, and arguments about technical work are won by the person with the most technical vocabulary.

What Does This Cost?

The US benchmark first, with a caveat that matters here.

No OEWS occupation covers AI engineering. The closest published proxy is Software Developers, SOC 15-1252, at a median annual wage of $135,980 (BLS OEWS, May 2025). Loaded at 1.4265 per BLS Employer Costs for Employee Compensation (ECEC, Dec 2025), that is roughly $193,975 as an estimate of annual employer cost. Where the role is closer to process work than to model work, Management Analysts, SOC 13-1111, at a $101,860 median (BLS OEWS, May 2025) loads to roughly $145,303. Both are proxies and may overstate or understate the actual role, and the loaded figures are estimates built from a national benefits average rather than a measured cost at your company.

F5 Hiring Solutions places AI roles from $600 per week, all-inclusive, inside the $375-$1,200 band that applies to every placement, with AI Solution Architects from $800 per week. All-inclusive covers salary, HR administration, payroll, equipment, compliance, and F5's management, with no setup fee, recruiting fee, or termination cost. Shortlists arrive in 7-14 business days from a network of 85,500+ pre-vetted professionals. That timeline is F5's own published commitment, not an independently measured benchmark, and it should be read alongside any vendor timeline the same way.

Replacement is 7-14 days at zero cost, at any point. F5 Hiring Solutions has served 250+ US companies with a 95% client retention rate, measured as clients continuing beyond the first three months.

For a buyer who cannot assess the work, the replacement mechanism is the part that matters most, because it is the only one that limits the cost of being wrong. F5 Hiring Solutions places full-time professionals across AI engineering, LLM engineering, AI/ML engineering, and AI solution architecture. It does not run a bench outside those roles, so a request outside them is a no rather than a stretch. The wider comparison of routes, including the ones F5 is not, is in the four routes to hire AI experts.

The Bottom Line

Stop trying to evaluate the work and start evaluating the working. Scoping under ambiguity, honesty about failure, and volume relevance are all judged in plain English, and they predict outcomes better than a technical screen you cannot interpret.

Borrow judgment for the two things you genuinely cannot assess, from someone with no stake in the answer, briefed to answer one narrow question in two hours.

Then structure the engagement so that the real decision is made on a number rather than an impression, and record that number before anyone starts.

To do that with a full-time professional rather than a contractor, hire remote AI and ML engineers from India.

Schedule a 15-minute call: https://calendly.com/joel-f5hiringsolutions/f5