Is Your Data Ready for an AI Engineer?

Run five checks on yourself before hiring: can the data be exported, is it labelled, is there enough of it, do you have permission to use it, and does someone still understand what the fields mean. Fail two or more and the hire will spend its first quarter on procurement instead of building.

The expensive version of this article is hiring someone and finding out.

Every check below can be run in an afternoon by someone who already works at your company. None of them requires technical skill, a consultant, or a tool. And the failure they prevent is the most common one in applied AI: an engineer joins, discovers the data is not reachable or not usable, and spends the quarter in meetings while the project's reputation quietly degrades. The delay then gets attributed to the technology, or to the person, and almost never to the thing that actually caused it.

The broader pattern of AI initiatives failing on organisational rather than technical grounds is covered in why AI projects fail on the talent gap. This article is the narrower, more actionable half: the audit you run on yourself first.

Check One: Can You Actually Get the Data Out?

Not "do we have the data". Can you export it, in bulk, without asking anyone's permission each time.

The distinction matters because most companies genuinely do have the data and cannot reach it. It sits inside a vendor platform whose export is a CSV of the last 90 days, or an API with rate limits that make a full extract take weeks, or a reporting layer that shows aggregates and not rows. In each case the information exists and is unavailable for the purpose.

The test is concrete. Ask whoever administers the system to produce a full export of one year of the relevant records, to a file, this week. Not a report. A file.

Three outcomes. It arrives, and you pass. It arrives partially, in which case you have learned exactly what the limit is and can plan around it. Or you discover it requires a contract change, a paid API tier, or a vendor conversation, which is a procurement task with its own timeline and should start now rather than in month two of an employment contract.

One thing worth naming: an export you can only get once is not the same as an export you can get repeatedly. Anything that runs in production needs the second kind. If the vendor will hand over a one-time dump but has no ongoing feed, the system you build will be a snapshot that ages.

Check Two: Do Examples of Correct Answers Exist?

This is usually called labelled data, which makes it sound like a technical asset you either bought or did not. In practice the question is softer and more answerable: somewhere in your history, is there a record of what the right answer was?

Resolved support tickets carry the category someone eventually assigned. Approved invoices carry the coding a person chose. Closed claims carry the decision. Past hiring decisions, completed inspections, reconciled accounts, all the same shape. These are examples of correct answers, generated as a by-product of doing the work, and they are usually enough.

What actually blocks a project is not the absence of a labelling tool. It is the absence of agreement about what correct means. If two experienced people in your company would categorise the same item differently, and neither is wrong, then the definition is not settled and no amount of engineering settles it for you.

Test that directly. Take twenty real items. Have two people who do the work classify them independently. Compare. If they disagree on more than a handful, you have found a scoping problem that will otherwise surface in month three as "the model is not accurate enough". The model would be reproducing a disagreement that was already there.

Check Three: Is There Enough of It?

There is no universal number, and anyone who gives you one is guessing. Volume requirements depend on how varied your inputs are, which is specific to your business.

The usable version of the question: can you assemble a few hundred real examples that cover both the normal cases and the awkward ones?

The awkward ones are the point. Ordinary examples are easy to find and teach a system very little. The edge cases, the ones your experienced staff handle by instinct, are what determine whether the system works in production. If you cannot locate them, that is not a volume problem. It usually means those cases were handled by a person who never wrote down what they did, which is check five arriving early.

A related trap is data that is plentiful but unrepresentative. Two years of tickets from before a product change describe a world that no longer exists. Volume from the wrong period is worse than less volume from the right one, because it looks sufficient.

Check Four: Are You Allowed to Use It?

The one most often skipped, and the only one that can remove the project entirely rather than delay it.

Look in three places, and look at the actual documents rather than asking someone's recollection.

Your customer contracts. Many contain limits on secondary use of data your customers put into your product. Training or processing it for a new purpose may sit outside what they agreed to.

Your vendor agreements. Software contracts frequently restrict processing data outside the platform, exporting in bulk, or sending it to a third-party service. That last one matters if the plan involves a model API.

Your regulatory position. Health, financial, education, and children's data carry constraints that apply regardless of what your contracts say, and they differ by jurisdiction.

This check has a specific characteristic worth planning around: it is the only one where the answer can be a permanent no. The others produce delay or extra work. This one can end the project, which is exactly why it should happen before an employment contract rather than after.

Check Five: Does Anyone Still Know What the Fields Mean?

Data outlives the people who designed it. Five years on, a table has a column called status_2 holding values nobody can explain, a flag that meant one thing before a migration and another after, and a date field that is sometimes the event and sometimes the entry.

Someone has to be able to answer questions about this, and that someone has to still work at your company and have time.

The test: pick ten fields the project would rely on and ask a colleague to explain each. Not what it is called. What it means, when it is null, and whether the meaning ever changed.

If the honest answer is that the person who knew has left, you are not blocked, but the first two months of any engagement will be archaeology. Budget for it, or reconstruct the knowledge before hiring, which is cheaper.

Check How to run it this week What a fail costs you
Exportable Ask for a full one-year export to a file, not a report A procurement timeline, discovered in month two instead of now
Repeatable export Confirm the same export can run on a schedule A production system built on a snapshot that ages
Examples of correct answers Two staff classify the same twenty items independently, then compare A disagreement inside your business, surfacing later as model inaccuracy
Sufficient and representative volume Assemble a few hundred examples including the awkward cases A system that works in testing and fails on real inputs
Legal permission Read customer contracts, vendor agreements, and your regulatory position The only check that can end the project rather than delay it
Institutional memory Ask a colleague to explain ten fields, including when each is null Two months of archaeology inside a paid engagement
Who should NOT use F5 Hiring Solutions Companies needing a W-2 US employee, on-site presence, fractional or part-time work, an engagement under six months, or a self-serve platform for browsing profiles. F5 Hiring Solutions places full-time professionals from India and the Philippines through a concierge process

What the Score Means

Five passes. Hire. Your first engagement will be spent on the actual problem, which is the only version of this that pays for itself.

One fail. Proceed, with the fix scheduled before the start date and someone named against it. A single known blocker with an owner is a plan.

Two or more fails. Do not hire yet. Not because the project is wrong, but because you would be paying engineering rates for work that is procurement, legal review, and internal archaeology. Those are real tasks and they are cheaper done by the people who already have the relationships and the authority.

That last case is the honest conclusion of this article and it applies to more readers than it sounds like. It is also reversible: most companies that fail two checks can pass them within a quarter, using people already on payroll, and then hire into a situation where the engineer builds from week one.

If the exercise suggests the whole hire is premature rather than just early, the prior question is argued in whether an AI engineer is worth the cost at all.

Who Should Run This

Someone on the business side with authority over access decisions. Not the future engineer.

Four of the five checks are about permissions, contracts, vendors, and institutional memory. A new hire has no authority over any of those, no history with the systems, and no standing to escalate. Handing them the audit as a first task converts a two-week internal exercise into a two-month one, performed by the most expensive person available.

If nobody internally can run it, that is itself a finding, and it usually means check four has never been examined. Where the buyer has no technical staff at all, hiring AI talent when you have no technical team covers how to borrow the judgment you need without borrowing the decision.

What This Has to Do With Cost

Because the cost of skipping it is a salary, not a delay.

The closest published US benchmark for this kind of role is Software Developers, SOC 15-1252, at a median annual wage of $135,980 (BLS OEWS, May 2025), which loads to roughly $193,975 at the 1.4265 multiplier from BLS Employer Costs for Employee Compensation (ECEC, Dec 2025). Where the work leans analytical, Data Scientists, SOC 15-2051, at a $120,230 median (BLS OEWS, May 2025) loads to roughly $171,508. Both are proxies, since no OEWS occupation covers AI engineering directly, and either may overstate or understate the real role. The loaded figures are estimates built from a national benefits average, not a measured cost at your company.

One blocked quarter is a quarter of that, spent on procurement. The audit above costs an afternoon.

F5 Hiring Solutions is a managed remote workforce company placing full-time professionals at $375-$1,200 per week, all-inclusive, with AI roles from $600 per week. All-inclusive covers salary, HR, payroll, equipment, compliance, and management, with no setup or recruiting fee. Shortlists arrive in 7-14 business days, which is F5's own published commitment rather than an independently measured benchmark.

None of that helps if the data is not reachable. Run the audit first.

The Bottom Line

Five checks, an afternoon, run by someone who already works there: exportable and repeatably so, examples of correct answers, enough of the awkward cases, legal permission, and someone who can still explain the fields.

Two or more fails means wait. That is a real recommendation, not a hedge, and acting on it is cheaper than any other way of learning the same thing.

When the checks pass, hire remote AI and ML engineers from India into a problem that is ready for them.

Schedule a 15-minute call: https://calendly.com/joel-f5hiringsolutions/f5