Skip to main content

Our AI pilot didn't produce anything useful. How do we tell if the model was the problem or the data was?

AI Integration Published August 28, 2026
Short Answer

Run the same task by hand on the same records, using only what the system was given. If a smart person with that input also gets it wrong, the input is the problem. If they get it right without effort, the model, the prompt or the plumbing is the problem. Most failed pilots fail on inputs and unclear task definition, not on model capability.

The hand test separates the two in an afternoon

Pull thirty records the system handled badly. Give a capable person exactly what the system received - no more context, no access to the CRM, no institutional memory - and ask them to do the task. This takes a few hours and settles an argument that otherwise runs for months.

If the person struggles, the model was never going to succeed either, because the necessary information was not in the input. That is an integration problem, and it is solved by connecting more of the business rather than by swapping models. This is the whole premise behind operational AI: capability is rarely the constraint, context is.

What a data failure looks like

Data failures have a signature. The output is confidently wrong in a consistent direction, and it is wrong on exactly the records where a field is empty or a code is ambiguous.

  • Missing context. The model saw a job record but not the call that produced it, or the call but not the invoice.
  • Inconsistent taxonomy. Two dispatchers use the same job type for different work, so any category-level conclusion is noise.
  • No ground truth. Nobody can say what the right answer was, so nobody can measure improvement.
  • Unrepresentative sample. The pilot ran on your cleanest location and production runs everywhere.

What a model or prompt failure looks like

Model failures look unstable rather than consistently wrong. The same input produces different answers on different runs. The output is right but in the wrong format. Instructions in the middle of a long prompt get ignored while the ones at the end are followed. The correct answer is present but buried in three paragraphs of hedging.

These are fixable, and they are the cheap kind of problem. Output structure, instruction placement, retrieval of the right reference material and choice of model are all adjustable in days. Rebuilding a category scheme across two years of history is not.

The third failure that nobody names

A large share of pilots fail because the task was never defined well enough for anyone to succeed at it. "Analyze our calls" is not a task. "Flag calls where a customer asked for a service we offer and no appointment was booked" is a task, because you can look at a call and say whether the flag was right.

Before deciding what went wrong technically, ask what decision the output was supposed to change and who was supposed to make it. If there is no answer, the pilot did not fail. It was never a test. Scoping that decision first is the first step in any real AI systems integration engagement.

Topics: pilots · diagnostics · data quality · failure modes

Have a version of this question about your own business?

The useful answer usually depends on which systems you run and how they're connected. That's a conversation, not a blog post.

Related Answers

People who read this also asked

Browse the Answer Hub →

AI is easy to access. Making it useful is hard.

Bluefrog makes AI useful by integrating it with the way your business actually works — your software, your calls, your customers, your marketing and your revenue.

Technology development since 1997 · AI integration platforms since 2001