Skip to content
NewEraAI

Adoption

Why AI pilots die in month three

Not because the technology failed. Because nobody owned it, the success criteria were never written, and the pilot was built on a process that was never defined in the first place.

15 July 2026 · 5 min read · New Era AI

The pattern is remarkably consistent. Month one: enthusiasm, a demo that works, a slot in the leadership meeting. Month two: it is still running, usage has dropped, someone raises a concern about accuracy. Month three: nobody mentions it, the subscription renews silently, and six months later somebody cancels it during a cost review.

The post-mortems almost never blame the technology, and they are right not to. The model did roughly what models do. What failed was everything around it.

Here are the four causes, in the order they do damage.

1. Nobody owned it

Pilots get sponsors. Sponsors approve budget and attend the kick-off. What pilots need is an owner: one named person whose job description now includes this thing working.

The distinction sounds like semantics until month two, when the tool starts producing an output that is 85% right. At that point somebody has to decide what happens to the other 15%. A sponsor cannot make that call from a monthly steering meeting. An owner makes it on a Tuesday, changes the rule, and moves on.

Without an owner, the 15% becomes everyone's grievance and nobody's task. Usage quietly stops, because using something that is sometimes wrong is more work than not using it, unless someone is actively closing the gap.

The fix is unglamorous: name the owner before the build starts, give them the authority to change the rules, and put their name on the roadmap next to the item. If nobody will accept the role, that is important information (it usually means the problem is not painful enough to justify the project.

2. Success was never defined in a way that could fail

Ask what success looks like at kick-off and you will usually get "save time" or "improve the customer experience". Neither of those can be false. That is the problem.

A criterion that cannot fail cannot be met either. Three months in, the honest question) is this working? (has no answer, so the conversation defaults to vibes, and vibes about new technology decay.

A usable criterion has a number, a baseline and a date. "Time from enquiry to first outbound reply, currently a median of 51 minutes with a tail past 24 hours, to a median under 10 minutes with no tail beyond 4 hours, measured at the end of quarter two." That can fail. Because it can fail, it can also be defended in a budget meeting, which is the actual reason projects survive.

Note the baseline. Most pilots skip it, and then cannot prove improvement even when it happened.

3. It was built on a process nobody had written down

This is the deepest cause and the one that produces the most wasted money.

Automating a decision requires the decision to exist. In most small and mid-sized businesses, the important decisions) which enquiries to chase, how to price an unusual job, when to escalate (live in one experienced person's head as a set of instincts with a long tail of exceptions.

A pilot built on top of that produces confident, plausible outputs that senior people quietly disagree with. Nobody can articulate exactly why it is wrong, because the correct rule was never stated. Confidence erodes. The tool gets used for the easy cases only, which were never the expensive ones.

This is why an honest audit spends more time writing down how work currently happens than looking at software. The written rule is the deliverable. Occasionally the business reads it, fixes the process with a checklist, and no longer needs the pilot) which is a success, however unsatisfying it feels to a technology supplier.

4. The output had nowhere to go

The mechanical version of the same problem. An agent answers calls beautifully and writes transcripts into an inbox. A document tool extracts data into a file nobody imports. A scoring model ranks leads in a dashboard the sales team does not open.

Each of these is a pilot that technically works and operationally does nothing. The work product exists; it just never reaches the system where work actually happens.

The rule we use: nothing gets built until its output has a named destination and a named consumer. Not "it will feed the CRM" (which field, on which record, triggering which sequence, reviewed by whom. If those questions do not have answers, the integration is the project, and the clever part can wait.

The month-three test

There is a simple diagnostic you can run on any pilot at week two, long before the money is spent.

Ask four questions and require four specific answers:

  1. Who owns this? A name, not a department.
  2. What number moves, from what, to what, by when? With a baseline you have actually measured.
  3. What is the rule this thing is applying, in writing? If it takes forty minutes to explain, it is not written down yet.
  4. Where does the output land, and who acts on it? A field, a record, a person.

A pilot that answers all four survives contact with month three more often than not. A pilot that answers none of them will fail regardless of which vendor you pick, and it is much cheaper to discover that now.

Why the retainer is structured the way it is

This is also the reason our consulting retainers are sold by days per quarter rather than as a single project fee, and why build work is quoted separately.

A one-off workshop produces a plan. Nobody comes back to it, the business changes, and the plan is a PDF. Quarterly days exist specifically to catch month three) to sit down when the enthusiasm has worn off, look at the number that was supposed to move, and either fix the thing or kill it honestly.

Killing things honestly is underrated. A pilot stopped in month three with a written reason is a cheap lesson. The same pilot left running quietly for eighteen months is a line item that makes the next proposal harder to get approved.

What to do differently on the next one

If you have a pilot that faded, the useful move is not to try a different vendor. It is to run the four questions against the previous attempt and see which one has no answer. In our experience it is question three about two-thirds of the time.

Then take the smallest job that has all four answers, build only that, and let it run for a quarter before adding anything. Small and finished beats broad and abandoned, and it is much easier to defend when someone asks what the last twelve months of AI spend produced.

If you want a second pair of eyes on a pilot that is drifting, the free AI Opportunity Assessment is a reasonable place to start. It is a conversation about where the work leaks, not a pitch for a platform.

Next step

Start with the free AI Opportunity Assessment.

A short, no-obligation conversation about where enquiries, hours and revenue leak today. You do not have to pick a tier to have it, and what comes out of it feeds Discover, so the first paid day starts from evidence rather than a blank sheet.