Aug 10, 2026

What not to automate

AI projects that failed

Why AI Projects Fail Even When the Technology Works

One of the easiest questions to ask before investing in AI is:

Can the technology do the job?

It is also one of the least useful questions on its own.

An AI system can perform well in testing, hit its technical targets and still create very little value once it reaches the people who are supposed to use it.

The reason is usually everything around the model.

Does it fit the existing workflow?

Do people trust its output?

Can it reach the information it needs?

What happens when it gets something wrong?

And perhaps most importantly, does removing the easy work actually remove the workload?

Three well-known AI projects show why those questions matter.

The easy work disappears first

AI is generally strongest at repeatable, predictable work.

That means when a process is automated, the routine cases tend to disappear first.

What remains with the team are the exceptions.

The strange customer request.

The account with incomplete information.

The issue that does not follow the normal process.

The situation where somebody needs judgement rather than a standard response.

This creates an important distinction:

volume is not the same thing as workload.

A team could lose 60% of its incoming tickets while retaining a much larger percentage of the actual difficulty.

If headcount is then reduced purely because ticket volume has fallen, the remaining employees can end up handling a smaller number of much harder cases with less spare capacity than they had before.

The AI can be working exactly as intended while the overall operation becomes worse.

Salesforce: proving the edge cases matters

Salesforce provides a useful example of the sequencing problem.

The company has made an enormous investment in autonomous AI agents, with the ambition of handling increasingly large parts of customer support through AI.

During 2025, its support organisation reduced from around 9,000 people to 5,000, according to public comments from its chief executive. Salesforce has also said that hundreds of employees moved into other roles rather than simply leaving the company.

The broader lesson is less about Salesforce itself and more about what happens when organisations plan capacity around the work AI is expected to remove.

Support teams carry knowledge that is difficult to capture completely in documentation.

They know recurring quirks.

They understand which internal team to ask when something unusual happens.

They know the exceptions that technically should not happen, but somehow appear every week.

AI may absorb the standard interactions very effectively.

If the organisation reduces human capacity before understanding how reliably the system handles everything around those standard interactions, the hardest work becomes concentrated onto fewer people.

The relevant measure is therefore not simply:

How many tickets did the AI handle?

It is:

What work remained, and how difficult was it?

McDonald's: context changes what "accurate enough" means

McDonald's tested automated voice ordering with IBM across more than 100 US restaurants from 2021.

The idea made sense.

Use voice AI to take drive-thru orders consistently and efficiently, particularly during busy periods.

The system reportedly achieved accuracy in the low to mid 80% range.

In another environment, that might sound respectable.

A drive-thru is not another environment.

A voice assistant being used at home can ask somebody to repeat themselves.

At a busy restaurant, every clarification adds time for the customer, the cars waiting behind them and the kitchen preparing the order.

An incorrect response also does not disappear.

A member of staff has to notice it, correct it and deal with the customer.

The work has moved.

It has not necessarily gone away.

McDonald's ended the trial across participating restaurants in July 2024, while making clear that it still expected voice ordering to play a role in the future.

The lesson is not that voice AI cannot work.

It is that the environment determines how expensive an error is.

A 90% success rate can be excellent in one workflow and unusable in another.

MD Anderson: an accurate system nobody used

A very different example came from the MD Anderson Cancer Center and IBM Watson.

The project aimed to help oncologists process huge amounts of medical research and identify clinical trials that might be relevant to individual patients.

It ran for roughly five years.

Around $62 million was spent, according to a University of Texas System audit.

The system never reached routine use with real patients.

The problem was not simply whether Watson could produce an answer.

It was whether that answer fitted into how clinicians actually worked.

The system could not integrate properly with the hospital's electronic health record.

Clinicians therefore had to work across information held in different places.

There was also the question of trust.

If somebody receives a recommendation from AI but cannot quickly see the evidence behind it, they have a choice:

trust it blindly, or check the work themselves.

In healthcare, checking is unsurprisingly the safer option.

Once somebody needs to manually verify the recommendation each time, the AI can become another opinion to review rather than a piece of work removed.

An accurate answer only creates value when somebody can actually use it.

Three questions matter before the build

These examples come from very different industries, but the same questions appear repeatedly.

1. Will people trust the output enough to use it?

An AI system does not save time if employees check behind it every time.

Trust should not mean blind acceptance.

It means giving people enough evidence, transparency and confidence to know when they can rely on the output and when they should intervene.

2. Does it fit the way the work actually happens?

A demo can work beautifully in isolation and fail the moment it needs to interact with the CRM, inbox, customer database or internal approval process.

Integration should not be treated as the final technical step.

It is part of whether the idea works at all.

3. What happens when it gets something wrong?

Every AI system will eventually encounter something unusual.

The important question is where that exception goes.

Who receives it?

How quickly can they intervene?

What information do they receive?

How much work does correcting it create?

An automation that saves five minutes on every standard case but creates an hour of work whenever it fails may look better on a dashboard than it feels inside the business.

Plan the transition, not just the technology

There are a few practical ways to reduce these problems.

Run the AI and the existing process alongside each other before removing capacity.

Give users evidence behind important answers wherever possible.

Measure the complexity of the work that remains, not simply the volume that disappeared.

Test integrations before assuming they will work.

Design a clear route for exceptions and human intervention.

And watch what employees actually do once the system goes live.

If they keep bypassing it, checking every answer or maintaining their own manual process in parallel, that is useful information.

The implementation is telling you something.

A successful AI project is an operational change

Technical capability matters.

It simply is not enough.

The model can be accurate, fast and impressive in a demonstration while still being the wrong answer for the business around it.

Successful AI adoption means planning for the people, the workflow, the exceptions and the systems it needs to reach just as carefully as the model itself.

That is also why TUSTRA starts with understanding how a business operates before deciding what should be automated.

The aim is not to find something AI can do.

There are plenty of those.

The useful work is finding where it can be introduced in a way that people will actually use and that genuinely improves the business.

FREE EMAIL BRIEFING

The Tustra Briefing

AI in plain English for UK business.

One considered briefing covering what changed, why it matters, how businesses are using it, what to be cautious about and one practical action worth considering.

Important developments without daily noise

Practical UK business context

Honest case studies, security and adoption guidance

Free to join. Unsubscribe at any time. We will not sell your information.