AI put in but not used in the field - structural reasons
AI put in but not used in the field —— structural reasons
I have noticed it in the field of AI introduction support.The difference between "until introduction" and "until fixation" is not the difference in workload, but the difference in the structure of the interaction.
Repeatedly listen to the story that the PoC was successful, was applauded in the presentation, and then nothing changed.If we dig into the reason, we won't talk about technology.The only thing that always came out was that they hadn't decided who would use it and how.And it's neither malice nor negligence.Structurally, there is a reason why it is easy to do so.
Stop even if the PoC is successful
McDonald's has partnered with IBM to demonstrate drive-through voice AI in more than 100 stores since 2021.Three years later, in July 2024, all tests were completed.It wasn't a question of accuracy.Environmental noise, dialects, simultaneous speech of multiple customers - the design to the actual site was not up to date.What moved in the lab environment stopped in front of the voices of customers who spoke quickly in the Kansai dialect amid the sound of wind and rain.
The same structure applies to the $62 million investment in MD Anderson Cancer Center, which uses IBM Watson to support cancer treatment.The PoC moved by the repaired data could not correspond to the actual electronic medical record - missing, ambiguous notation, chronological confusion - and was terminated without any application to the actual patient.PoCs operate in a still-image environment.The production moves on the scene where it keeps moving.There was no one to fill in the difference between these assumptions.
Why is this happening?PoC is a place to check whether it moves or not.Move it under the condition of maintenance and confirm that it has "moved".There, there is no question of the design of who the person who runs the operation is, what to check, and what to judge.It will not be designed because it will not be asked.It is only when it is time to bring it to the production stage that the question of "who will use it and how will it be used" arises.However, at that time, the energy of the project has already been exhausted in the place of "moving".
Looking at the numbers, this is no exception
In Cisco's study (2025), only 5% of companies were able to get AI pilots to production.The remaining 95% has stopped somewhere.In Japan, only 21.9% of companies surveyed by Tis are adopting AI on a company-wide basis.According to the PwC Japan survey, the biggest challenge is “not knowing how to use it effectively”, surpassing concerns about security and cost.
! [International comparison of AI adoption retention rates] (../assets/ai-not-working-graph.png)
"I don't know how to use it effectively" does not mean that I don't know how to use the tool.I don't know where to incorporate it into my work.Moving the tool itself is not difficult now.If you paste the text into ChatGPT and type "summarize", it will return something like that.The problem is that it is not decided where to insert the "reasonable thing" into the work, who will use it, and how to judge the result of using it.
"I put it in, but it's not used" is not an exception, it's the majority.And this is not a matter of AI's ability.
Why are you stopping?
From on-site consultation and observation, I think that the reason for stopping will be narrowed down to three reasons.
Added AI to an existing flow
What we often see in the introduction scene is that AI is mounted on top of the existing workflow.The AI will do what the person in charge was doing manually.The person in charge confirms the output, corrects it, and completes it.This will increase the man-hours.This is because there are more work steps before and after using AI.
In a major retailer that introduced generative AI to automate LP production, the total man-hours increased by 1.3 times due to the increase in output confirmation and correction work, and operations were discontinued.The flow of "letting an AI make it and then fixing it" was slower than a skilled person in charge writing from scratch.The sentence generated by the AI is mixed with something that looks like it but is subtly misaligned.If you write from scratch, you can proceed at your own pace, but if you enter the step of "reading and judging what AI wrote", you will read the sentence unreliable from the beginning.
Why is this happening?If you try to introduce AI as a "tool", it is easy to add it to an existing flow.Because it is the idea of "doing what I am doing with AI".However, AI is not a tool, and it must be viewed as an opportunity to redesign the workflow to be effective.Only after redesigning the flow based on the assumption that the AI will move will the man-hours be reduced.It is only by changing the design that the machine detects and provides only the places that humans should judge, rather than confirming the results produced by AI from scratch.
I wasn't sure who would review what
The operation "AI output is reviewed by humans" is correct.The problem is what's inside.When the operation starts under the instruction "Please check all", the person in charge reads the text issued by the AI in full.Check all.This is slower than manual work.Moreover, since it is not clear what to look at, the correction is made with the feeling that "I feel somewhat suspicious", and the judgment axis of quality continues to shake.In one case, the practice of "reading everything just in case because it is written by AI" became established, and the person in charge began to feel that "my work has increased even though I am using AI".
At the site where the review is functioning, the items to be confirmed are specifically determined.Since there are specific checkpoints such as "Is the numerical value correct?" "Is the description of proper nouns aligned?" "Is the tone in line with the guidelines?", confirmation is fast and the judgment axis is not blurred.It has changed to the task of "matching" rather than "reading", so it does not get tired.
Why is this happening?From the fact that there may be errors in the output of AI, it is easy to jump to the conclusion that "we must check all of them".However, areas prone to errors can be predicted in advance.Proprietary nouns, numerical values, and tone shifts - these patterns are known to be disadvantageous to AI.The design of confirming only that ensures quality while reducing the burden.
I didn't know what to do when I failed
What to do when the AI output is strange.When the operation starts without an answer to this question, it stops at the first failure.The person in charge said, "Is this okay?"The moment I think about it, my hand stops.Check with your supervisor.It will take some time to confirm.Eventually, it will be decided whether it is necessary to use AI.
Not in the field where fallbacks are designed.The procedure "If this output is wrong, replace this template and proceed manually" is determined from the beginning.So even if you fail, you won't stop.Does not accumulate mistrust of AI.The person in charge can use it with a sense of security that they can deal with it even if they fail, so continue to use it.
Why didn't you design this from the beginning?I think it is because it is felt that assuming failure will hinder the promotion of the introduction.In the air of "I want to proceed on the premise that it will work", "What if I fail?"The question is not welcome.However, if you enter the production without assuming failure, the entire project will be questioned at the first failure.Deciding what to do in the event of a failure is also about preserving trust in the project.
All you need to “get settled” is a pre-design
Well-functioning sites have something in common.The design precedes the introduction of the a.i.
The decision line is set
There is an agreement in advance to "check only this item in this output".Therefore, the confirmation is not a full re-read, but a point match.The side that receives the results from the AI knows what to look for.
In order to decide this, it is necessary to organize which decisions can only be made by humans in the work.It is a task to ask the question "Can this decision be left to the machine?" for all steps of the work.This is not about AI, but about business design.Many organizations have the opportunity to ask this question for the first time by incorporating AI.
Embedded in workflows
Rather than "use by the person who wants to use it", the execution point of "always pass this tool at this step" is fixed.Habits do not come from freedom, but from coercion.Intentionally creating a situation where you use because you don't have a choice creates a habit of using it for the first time.
The introduction method of "you can use it freely" is difficult to succeed.Use advanced personnel, but do not use non-advanced personnel.There is a difference in proficiency, and it does not have an effect on the whole organization.The tool has become "the user's thing" and is not rooted in the workflow.
The route of failure is clearly indicated
It is a mechanism that does not stop even if it fails once.When the AI output is strange, the route "then manually" is clearly indicated, so the person in charge does not continue to have anxiety.In sites where AI is incorporated into work for the first time, the psychological hurdles of the person in charge are high.The anxiety that "I made a mistake because of AI" makes me hesitate to use it.With fallback, the person in charge can use it with a sense of security that they can return manually at the worst.
These designs cannot be seen from the outside unless they enter the site and organize their work together.It is not something that can be written in the requirements document, but something that you can see while moving together."At what step does the judgment occur?" "What information is needed for that judgment?" "Who is making that judgment?" - these do not just come out by asking the people on the scene.It's the tacit knowledge you see only when you observe while working together.
The invisible problem of 'translation costs'
I think that the concept of "translation cost" is useful when considering the structure in which the introduction of AI fails.
There is always a need for translation between technology and work.Translate work challenges into technical requirements.Translate technical judgments into business language.Each time this round trip occurs, the context is lost.The desire to "solve such tasks with such tools" has gone through many stages of translation before being conveyed to the development side.At the time of communication, the original intention is often diluted.
I think the problem with the outsourced model is that this translation becomes asynchronous.It will take some time for the answer to come back.In the meantime, the scene moves.Requirements become obsolete from the moment they are written.The state of the organization, the tools available, and the work required change each month.It is not possible to respond to changes in the form of suggestions from the outside and leaving.
The way to reduce this translation cost is to get into the field.Judge together every day.Respond on the spot when requirements change.We will also make a design decision to bring PoC to production.That's what it means to accompany rather than outsource.In organizations where AI adoption is well underway, people in this role are either inside or coming in from outside.
Things to do before choosing a tool
Before choosing a tool, it is necessary to linguize the current business.
How long is it going to take?Which decisions are happening every time?What information do I need to make that decision?If you enter the AI without this inventory, you can not determine where to insert it, nor can you determine the indicators to measure the effect.
Specifically, start by writing down the judgments that are occurring repeatedly in the work."Check the reply text of this email", "Check if this number exceeds the threshold", "Choose the one that meets the conditions from this list" --When you write down these judgments, you can see what is left to the AI and what is not.AIs are good at organizing and presenting information.The judgment itself is human.Therefore, "leave it to the AI" means "leave it to the collection and organization of information necessary for judgment".
A common failure is to start introducing it with the policy of "try it first and continue to use it if it seems like it can be used".Since there is no criterion for judging whether it is likely to be usable, time passes without knowing whether it is going well.After half a year, it becomes a state of "some people use it, but the effect is not visible".
In order to measure the effect, it is necessary to understand the state before introduction with numerical values."How many minutes does this task take?" "How many do you process per month?" "How often do you make mistakes?" - this can be compared with after the introduction of AI.What is not measured cannot be improved.
Rather than "doing what AI can do", the question "can this decision be left to AI?" separates the introduction that will take hold and the introduction that will stop.What decisions are being repeated every day at your site?It's up to the AI.The quickest way to start with that question is to look around.
Reference: - [Cisco survey (Uravation 2026)] (https://uravation.com/media/ai-adoption-failure-cases-by-industry-2026/) - [McDonald's AI ends (CNBC 2024)] (https://www.cnbc.com/2024/06/17/mcdonalds-to-end-ibm-ai-drive-thru-test.html) - [S&P Global AI failure rate survey (CIO Dive)] (https://www.ciodive.com/news/AI-project-fail-data-SPGlobal/742590/) - [PwC Japan Generated AI Survey Spring 2025] (https://www.pwc.com/jp/ja/knowledge/thoughtleadership/generative-ai-survey2025.html)
---
*Machine-translated (MyMemory API) from a Japanese original at [nomuraya-hub.pages.dev](https://nomuraya-hub.pages.dev/). Pre-review draft. I am the same author writing under different pen names — "nomuraya / shimajima / 中翔" — depending on the medium.*