Week 16: ChatGPT Wrote My A3. Here's Why It Nearly Killed the Project.

ChatGPT Wrote My A3. Here's Why It Nearly Killed the Project.

I ran an experiment. A real one. Not a hypothetical thought exercise for a LinkedIn post -- an actual test with a genuine operational problem, a genuine A3 template, and a genuine AI tool.

I wanted to answer a question that I think every CI practitioner needs to confront in 2026: Can AI do structured problem solving?

Not the kind of problem solving where you Google an answer. The kind where you investigate a real process, identify a real root cause, and develop a real countermeasure. The A3 kind. The kind that requires going to the Gemba, talking to people, and understanding the physics of what is actually happening.

Here is what happened.

The Setup

The problem was real. A recurring quality defect on a production process. Specific product, specific defect mode, intermittent occurrence. The kind of problem that a Black Belt project would normally take three to four weeks to investigate.

I fed the AI the problem statement. I gave it the defect data -- six months of quality records, inspection results, process parameters, and production logs. I asked it to complete an A3: background, current condition, target condition, root cause analysis, countermeasures, and implementation plan.

It produced the A3 in eleven minutes.

The output was impressive. Structurally correct. Well-organised. Clear language. Professional formatting. The background section accurately summarised the problem. The current condition section presented the data clearly. The analysis section identified correlations in the data that were statistically valid.

If you were reviewing this A3 in a tollgate meeting and you did not know it was AI-generated, you would probably approve it. It looked right. It sounded right. It hit all the structural markers of a competent analysis.

There was just one problem.

The root cause was wrong.

Where the AI Went Wrong

The AI identified a correlation between the defect occurrence and a specific process parameter -- temperature variation during a particular phase of production. The correlation was real. The statistical analysis was correct. When temperature varied beyond a certain range during that phase, defect probability increased significantly.

Based on this correlation, the AI proposed the root cause: inadequate temperature control during that production phase. The countermeasure: tighten temperature specifications and install additional monitoring.

Logical. Data-driven. Statistically sound. And completely wrong.

I know it was wrong because I went to the Gemba.

I stood at the production line. I watched the process. I talked to the operators. And within two hours, I understood what the data could not tell me.

The temperature variation was not the cause. It was a symptom. The actual root cause was a maintenance issue with the heating element in the upstream equipment. When the element degraded -- which happened intermittently based on a wear pattern that was not captured in any dataset -- it caused inconsistent heat transfer. This inconsistency propagated downstream and manifested as temperature variation in the phase the AI had flagged.

The AI saw the temperature variation and called it the root cause because it was the strongest statistical signal in the data. It could not see the heating element. It could not feel the vibration that the operators had noticed but never formally reported. It could not know that three months ago, a maintenance technician had flagged the element as "approaching end of life" in a handwritten note on a maintenance board that was not connected to any digital system.

If we had implemented the AI's countermeasure -- tighter temperature specifications and additional monitoring -- we would have created a very expensive alarm system for a symptom while the actual cause continued to degrade. The defect would have continued. The monitoring would have flagged every occurrence. And we would have been managing a problem instead of solving it.

The Gemba Cannot Be Googled

This is the fundamental limitation of AI in structured problem solving. And it is not a limitation that will be solved by better models, more data, or more sophisticated algorithms.

AI can only analyse what is in the data. The root cause of many operational problems is not in the data. It is on the shopfloor. In the conversations with operators. In the physical observation of the process. In the tribal knowledge that has never been digitised and probably never will be.

The heating element degradation was not in any dataset the AI could access. The operator's observation about vibration was in his head, not in a database. The maintenance technician's handwritten note was on a physical board, not in a CMMS system.

The Gemba -- the real place where the real work happens -- contains information that no data model can capture. Not because the data is too complex. Because the data does not exist in digital form. It exists in human experience, physical observation, and contextual understanding.

AI guessed the root cause. It made an educated guess based on the strongest statistical signal available. And educated guesses, no matter how statistically sophisticated, are not root cause analysis.

Root cause analysis requires investigation. Physical, human, contextual investigation. Going to where the problem happens. Watching it happen. Talking to the people who deal with it every day. Understanding the system, not just the data.

What AI Can and Cannot Do

Let me be fair to the technology. Because this is not an anti-AI argument. This is a precision argument.

What AI did well:

Data organisation. The AI structured six months of data clearly and comprehensively. It identified patterns that would have taken a human analyst days to find. That is genuinely valuable.

Statistical analysis. The correlations it identified were real and correctly calculated. The statistical rigour was sound. A practitioner doing the same analysis manually would have reached the same statistical conclusion -- and might have made the same mistake.

Structure and formatting. The A3 was well-organised, clearly written, and professionally presented. The framework was followed correctly. The logic was coherent within the constraints of the data available.

Draft generation. As a starting point for investigation, the AI's output was excellent. It identified where to look. It highlighted the variables worth investigating. It gave the practitioner a head start.

What AI cannot do:

Go to the Gemba. AI cannot stand at a production line, feel a vibration, smell an abnormality, or observe the subtle physical cues that experienced operators detect intuitively.

Access undocumented knowledge. The most important information in most organisations has never been digitised. It lives in the heads of the people who do the work. AI cannot access it because it does not exist in any system the AI can reach.

Understand process context. AI can identify that temperature varied. It cannot understand why temperature varied without access to the physical cause-and-effect chain that created the variation. Correlation is not causation, and the gap between them is filled by human investigation, not algorithmic analysis.

Challenge its own assumptions. The AI presented its root cause with confidence. It did not flag its own uncertainty. It did not say "this is a statistical correlation that may or may not be causal -- someone should go verify." It stated its conclusion as if it were established fact. And someone without process experience might have accepted it.

The Real Danger

Here is what concerns me about the current enthusiasm for AI in CI.

Someone less experienced than me might have taken the AI's A3 at face value. They might have implemented the countermeasure. They might have reported the root cause to management. They might have closed the investigation, signed off the A3, and moved on.

The defect would have continued. The actual root cause would have gone unaddressed. And the organisation would have learned a dangerous lesson: AI-generated analysis is sufficient. No need to go to the Gemba. No need to talk to operators. No need to investigate physically.

That lesson, once learned, is extremely difficult to unlearn. Because it is efficient. It is fast. It produces professional-looking output. And it feels like progress.

I have spent 20+ years building a career on the principle that you cannot understand a process without going to where the process happens. That principle is under threat. Not from AI directly, but from the temptation to use AI as a shortcut that bypasses the human investigation that actually solves problems.

The Right Model

AI should draft. Humans must think.

Use AI to analyse the data. Use AI to identify patterns. Use AI to generate the first draft of the A3 -- the background, the data summary, the statistical analysis. That saves time and adds rigour.

Then go to the Gemba. Verify the statistical findings against physical reality. Talk to the operators. Watch the process. Challenge the AI's conclusions with contextual knowledge that no algorithm can access.

The combination is powerful. AI-speed data analysis with human-depth investigation. But the human investigation is not optional. It is not a nice-to-have. It is the part that separates a statistically sophisticated guess from an actual root cause.

The A3 that I completed after going to the Gemba -- the one with the correct root cause -- took three weeks. The AI draft that nearly killed the project took eleven minutes.

Three weeks beats eleven minutes when the alternative is solving the wrong problem.

Here is my question for every CI practitioner using AI tools: When was the last time AI gave you an answer that you verified at the Gemba before acting on it? And when was the last time that verification changed the answer? If you have never verified, you have never actually solved the problem. You have accepted a guess.


Build the problem-solving capability that AI cannot replace. The Stormholt CI Practitioner Mastery Programme teaches the thinking behind the tools -- 16 modules from real deployments.

View the programme


New to Stormholt? Start here — it locates your stage in minutes.

Building capability? Stage III — Skill — the tools, the courses, the simulations.

Back to blog

Leave a comment

Please note, comments need to be approved before they are published.