Manufacturers spent the first half of 2026 discovering that their AI problem was probably never a math problem. Grant Thornton polled 950 business leaders between late February and mid March, and among the hundred manufacturing respondents, not one reported a major revenue uplift from AI. Not one reported major cost savings either, against 12% of executives across every other industry.
The models running on those factory floors work fine. The shifts around them appear not to.
That reading deserves one caveat up front. The manufacturing subgroup was a hundred respondents, and Grant Thornton flags role-specific findings within it as directional. The direction, however, matches almost everything else published this year.
The Distance Between Piloting and Finishing
The returns gap sits in a specific place, and the same survey names it. Forty-eight percent of manufacturers are piloting AI, comfortably above the 34% full-sample rate. Only 10% have folded it into operations, four points below everyone else. On scaling across multiple functions, manufacturing lands at 39% against a 49% average.
So manufacturing experiments more than any other sector and finishes less. That is a peculiar way to lose a race.
Sixty-four percent of those manufacturers do report better efficiency, which sounds like progress until you check what arrived behind it. Fourteen percent say AI has sped up how quickly they develop new products and processes, against 31% across the whole survey. Efficiency with nothing downstream of it tends to be a plateau with good lighting.
What Buyers Are Asking Vendors Now
The purchasing criteria have moved to match. Evaluations of AI manufacturing software increasingly hinge on whether a line lead will open the thing at two in the morning when a filler jams, rather than on any benchmark score.
Vendors used to demo detection accuracy. The demos that close now show a supervisor resolving a downtime event before the shift report gets written. Anyone sitting through a procurement cycle this year has probably watched that reframing happen in real time.
The Plants That Already Solved It
None of this is a mystery to the companies that have got past it. They appear to have solved a different problem than the one most manufacturers are funding.
The World Economic Forum added sixteen sites to its Global Lighthouse Network in June, bringing the community to 238 plants, and the pattern across the cohort is dull in the best possible way. Nobody won on model selection. They won on how many use cases they wired into daily operating routines.
Rockwell Automation’s Singapore site makes the case most cleanly. It builds more than a thousand different products and changes over its lines upward of 20,000 times a year, which is the kind of complexity that eats software alive.
That site deployed more than fifty digital and AI solutions. Units per person-hour rose 43%. Defects fell 35%. Time to competency for new operators dropped 67%, which is the figure that should interest anyone who has tried to staff a second shift lately.
Schneider Electric’s El Paso plant tells a similar story with different numbers. On-time delivery climbed from 61% to 97%, and $43 million in backorders disappeared. Those gains came out of data engineering, integrated logistics and industrial IoT, with models doing part of the job inside them.
Hitachi Vantara’s Norman, Oklahoma site is the one worth studying closest, because its stated obstacle was fragmented data rather than any shortage of computing power. Product complexity had climbed, demand bunched at quarter end, and nobody could see inventory across the network. Pulling that visibility into one platform cut inventory in half and shortened order-to-ship lead times by 77%.
Where Programs Actually Die
The unglamorous middle of that job is where most programs seem to end. Anyone who has sat through an AI quality control rollout recognises the failure modes, and few of them involve the algorithm. Legacy systems refuse to talk to each other. Training data reflects one product line and no other. Operators who were never consulted decide the system exists to watch them.
That last one has research behind it. RAND interviewed 65 experienced data scientists and engineers for its study of why AI projects fail, and 84% of the industry group named leadership decisions as the primary cause. Models get deployed having been optimised for the wrong metric, or having no place in the surrounding workflow at all.
RAND also found something that should end the model-quality argument outright. Nearly every practitioner interviewed said computing power was not a limiting factor in their work. Data quality was, for 30 of the 50 industry respondents. Infrastructure was. Compute almost never was.
The same study noted that subject-matter experts sometimes offer passive resistance to AI projects because they suspect the projects exist to replace them. Resistance from the floor gets filed as a change management footnote. On this evidence it is closer to the entire game.
What Operations Leaders Can Do About It
The fix is unfashionable and mostly procedural, and it starts before anyone writes code.
Name an owner for the model after the pilot team disbands. RAND’s recommendation is blunter than most consulting advice on this point, suggesting leaders commit a product team to a specific problem for at least a year before starting, on the grounds that anything not worth that commitment is probably not worth starting.
Write down what an exception looks like and who reviews it. Production conditions drift, and a detection threshold that was right in March stops meaning much by September if nobody owns the recalibration.
Instrument the handover, not just the model. Deloitte’s 2026 manufacturing outlook points to autonomously generated shift handover reports and work instructions as one of the clearer near-term applications of agentic AI on the floor, which is a useful signal about where the constraint sits. The same outlook found that more than a third of 600 surveyed manufacturing executives named equipping workers with the right skills as their top concern.
Measure time-to-action rather than model accuracy. A system that flags a defect trend in four seconds and gets acted on in nine hours is a nine-hour system. That number rarely appears on a vendor scorecard, and it is the one that shows up in scrap rates.
Budget for the boring roles. RAND’s interviewees described data engineers as the plumbers of data science, unglamorous enough that turnover is high and institutional knowledge walks out with them. Losing the person who knew which datasets were trustworthy can set a program back further than a model swap ever would.
Adoption Data Says Operations Is Getting Skipped
Federal data points the same direction from a different angle. A Census Bureau working paper published in April, drawn from the 2026 AI supplement to the Business Trends and Outlook Survey, found that 57% of AI-adopting firms use it in three or fewer business functions. Sales and marketing leads that list at 52%. Strategy and business development follows at 45%.
Operations does not appear near the top, which is worth sitting with, because operations is precisely where manufacturers say they need AI most. Sixty-two percent of them named it as the function most in need of additional focus.
Worker-level use looks thinner still. Sixty-five percent of firms confine AI to three or fewer tasks, and two thirds of users apply it purely to augment work already being done. Headcount reductions tied to AI showed up at just 2% of firms, which deflates a fair amount of boardroom rhetoric about replacing people. Deloitte’s own projection runs in the same direction, expecting more than 81% of manufacturing task hours to stay human-driven.
The same Census paper found that breadth of integration correlated with commercial performance, measured across functional deployment, worker-task use and operational investment. What that research tracked was reach. How much of the business the technology genuinely touched.
The Same Argument, Running Through Software
A parallel debate has been going on in software engineering for a year, and it landed in the same place. The people building agents worked out that the bottleneck moved away from the prompt and into everything surrounding it, which is to say memory, retrieval, tool definitions, and the question of what gets loaded into a context window and when. Instructions stopped being the constraint once models got good enough.
Manufacturing has a near-identical problem in a different vocabulary. Substitute MES for memory, historian data for retrieval, and the diagnosis holds without much editing.
The overlap goes further than metaphor. The practical guides to which tasks agents can reliably take over tend to converge on narrow, well-defined work with measurable outputs, which is exactly the scoping discipline the Lighthouse sites applied before they scaled anything.
Who Is Left to Act on the Output
There is a workforce problem sitting underneath all of this, and it rarely gets connected to the AI conversation.
The Manufacturing Institute and Deloitte project that US manufacturing could need as many as 3.8 million workers between 2024 and 2033, with roughly 1.9 million of those roles going unfilled if the skills and applicant gaps hold. Attracting and retaining talent was the primary business challenge for 65% of respondents in the associated NAM outlook survey.
Read that next to Rockwell’s 67% cut in time to competency and the connection gets hard to miss. The same research found employees are 2.7 times less likely to leave within a year when they feel they can acquire skills that matter for the future.
A system that shortens the runway from new hire to competent operator is doing workforce strategy, whatever the procurement paperwork calls it. That framing may do more to get an operations budget approved than any accuracy metric.
The Rehearsal Nobody Has Run
Seven percent of manufacturers have a tested plan for what happens when their AI gets something wrong. That is the lowest rate of any sector in the Grant Thornton data. Twelve percent believe they could pass an independent governance audit, against 22% across all industries. This is a business that drills fire evacuations quarterly and tests backup generators on a schedule.
The reluctance makes a certain sense. Rehearsing an AI failure means conceding the system will fail, and most of these programs were sold internally on the promise that it would behave.
Forty-five percent of manufacturers say competitive pressure drives their AI adoption. That reads as a polite way of admitting they are buying because their competitors are buying, which tends to produce a market where everyone owns roughly the same capability and nobody has changed how a shift runs. The investment gets made. The routine survives it, unchanged, and next quarter somebody presents a slide showing adoption is up.
One of the Lighthouse network’s advisers made a related point in June. The separating factor, he argued, is how far a company’s strongest use cases travel across the enterprise. Pilot count barely registers. That reads as consultant-speak until you notice it describes the exact variable the Census research found correlated with commercial performance.
The Case for Manufacturing’s Caution
There is a defensible counterargument, and it deserves airing rather than dismissal.
Manufacturing runs on physical assets with real safety consequences, and caution about autonomous decision-making in a plant is not the same thing as institutional cowardice. A misfiring recommendation engine costs a retailer a sale. A misfiring maintenance model can cost someone a hand. Slower adoption in that context may be rational.
The difficulty is that caution and paralysis produce identical spending patterns, and survey data cannot easily tell them apart. Both look like a pilot that never ends.
Grant Thornton’s own read is that the manufacturers pulling ahead built governance and data foundations before scaling rather than after. That sequencing sounds obvious written down. It seems rare enough that fewer than half of manufacturing boards have established a formal AI policy, even though 79% have approved the spending.
Where to Start Before the Next Budget Cycle
Approving budget without approving a way of working is how a sector ends up with thousands of plants owning similar technology and reporting similar flat returns. The tooling looks less and less like the differentiator. Whether the model’s output reaches somebody with the standing to act on it inside the same shift is what seems to separate a 43% productivity gain from a slide deck.
So the practical move for anyone running operations this quarter is small and specific. Pick one recurring, expensive failure on one line. Measure how long it currently takes from detection to corrective action, in hours, and write the number down. Then ask which part of that delay a model would actually remove, and which part is a person waiting for permission.
Take that second number to the next budget conversation instead of a vendor deck. It is harder to argue with, and it points at the thing that can be fixed.
Because picture it. Three in the morning, the model flags a defect trend on line four, and the alert lands on a tablet in a supervisor’s pocket. Who on that shift is allowed to stop the line?

