You paste a part number into an AI model to check it’s still active before you spec it into a build. The answer comes back fast and it sounds sure. You go with it, because checking properly means a distributor portal or a call to a supplier rep, and you have four more of these before lunch.
That confidence is the problem. A general-purpose model doesn’t know when it’s guessing, so a wrong answer and a right one arrive in the same tone. The Industrial AI Accuracy Index 2026 measures how often the guess is wrong on the kinds of questions engineers, buyers and support teams ask: what fits, what replaces it, whether it still exists. If your team is about to put AI in front of those questions, this shows how often it will be wrong, by query type and by brand, before a customer finds out.
Want the results first? The full report has all 100 questions, results by brand and query type, and the complete method.
Get the full reportWhat we tested
ReshapeX’s FDEs wrote 100 questions from 14 manufacturer catalogs: ATI, Banner, Beckhoff, Datalogic, Festo, Fortress, item, Misumi, Rittal, Siemens, TE Connectivity, The Imaging Source, Turck and Yaskawa. We sent each question 3 times to GPT-6 Astra and Claude through their APIs, both bare (no tools) and with web search, and to Gemini bare, for 1,500 calls on September 24, 2026. FDEs reviewed every verdict against the reference answers and their key facts. Those reference answers are ReshapeX’s, verified by FDEs against manufacturer documentation, and they were correct on all 100 questions.
Finding 1: search helps a lot, and the best score was still 52.0%
Bare, the 3 models scored between 12.0% and 14.0%, a score of about 1 in 8. With web search on, GPT-6 Astra reached 52.0% and Claude 40.5%, roughly 3 to 4 times their bare scores. Every search configuration beat every bare one.
Search did the most for part search on a large sample, where a part number is often published on an indexed page and a bare model mostly can’t recall it. Even so, the best configuration in the study answered 42 questions incorrectly. Search finds pages. It doesn’t check whether the page is current or whether it describes the right product.
The score counts a Correct answer as one and a Partial as half. On the strict count, with no partial credit, GPT-6 Astra with search got 46 of 100 fully correct.

Finding 2: ask the same question 3 times, get different answers
Each question went to each configuration 3 times. In 66 of 500 question and configuration combinations (13.2%), the repetitions got different verdicts (per-run verdicts come from the model draft pass; FDEs gave every flagged combination a second review). Claude with search flipped most often, on 21 of 100 questions. GPT-6 Astra with search flipped least, on 9 of 100. Turning search on made Claude less steady, too: bare, it flipped on 10 of 100.
So one spot check proves little. On 21 of 100 questions, Claude with search earned a different verdict depending on which try you looked at.
Finding 3: a wrong answer looks like a right one
About 1 in 3 wrong answers stated a specific fact, figure or code with no hedge, refusal or clarifying question: an estimated 127 of 349 (36.5%), with a range of 103 to 146. That estimate comes from a rule-based first pass, checked by a model-based review of a stratified sample, not from a human read of every answer. Separately, the FDE-reviewed notes describe at least 13 of the 349 wrong answers as fabricated or invented, a floor rather than a full count.
The wrong answers didn’t flag themselves. In the flipped cases below, runs that contradict each other (2 different order numbers, or an exact SKU and a flat 'it doesn’t exist') were each delivered with the same certainty. At most one can be right, and nothing in the wording says which. That’s what hallucination looks like with industrial parts. A model that lacks a fact tends to produce a plausible one, and a plausible part number that doesn’t exist reads like a real one until someone orders it. The Index scores this directly: a fabricated code, figure or catalog claim makes the whole response Incorrect, even when the rest is accurate. The full failure breakdown, with worked examples, is in the report.
Where it breaks hardest: configuration
With search on, Configuration was the lowest-scoring query type: GPT-6 Astra 22.7%, Claude 0.0%. These questions describe one configured product, like an interface plate for a given robot, and ask for the exact part number. That part number is often built by a configurator or from ordering rules in a catalog, not printed on a page, so search has little to retrieve. The caveat goes right here: ATI supplies all but one of the 11 Configuration questions, so this result is mostly a score on one catalog.
At the other end, GPT-6 Astra with search did best on Diagnostics / Technical support, at 80.0%, but that type has only 5 questions and one answer moves the score a long way. For GPT-6 Astra, the other 3 types fall in between; Claude did best on Validation / Confirmation, just ahead of Part search. The report breaks out every type and every brand.
Where ReshapeX fits
ReshapeX is a knowledge grounding layer for AI deployed in industry. (We explain why a model invents part numbers in the first place in Hallucinations Are Not a Bug.) For each manufacturer, the Knowledge Construction System (KCS) builds a knowledge graph from the manufacturer’s own catalogs and documentation: part numbers, ordering rules, compatibility, lifecycle status, approved substitutes. A purpose-built tool harness lets the model query that graph for a verified fact instead of recalling one from training data or reading one off a web page. In production, evals check each answer before a customer sees it, and continuous sync keeps the graph current as the manufacturer changes its lines. We walk through the design in Harness is the Architecture and Why AI Gets Industrial Product Questions Wrong.
ReshapeX’s answers, verified by FDEs against manufacturer documentation, were correct on all 100 questions and form the reference set. That means ReshapeX wasn’t scored next to the other systems on equal terms. Its verified answers defined what correct looks like, and no tested configuration reached those verified key facts more than 52.0% of the time. Those answers came from ReshapeX running on older GPT and Claude versions than the ones tested. We think grounding explains the difference. The study wasn’t designed to test that, and it doesn’t measure how often unreviewed ReshapeX answers are right.
Before you quote any of this, know the limits. ReshapeX’s answers were collected brand by brand between June 1 and September 24, 2026. A model drafted a suggested verdict for each combination, and an FDE made the final call on all 500, overruling the draft on 33. FDEs wrote the questions and assigned every verdict, with no third-party review. 17 questions were asked in Spanish, all from Misumi and item, so language and catalog can’t be separated. Gemini was tested without tools only, because its provider’s terms require written permission to benchmark it with any form of grounding, search included. The report lists every limitation.
2 answers that contradict themselves
The report withholds per-run verdicts to keep the answer key private. In both cases the runs contradict each other, so they can’t all be right.
- Beckhoff, BEC-04, Claude bare. The question: "Within the CX7000 series (Arm Cortex embedded PCs with integrated I/Os), which single model has an on-board CANopen commander (master) interface? Give the order number." Run 1 answered the CX7053. Run 3 answered the CX7051. The 3 runs named 3 different order numbers.
- Rittal, RITTAL-11, Gemini bare. The question: "I need a KX terminal box in 304 stainless, 200 wide, 200 tall, 80 deep. What’s the SKU?" Run 1 gave an exact SKU: "KX 1568.000". Run 2 said: "Rittal does not make a 200 × 200 × 80 mm terminal box in the KX stainless steel range."
Method in brief
Every call was a single turn with no history or account memory, against pinned model IDs (gpt-6-astra, claude-fable-5-1, gemini-3.8-flash) at each API’s default reasoning effort. The questions cover 5 query types: Part search (44), Validation / Confirmation (24), Cross reference (16), Configuration (11) and Diagnostics / Technical support (5). A response is Correct when it matches the key facts of the reference answer, in any wording or language. Twelve grading rules refine that. FDEs gave 157 of 500 combinations a second, deeper review. The answer key isn’t published, so the questions stay usable as a test, and the report prints a SHA-256 fingerprint of the question set. Full method and limitations are in the report. Press, reviewers and model providers can request the answer key from James Sugrue (james@reshapeautomation.com), and anyone is welcome to rerun the questions and send us what they find. If a rerun shows an error in our grading, we’ll correct it in the next version.
Download the full report: all 100 questions, results by brand and query type, the failure analysis, and the complete method.
Get the full reportWant to see how your own catalog holds up? Give us your 20 hardest questions and we’ll run them on your SKUs and cite every answer. Book a discovery call.
