We asked four AIs the same 100 industrial parts | ReshapeX

[Skip to main content](#main-content)[Skip to navigation](#navigation)[Skip to footer](#footer)

August 27, 2026

ReshapeX Team

# We asked four AIs the same 100 industrial parts questions

[← Back to blog](/en/insights)

You paste a part number into ChatGPT to check if it’s still active before you spec it into a build. It answers fast and it sounds sure. You use the answer anyway, because checking it properly means digging through a distributor portal or waiting on a supplier rep, and you have four more of these before lunch.

That confidence is the problem. A general-purpose AI model doesn’t know when it’s guessing, so a wrong answer and a right one arrive in the same tone.

We wanted to know how often that guess is wrong on the questions our customers actually ask: is this part discontinued, what replaces it, does this drive match that motor, will this configuration actually work together. So we ran the same 100 questions, across 14 industrial brands (ATI, Banner, Beckhoff, Datalogic, Festo, Fortress, item, Misumi, Rittal, Siemens, TE Connectivity, The Imaging Source, Turck, and Yaskawa), through our own knowledge grounding layer and through ChatGPT, Gemini, and Claude. Our engineers graded every answer against the verified correct answer: the right part number, the right item ID, the right lifecycle status, the right spec.

ReshapeX got all 100 right. The three general-purpose models landed between 47% and 54%, and none of them cleared 10% on the questions that asked them to assemble a compatible set of components.

## Why general models miss on industrial questions

This isn’t a knock on how these models are built. They’re trained to be broadly useful, not to hold a live, discontinued-when, replaced-by-what map of every industrial catalog. When a general model doesn’t have the fact, it doesn’t say so. It produces the most plausible-sounding answer instead, which is the core problem with using an ungrounded model for parts and specs: a hallucinated part number reads exactly like a real one until someone orders it.

The pattern showed up hardest in the categories that require chaining several facts together instead of recalling one. Part search, going from a spec to a part number across 44 questions, put the general models between 52% and 61%: better than a coin flip, worse than you’d want to build a bill of materials on. Validation, confirming a part is correct or still current, was their best category at 59% to 71%, and that’s still a real miss rate on a question that’s supposed to be a quick check. Cross reference, finding an equivalent to another supplier’s part or a legacy part, dropped to 38% to 44%. Configuration, assembling a set of compatible components, is where it fell apart: 0% to 9%. That’s where guessing compounds. Get one component wrong and the whole assembly is wrong, and a model with no ground truth to check against has no way to know it happened.

ReshapeX held 100% across every one of those categories, on every brand. The answer isn’t generated from a pattern in training data. It comes from a knowledge graph built from the actual catalogs for each brand: drive configurations, fieldbus module compatibility, firmware revisions, approved substitutes. Evals check it before it reaches you or your customers, and continuous sync keeps it current as manufacturers update their lines.

## What this means for the person doing the asking

The engineers and planners who do this lookup today, by habit, by portal, or by calling a rep, already know a wrong part number is expensive: a delayed line, a return, a redesign. They’re the ones who catch it when a general-purpose model hands back a guess with the same confidence as a fact.

That’s the gap a grounding layer closes. It gives them an answer they don’t have to run down a portal to check five times a day, so their judgment goes where it’s actually needed: a nonstandard application, a tradeoff between two viable substitutes, a customer relationship. The expertise is the hard part to build in this business. This is about putting it to use faster, not about who’s doing the asking.

## See the full breakdown

If you want the brand-by-brand and category-by-category numbers behind this, they’re in the graphic below. It’s the same 100 questions, the same 14 brands, and the same four systems, scored the same way.

![Bar chart comparing ReshapeX, ChatGPT, Gemini, and Claude accuracy across part search, configuration, cross reference, diagnostics, and validation query types, and across 14 industrial brands. ReshapeX scores 100% in every category and every brand; the three general-purpose models range from 0% to 100% depending on category and brand, averaging 47% to 54% overall.](/images/blog/ai-industrial-accuracy-index-comparison-graphic.png)

Want to see how this performs against your own catalog and your own questions? Schedule a meeting via the agent on our homepage and we’ll show you.

## Give us your twenty hardest questions.

We’ll demo on your SKUs, run your evals, and cite every answer.

*   Real Examples
*   Working Demo
*   Your Data

Schedule a meeting Talk to the agent first