The Untrainable Corner of Industrial Commerce | ReshapeX

[Skip to main content](#main-content)[Skip to navigation](#navigation)[Skip to footer](#footer)

June 18, 2026

Juan Aparicio

# The Untrainable Corner of Industrial Commerce

[← Back to blog](/en/insights)

Sarah Guo published an essay last week about what she calls the mid-2026 investor despair. The logic of the despair goes like this: the models keep getting better at everything, so every company built on top of one is a thin wrapper waiting to be absorbed, and the only value that survives is compute and frontier weights. Put everything into Anthropic and Nvidia and go home.

She thinks the despair is half right. So do I. And the half that is wrong happens to describe the industry I work in better than anything else I have read this year.

Her core mechanism is worth restating plainly. Anything you can measure becomes a benchmark. Anything that becomes a benchmark gets trained against. Anything that gets trained against commoditizes. This is why coding fell to AI first: a compiler is a free verifier, a test suite is a free verifier, and when the answer checks itself for nothing, the labs can grind against the check until they beat it. Devin solved 13% of a standard software benchmark in 2024 and got dismissed. The best agents now sit in the high eighties.

But there is a number inside that story that almost everyone skipped. Researchers at MIT measured the actual effect of coding agents across more than 100,000 developers. The amount of code written went up roughly 180%. The amount that actually shipped went up about 30%. Writing got cheap. Everything else still runs through people, systems, and trust, and none of those move at the speed of a model release.

The conclusion she draws is the one that matters: the durable value in AI is sliding toward work whose correctness is private and expensive to establish, locked inside systems you have to be trusted to enter. She calls it the untrainable corner. A better model does not make private ground truth public. It does not hold the license, own the files, or sign its name to the outcome.

Now apply that test to industrial commerce.

Start with the verifier. Code has a free one. Industrial product knowledge has none. When an agent recommends a drive configuration for a specific motor load, there is no compiler to catch the error. It surfaces weeks later when an incompatible part arrives at a facility, or it never surfaces at all because the engineer trusted the confident answer. I wrote a full piece on this earlier in the spring, and Guo’s essay supplies the economic consequence: tasks without free verifiers are precisely the ones the labs cannot train their way through, because there is nothing to grind against.

Then the ground truth. Turck offers more than 5,000 proximity sensor part numbers in a single product category. Festo’s catalog runs past 33,000 SKUs. A serious automation distributor carries dozens of brands. Two part numbers that differ by one character can be incompatible for a specific application. Cross-references between brands are not symmetric. Compatibility depends on firmware revisions that are rarely documented in one place. And the most valuable layer, the exceptions and workarounds learned across hundreds of customer installs, exists nowhere except in the heads of people who have been doing this for fifteen years.

Distributors are not the only ones sitting on this kind of ground truth. Walk into any system integrator that has been in business for twenty years and ask where their proposal knowledge lives. The answer is a file server in the corner. Thousands of proposals, panel designs, sizing calculations, and commissioning notes, each one encoding a judgment call somebody made and defended: why this drive over that one for this washdown environment, why the safety architecture was structured that way, what failed at startup last time and what they changed. Every proposal that won is a record of what good engineering looked like for a specific customer with a specific constraint. None of it is structured. Almost none of it is searchable. And the engineer who wrote the best of it is retiring in three years.

The OEMs have the same problem from the other side. A machine builder with configurable product lines carries its own combinatorial catalog, its own compatibility rules, its own cross-reference battles against competitors, plus something the distributors do not: a fleet of machines in the field for twenty or thirty years, each one a slightly different configuration, each support call requiring someone who remembers what was shipped in 2009 and what the approved retrofit path is. That knowledge holds up entire aftermarket revenue streams, and it lives in the same place it lives everywhere else in this industry. In people.

That is private correctness in its purest form, distributed across three kinds of companies that together make up most of industrial commerce. No frontier model reaches any of it by being 2% better on a public leaderboard, because the knowledge was never public to begin with.

Then the walls. You do not get to verify whether AI works inside a distributor’s operation, an integrator’s proposal process, or an OEM’s support organization from the outside. You get inside after the security review, the catalog access, the contract with your name on the outcome. And then there is the deadbolt, which Guo identifies correctly as the user. A majority of American doctors now open OpenEvidence every day, and no amount of compute buys that habit. The industrial equivalent is the inside sales rep or applications engineer who has been burned once by a confident wrong answer and now checks everything manually. Trust in this industry is built interaction by interaction, and it transfers to whoever earned it, not whoever benchmarked best.

So industrial commerce sits about as deep in the untrainable corner as a market can sit. Which raises the obvious question: if the model cannot eat this, what does the work actually look like?

Guo’s answer is the most honest sentence in the essay. The job is the unglamorous translation: arranging a company’s private reality so a model can act on it, handing the model the tools to act, working with the customer as their reality changes. Domain-specialized engineers next to the customer, doing work that never fully ends.

That is not a description of software in the traditional sense. It is a description of knowledge construction. In our world it means an engineer sitting next to the person who knows the Siemens line better than anyone alive, or the estimator who has priced a thousand panels, or the OEM support veteran who can identify a 2011 machine configuration from a photo, and encoding what they know into a structured layer the agent can traverse with purpose-built tools instead of retrieving text and hoping. Every inferred relationship gets validated against manufacturer data and against the expert’s own judgment. The construction happens before the customer ever asks a question. The tools the agent holds are generated from that verified layer, which is why the answers hold up under real queries instead of just demo queries. And the translation continues for as long as catalogs change, proposals accumulate, and machines stay in the field, which is forever.

There is one more move in her essay, and it is the one I keep thinking about. If work cannot be scored from the outside, someone on the inside has to decide what a good answer even is. Enough of those decisions, written down, become the benchmark for the entire field. Harvey is doing this for legal work. OpenEvidence is settling what a safe clinical answer looks like. The authority to define good lands with whoever the field already uses, and it cannot be authored by a foundation lab, however smart the model gets, because that standing only exists inside the field.

Nobody has written down what good means for industrial product knowledge. There is no public standard for what an acceptable answer to a cross-reference question looks like, what precision threshold a BOM recommendation has to clear, or how you score an agent that needs to say "I do not know" instead of guessing on a safety-rated configuration. Every deployment we run involves thousands of evaluation queries against real catalogs before anything goes live, reviewed by people who sell, spec, and support these products for a living. That accumulating body of judgment, what this industry on this kind of question accepts as correct, is quietly becoming the thing Guo says the benchmark eventually becomes: the standard everyone else gets measured against.

The despair says the application layer is dead. The truth is narrower. The legible application layer is dead. The wrappers are getting absorbed. What survives is the work the models cannot score: private ground truth, earned trust, translation that never finishes, and the slow accumulation of written-down judgment about what good means in a specific corner of the world.

Industrial commerce, the distributors, the integrators, and the OEMs together, is one of the deepest corners there is. But sitting in a deep corner is not the same as winning it. Winning it has requirements, and they are hard ones.

You have to be grounded. Not a model with a clever prompt, but an agent reasoning over a structured knowledge layer built from the actual catalog, the actual compatibility rules, the actual proposal history.

You have to be right. Not 80%, not 90%. In an industry where a wrong part stops a line, the threshold is 99.9%, and the last few points are the hardest and the most expensive to reach.

You have to put people on the ground. An army of forward-deployed engineers who sit next to the rep, the estimator, and the support veteran and pull the knowledge out of their heads before it retires with them. There is no scraper for this. There is no shortcut.

And you have to come from the industry. You cannot translate a world you do not understand. The judgment about what good means in industrial commerce belongs to the people who have lived in it.

That is the work. That is the corner. That is us. ReshapeX

## Give us your twenty hardest questions.

We’ll demo on your SKUs, run your evals, and cite every answer.

*   Real Examples
*   Working Demo
*   Your Data

Schedule a meeting Talk to the agent first