← Blog
July 9, 2026· Lee Konstanty, Product and Partnerships at arbitrcompanyai governancetrust layercontent operations

Why we built arbitr

The rule that keeps AI safe was written in 1980: automate the work, not the judgment, and arbitr is the platform that finally enforces it, everywhere AI touches your business.

Thesis: your experts set the bar. Your AI has to clear it. Nothing ships until it does.

The AI that works in the demo and fails in production

Ask almost any leadership team about AI in 2026 and you will hear two things in the same breath: it is the most important technology shift of their careers, and most of what they have tried has quietly gone nowhere.

The numbers back up the mood, and they come from independent studies using different methods that land on the same place. Gartner predicts that over 40% of agentic AI projects will be canceled by the end of 2027, citing escalating costs, unclear business value, and inadequate risk controls. S&P Global Market Intelligence reported that the share of companies abandoning most of their AI initiatives jumped to 42% in early 2025, up from 17% a year earlier. MIT's Project NANDA, in its 2025 State of AI in Business study, found that only about 5% of integrated pilots were capturing real value. The exact figures are debated. The direction is not: most enterprise AI is not reaching production value.

It is tempting to read those figures as a verdict on the models. It is not. The models are extraordinary and getting better on a monthly cadence. The failure is happening one layer up, at the point where AI output meets something real: a live website, an outbound email, a contract, a regulated claim, a price, a customer, an action taken on your behalf.

Picture the specific failures. A product claim that overstates what a regulator allows, published to a live page in nine markets before anyone reads it. A price that generates correctly and localizes wrong. A support agent that takes an action on a customer's account it was never meant to take. Each one is fine in the demo, because a human is watching. Each one fails in production, because at machine speed and machine scale no human can watch every output, and there is no automated standard deciding what is safe to ship.

That is the gap. Not intelligence. Judgment, and control over what goes live.

The uncomfortable part is that this is not a new discovery. The industry reached this exact conclusion once before, 45 years ago, and then spent decades forgetting it.

The gap already has a name

In October 1980, the computational linguist Martin Kay circulated a report from Xerox PARC titled "The Proper Place of Men and Machines in Language Translation." His field was in one of its recurring winters. The promise of fully automatic, human-free systems had been made loudly and had failed publicly, and the mood was that the whole endeavor might be a dead end.

Kay's response was not to declare the technology hopeless. It was to reject a false choice. Full automation, he argued, "stands no chance of filling actual needs," but the answer was not to throw out the machine. It was to build "cooperative man-machine systems," where the machine did the relentless, repetitive work and remembered everything, while the human stayed in command of the judgment.

Automate the work, not the judgment. He wrote it down in 1980.

Kay happened to be talking about language, because that was his domain. The principle has nothing to do with language specifically. It applies anywhere a machine is trusted to produce something that has consequences: where do you let the machine move fast, and where must a human-defined standard hold? That is now one of the central questions in AI. Kay framed it clearly in 1980. What he lacked were the tools to build it properly, and those took another four decades.

We did not read this in a book. We lived it.

We know the principle holds because we have spent years operating at the frontier of the one domain where the stakes of getting language wrong are unambiguous and expensive: professional localization for global brands, manufacturers, and financial institutions. That work produced two things most AI companies do not have.

The first is proof. For decades, the only way to move fast without shipping garbage was exactly Kay's model: let the machine do the work, keep expert judgment as the bar, and never ship what has not cleared it. We did not theorize that augmentation beats automation. We ran a business on it.

The second is real technology, not a thin layer over someone else's model. At arbitr, we build our own small language models and our own agents, purpose-built for specialist domains. Our models start from a foundation trained on more than a decade of expert human data across ten languages, and include specialist models built from the ground up for demanding sectors like financial services, and, tellingly, a model whose entire job is to score another model's output for accuracy and adequacy and flag where it falls short. That last one matters: it is Kay's principle turned into a model. The bar is not a slogan. It is something we train, run, and stand behind. We are not wrapping a frontier model and hoping. We build our own models and agents, we are an IBM partner, and we build the layer that governs all of it.

Which brings us to the box.

Breaking a thirty-year-old box

Here is the lesson of those decades, and it is almost always misread.

The value a mature operation builds was never the software it ran on. It was the asset underneath: years of expert-approved decisions, domain-specific terminology, brand rules, compliance constraints, and the models and data trained on all of it. That asset compounds. It is among the most valuable things an organization owns.

For thirty years, legacy tooling locked that asset in a box. It served one narrow use, sat behind an access layer designed before the modern web, and could not be extended, repointed, or moved. Customers were frozen at whatever the tool did the year they bought it, paying to maintain aging infrastructure while the asset inside it stayed trapped.

The investment was not a mistake. It produced something genuinely valuable. It has just been trapped behind primitive technology, and the job now is to set it free and point it at a far larger problem than the one it was built for.

The same problem, now everywhere

Because that larger problem is no longer about translation, and it is no longer about any single domain. AI turned Kay's question into the daily operating question of the whole business.

You might be here to build and deploy a custom model. To create content at a scale no review team can keep up with. To localize or dub across dozens of markets. To audit whether what your systems produce actually complies with your rules and your regulators'. To let agents take real actions on your behalf. These look like different projects. They are the same project. Every one of them puts a machine in a position to produce or do something with consequences, faster than a human can check it, and every one of them lives or dies on the same thing: can you let the machine do the work while guaranteeing a human-defined standard is met before anything ships or acts?

Whatever brought you here, that is the problem arbitr is built to solve. Translation is simply where we proved we can solve it.

The ladder

We did not build arbitr as a better version of any one tool. The tools are rungs on a single ladder. The ladder measures one thing: how much judgment the machine is trusted with, and how much is at stake if it is wrong.

Reuse of approved work → adaptation → content creation → governed content creation → risk management for agentic systems

You do not have to start at the bottom. You can enter the ladder wherever your problem is today. What matters is that every rung asks Kay's question with higher stakes than the last, and that legacy tooling answered it for exactly one rung and then locked the door, and that arbitr is built to be the vehicle for the whole climb.

What arbitr actually is

If the problem is judgment and control, the product has to be the layer that sits between AI and everything real, enforcing a standard you define. That is what arbitr is: the trust layer between AI and your live systems.

Today it governs content and publishing. It connects to your CMS, your content and localization stacks, and the publishing surfaces where AI output goes live, through APIs, MCP, and whatever comes next. It checks every change, whether a human or a machine made it, against the rules you approved: your locked fields, your regulated phrasing, your brand terms, your market and compliance requirements. Safe changes flow straight through. Risky ones are pulled aside with the reason and routed to the right person. Approval becomes the exception, not the bottleneck. The same control is designed to extend outward, to email, to your other systems, and up the ladder to what agents are allowed to do. Your experts set the bar. Your AI has to clear it before anything ships.

Three commitments follow, and each is a deliberate reaction to how the old box failed:

Governance, not gatekeeping. AI has democratized building. Anyone can spin up an agent or a model now. The differentiation is no longer who is allowed to build. It is who can hand people the guardrails to build safely, and arbitr exists to enable that, not to stand in front of it.

Sovereignty as architecture, not a feature. In regulated environments, data residency and auditability cannot be a checkbox bolted onto someone else's cloud. The standard has to be able to run where your data already lives, with a full audit trail of every decision, structured so it can serve as evidence when a regulator asks. That is why arbitr is a control layer over your systems, not another destination you pour your work into.

An open platform, not a box. Legacy tooling locked a generation of customers into a fixed feature set for thirty years. We are building the opposite: frequent releases, integrations that reach wherever your systems already are, and the ability to stand up models and agents specific to your needs rather than the handful someone else decided to ship, so arbitr is built to grow toward where you are going, instead of freezing you where the tool happened to be the year you bought it.

Where the journey has arrived, and where it goes next

This is not a description of a product that might exist one day. The working platform is live, and arbitr today is a content governance platform for AI publishing, ready to use at getarbitr.straker.ai. That is the honest starting line: the place the journey has actually reached, ready to walk.

The top of the ladder, full risk management for agentic systems, is the direction: a world where the same discipline that checks a price, a claim, or a translation also governs what an autonomous agent is allowed to do, with the human-defined standard, the audit trail, and the data sovereignty carried all the way up. We are not going to point you at a finished future and pretend. We are telling you it is where this goes, and that the reason to start now is that the standard you set at the content layer is the same standard that will govern your agents.

Forty-five years ago, at a moment that looked a great deal like this one, Martin Kay told an industry rushing toward full automation that it had the balance wrong: let the machine do the work, and keep the judgment human. He was right, and the tools to do it properly did not exist yet. Now they do, and the thirty-year-old box that trapped the idea is finally worth breaking.

Your experts set the bar. Your AI has to clear it. Nothing ships until it does. That is why we built arbitr.

Timeline: the arc of an idea

Definitions

Sources