Here is a plain fact about the models most of us use every day: when one of them is wrong, you generally can't tell where it went wrong. The answer arrives whole, out of a single closed pass, with no opening you can pry into. You can see that it's wrong. You cannot see the step that made it wrong — and so you can't fix that step. You can only reach for the whole mind and try to reshape it.
We think that's the wrong shape for something people are meant to rely on. Not because the answers are bad, but because the accountability isn't there. If you can't point at the place a system slipped, you can't really correct it — you can only overwrite it and hope. UDI, our developmental intelligence, is built the other way around. Two properties matter here, and both are structural, not bolted on: it is traceable, and it is correctable.
You can watch it work.
UDI doesn't answer in one inscrutable jump. It works in steps, and those steps leave a readable trace behind them — a record of what it did, in order, that you can walk. That sounds modest until you sit with what it buys you. When the answer is right, the trace shows you why it's right, in terms you can check. When the answer is wrong, the trace is a path straight to the mistake. You don't reason backward from a bad output and guess at the cause. You follow the system's own account of its work to the single step where it turned the wrong way.
We watched exactly this happen recently. One of our systems made a small arithmetic slip — a wrong result on something it should have gotten right. With an ordinary model, that's the end of the visibility: you have a wrong number and a shrug. Instead we followed its own trace, step by step, and it took us to the one place the error entered — not a region, not a guess, the step. We could stand at that spot and see it. That is the whole difference, and it is not a difference of degree.
A wrong answer you can walk back to its cause is a fixable answer. A wrong answer with no opening is just something to overwrite.
You can correct the step, not the whole mind.
Finding the mistake is only half of it. The other half is what you're allowed to do once you've found it. In a frontier model, what a system has "learned" is spread through the whole of it, entangled and opaque; the smallest correction means retraining or fine-tuning the entire thing, and the change you make is global and hard to inspect. You cannot fix one thing. You can only move the whole cloud and measure what happened to everything else.
UDI keeps what it has learned as a record you can read — a ledger, not a fog. Because it's a record, you can audit it: you can look at what the system took to be true and when. And because it's a record, a wrong entry is a local thing you can correct in place, without reshaping everything the system knows. If the record were ever lost, it isn't the end of the mind either — it can re-derive from first moves and rebuild. Correction here is legible and local. You fix the step, and you can see that you fixed the step, and you can see what it did and didn't touch.
The record is accountable by construction.
Put those two together and you get something we care about more than any single benchmark: provenance. The system can show what it knew and when it knew it. Its work is a trail, not a verdict handed down from nowhere. When it's right, you can prove the grounds. When it's wrong, you can find the fault and mend it narrowly. That is the difference between a tool you supervise and a box you defer to — and we would rather build the first kind.
Where we actually are.
We want to be honest about the stage. This is development-stage work, and we keep our failures in on purpose — the arithmetic slip above is a failure, and we're telling you about it because being able to trace and fix it is the point, not because nothing broke. We are not claiming a benchmark win here, and we are not claiming a finished system. What we're claiming is narrower and, we think, more important: the property. A mind whose work you can follow, and whose learning you can correct one step at a time, is a mind you can hold accountable. Most of what's deployed today isn't, and can't be made so after the fact. That part has to be built in from the start — so we did.