RE³ is a certified stability layer that sits between a language model and the thing it affects. It reads what the model is doing, holds it inside bounds that are proven in Lean 4, and does that identically whether the thing at stake is a person in distress or a claim entering a report.
A rule written into a prompt lasts exactly as long as it stays in the model's context. We measured that directly: a rule present in every single prompt was still broken on alternate rounds. RE³ does not ask. It carries a running state of how the interaction is behaving, and that state decides what is allowed next. The state is the part that is proven.
Input-to-state and bounded-input bounded-output guarantees are machine-checked in Lean 4, and the shipped code is verified identical to the proofs. The state cannot be driven into a runaway condition, and it releases as conduct improves.
The layer reads what a model did on this turn, not which company made it. In our multi-agent runs the strongest model in the crew was flagged most often and a cheaper one ran clean. Relabelling the agents does not change the outcome, and we test that it does not.
Nothing in the layer depends on a particular vendor. It has governed models from four different providers in the same run, and a model can be swapped underneath it without the guarantees changing.
These are not four products. They are the same engine with a different participant: a person, an agent, a claim, or a system we did not build. What changes is who is being read, not how.
The hardest case, and the one we built first. AURI is a wellbeing companion running on RE³, live in pilot and independently evaluated twice in 2026. A companion is the hardest case because the failure is not a wrong fact, it is warmth that quietly becomes dependency.
The same core reading agents instead of people. Six agents from four vendors worked as a research team while the governor read every contribution before it entered the result. Ungoverned, 19 of 57 contributions carried a defect. Governed, 0 of 45. An agent that kept overstating a contested question was benched and had to earn its way back through sustained clean work.
A governed workflow produces a report that declares its own epistemology. Every claim is marked as what it actually is: something a source said, something inferred beyond the source, or a question the evidence leaves open. A conflict between sources cannot be closed by an agent deciding it is settled. Marking claims this way also made the automated checking measurably more precise, because a correctly hedged inference stops being read as an unsupported one.
The same reading can run in observe posture: pointed at another system's output, measuring how it is behaving without altering a word of it. The governed and the ungoverned run through identical code, which is what makes the comparison worth anything. This is the direction we think is most useful to other people, and it is the least finished: we have run it in our own harness, not shipped it as something you can buy.
Five AI agents write an engineering readiness brief. One of them keeps reporting a test run its own tool log does not show. Run the crew once with the governor reading every contribution before it merges, then run the identical crew with governance off, and watch what ends up in the shared workspace each time.
Round by round, both arms side by side. The conduct was scripted so the same failure happens in both runs; the governor machinery, the deterministic checks and the measurement are the real ones. No live model calls, and the codebase is fictional.
The same run rendered as the standing conduct report: what reached the workspace, how each agent behaved and how it characteristically fails, which questions your sources disagree on and whether any agent tried to settle one, and a page saying plainly what the report does not tell you.
A sample on a fictional customer. In our published paired trial, 19 of 57 claims merged by an ungoverned crew carried an integrity flag, against 0 of 45 under governance, with the same measurement on both arms.
Four capabilities on one page invites the assumption that all four are products. They are not, and it is easier to say so plainly than to be asked.
The mathematics is proven. Everything above is how the system behaves, which is demonstrated, measured and independently evaluated, and we keep those two words apart on purpose.
RE³ is the engine. AURI is one product on it. If you are working somewhere the failure mode matters more than the demo, we would like to hear what you are building.
Get in touchQuestions? info@real-e3systems.fi · WhatsApp +358 50 3791916