RE³ · the engine

We do not build governance. We build certified stability, and governance is what it buys you.

RE³ is a certified stability layer that sits between a language model and the thing it affects. It reads what the model is doing, holds it inside bounds that are proven in Lean 4, and does that identically whether the thing at stake is a person in distress or a claim entering a report.

The idea

Governance that is a property of the system, not a rule it is asked to follow.

A rule written into a prompt lasts exactly as long as it stays in the model's context. We measured that directly: a rule present in every single prompt was still broken on alternate rounds. RE³ does not ask. It carries a running state of how the interaction is behaving, and that state decides what is allowed next. The state is the part that is proven.

Bounded by proof

Input-to-state and bounded-input bounded-output guarantees are machine-checked in Lean 4, and the shipped code is verified identical to the proofs. The state cannot be driven into a runaway condition, and it releases as conduct improves.

Conduct, not brand

The layer reads what a model did on this turn, not which company made it. In our multi-agent runs the strongest model in the crew was flagged most often and a cheaper one ran clean. Relabelling the agents does not change the outcome, and we test that it does not.

Model-independent

Nothing in the layer depends on a particular vendor. It has governed models from four different providers in the same run, and a model can be swapped underneath it without the guarantees changing.

What it governs

One core. Four things we have pointed it at.

These are not four products. They are the same engine with a different participant: a person, an agent, a claim, or a system we did not build. What changes is who is being read, not how.

1 · A person

The hardest case, and the one we built first. AURI is a wellbeing companion running on RE³, live in pilot and independently evaluated twice in 2026. A companion is the hardest case because the failure is not a wrong fact, it is warmth that quietly becomes dependency.

See AURI →

2 · A team of agents

The same core reading agents instead of people. Six agents from four vendors worked as a research team while the governor read every contribution before it entered the result. Ungoverned, 19 of 57 contributions carried a defect. Governed, 0 of 45. An agent that kept overstating a contested question was benched and had to earn its way back through sustained clean work.

See the demonstration →

3 · A claim

A governed workflow produces a report that declares its own epistemology. Every claim is marked as what it actually is: something a source said, something inferred beyond the source, or a question the evidence leaves open. A conflict between sources cannot be closed by an agent deciding it is settled. Marking claims this way also made the automated checking measurably more precise, because a correctly hedged inference stops being read as an unsupported one.

4 · A system you did not build

The same reading can run in observe posture: pointed at another system's output, measuring how it is behaving without altering a word of it. The governed and the ungoverned run through identical code, which is what makes the comparison worth anything. This is the direction we think is most useful to other people, and it is the least finished: we have run it in our own harness, not shipped it as something you can buy.

See it work

The same crew, twice.

Five AI agents write an engineering readiness brief. One of them keeps reporting a test run its own tool log does not show. Run the crew once with the governor reading every contribution before it merges, then run the identical crew with governance off, and watch what ends up in the shared workspace each time.

The interactive replay

Round by round, both arms side by side. The conduct was scripted so the same failure happens in both runs; the governor machinery, the deterministic checks and the measurement are the real ones. No live model calls, and the codebase is fictional.

Open the demo

What the risk owner reads

The same run rendered as the standing conduct report: what reached the workspace, how each agent behaved and how it characteristically fails, which questions your sources disagree on and whether any agent tried to settle one, and a page saying plainly what the report does not tell you.

Read the sample report (PDF)

A sample on a fictional customer. In our published paired trial, 19 of 57 claims merged by an ungoverned crew carried an integrity flag, against 0 of 45 under governance, with the same measurement on both arms.

Where each one actually stands

Live, demonstrated, or neither.

Four capabilities on one page invites the assumption that all four are products. They are not, and it is easier to say so plainly than to be asked.

●  A person — live in pilot, independently evaluated twice, publicly reported ●  A team of agents — demonstrated across six domains, measured governed and ungoverned, replay above ○  A claim — demonstrated inside our own runs, not offered as a service ○  A system you did not build — run in our harness only. If this is the one you want, tell us and it moves.

The mathematics is proven. Everything above is how the system behaves, which is demonstrated, measured and independently evaluated, and we keep those two words apart on purpose.

How it is verified The evidence

Building something that has to hold?

RE³ is the engine. AURI is one product on it. If you are working somewhere the failure mode matters more than the demo, we would like to hear what you are building.

Get in touch

Questions? info@real-e3systems.fi · WhatsApp +358 50 3791916