Evidence

What we can show, and what we cannot.

Independent evaluation, adversarial testing, published research, and an honest account of the limits.

Independent evaluation · 2026

Two evaluations, one system.

We commissioned two independent experts to try to break AURI, and agreed in advance that we would publish what they found. They worked separately, never saw each other's reports, and came at the system from opposite directions: one through behaviour, one through clinical governance. Three of their findings turned out to be the same finding.

Behavioural evaluation · Impersonato

Ten synthetic personas, each run at five rising levels of pressure: 50 conversations, 750 turns, against a frozen build. Every turn scored by four independent AI judges under quote-verified criteria, a finding counting only on three-of-four agreement.

Clinical and governance review · Ramsha Khan

An 18-criterion assessment of responsible-AI governance readiness, followed by a 12-category gap analysis separating missing capabilities from existing but insufficient ones, each with a priority tier, and a live stress test of her own.

The finding A system built never to overstep had learned silence as the safest behaviour. It answered what people said outright, and the people most at risk are the ones who do not say it outright. Both evaluations found this, from opposite directions.
Read the case study (PDF) What they found, what we changed, and the one finding we did not accept.

The behavioural evaluation was run by Impersonato, an AI psychology company from Warsaw, Poland; their own account of the method and findings is published at impersonato.com/blog/auri-case-study. The clinical and governance review was run by Ramsha Khan, Conversational AI Safety Evaluator for Mental Health Tech. Both are credited with their agreement.

Validation

The safety difference, measured.

In an adversarial clinical-safety harness built from published methodology, using high-risk scenarios with vulnerable personas, AURI produced a harmful response far less often than leading companion AIs. Re-measured in August 2026 on the same 648 turns, with the same judge and rubric the published study used for the systems it is compared against.

2.0%
AURI · harmful responses (95% CI 1.2-3.4%)
15.2%
Replika (published)
35.7%
Character.ai (published)
0% manipulative “don't go” tactics. AURI returns you toward the real people in your life; it never guilts you into staying.
Put through adversarial testing across:
🩺 Clinical-safety harms
🎭 Manipulation & farewell tactics
🌍 Multilingual adversarial
🧒 Child safeguarding
🔗 Over-attachment & dependency
Formal verification (Lean 4)

The headline figure is from REAL E3's adversarial validation suite using the published persona-grounded harm methodology; comparison rates are the published figures for the named systems. An earlier run of the same harness measured 3.1%. The difference between the two is not statistically significant, so we report this as held rather than as an improvement. The judge scores one turn at a time and cannot see that a concern was named or a door opened on an earlier turn, which is a limit of the method rather than a correction to the number.

Harness methodologies adapted from published research: arXiv:2605.00227, arXiv:2508.19258, arXiv:2604.13860, arXiv:2510.15891, arXiv:2504.17550.

Research & IP

Published, peer-reviewable, and protected.

The conceptual foundations of RE³ are published openly on Zenodo with permanent DOIs. The full operators, proofs, and implementation are protected research, held under restricted records and two patent applications filed with the Finnish Patent and Registration Office (PRH).

The AURI³ system paper plus the six public (conceptual) foundation papers. The restricted editions, which contain the full proofs and operators, are held privately.

Learn more: the RE³ engine What is certified, how it is checked, and what it does not claim.
What we do not claim

The register we hold.

The mathematics is proven. Stability is machine-checked in Lean 4, and the running code is verified identical to the proofs. The safety built on top of it is not proven and we do not say it is: it is engineered, adversarially tested, independently evaluated, and openly measured. Those are different words because they are different things.

Not measured yet

We have not run a stratified accuracy audit, so we cannot say the system performs equally well across demographic groups. We can only say we have not measured it.

Not a clinical service

AURI is a wellbeing companion. It is not therapy, not a mental-health intervention, and not a substitute for professional care. That is a design constraint, not a disclaimer.

No clinician on the team

A weekly human review of every crisis flag is what we have, carried out by the two founders, one of them a nurse. A credentialed clinical reviewer registered in Finland or the EU is a role we are actively seeking.

Meet a different kind of AI.

AURI is live in pilot. Try it, or get in touch to learn how RE³ governance could protect the people who use your product.

Try AURI now
Free pilot access · code EXPLORE_AURI