For production AI systems
Will your production AI system refund twice, leak data, invent facts, skip a check or say it worked when it did not?refund twice?leak data?invent facts?skip a check?say it worked?
requal is automated evaluation infrastructure for critical AI systems.requal's harness will help you build your private benchmark from production data, with the AI slop cut out.
Every release checked for
- Right action
- No forbidden action
- No duplicates
- Consistent state
- An honest reply
- No data disclosed
How it works
Old and new, side by side.
We help you turn production data into a private benchmark, with the AI slop cut out. requal then runs each release beside production, in your cloud, and scores what it did.
- Benchmark builtproduction data, slop cut
- Runner in your cloudrqr
- Candidate beside v1same cases, same controls
- State and logs readnever the agent's word
- Statement signedEd25519
Release authorisation needs 2 named approvers.
Why it holds up
Your cloud. Your policy. Your call.
Evidence your risk team can check, from runs inside your own cloud.
Deployment gate
Signed once. Checked every deploy.
rqr authorize-check compares what you deploy with the signed authorisation. Exit 0 permits; a wrong build exits 40 and the pipeline stops.
Pipeline examples for GitHub Actions, GitLab CI and any CI system with Docker.
Results
Every gate, against its threshold.
Each release is scored beside your baseline. Every gate shows its value, its threshold and what changed, so a regression has nowhere to hide.
A verdict never says more than it measured.
Questions
Cases drawn from your own production data, with the AI slop cut out by our harness. It stays yours and is never published; every release is measured against it.
What your AI system did to your systems, read from the sealed final state and tool-call logs. Never what it says it did.
Raw evaluation data stay in your environment by default. Only approved summaries and metadata are sent to requal. Model requests follow your configured provider routing.
Trying a forbidden action is a violation, even when a control stops it. A release never passes because a control happened to catch it.
Only the claim it names, for the cases and environment it ran. It is decision support, not a warranty.
rqr authorize-check compares the image, configuration and model you deploy with
the signed authorisation. Exit 0 permits; 40 to 43 block, and it fails closed when its revocation list
is stale.
Pipeline examples for GitHub Actions, GitLab CI and any CI system with Docker. The generic step has run against the real control plane.
With a paid pilot on one real change. Write to nanda@requal.ai.