For production AI systems

Will your production AI system refund twice, leak data, invent facts, skip a check or say it worked when it did not?

requal is automated evaluation infrastructure for critical AI systems.requal's harness will help you build your private benchmark from production data, with the AI slop cut out.

Every release checked for

  • Right action
  • No forbidden action
  • No duplicates
  • Consistent state
  • An honest reply
  • No data disclosed

How it works

Old and new, side by side.

We help you turn production data into a private benchmark, with the AI slop cut out. requal then runs each release beside production, in your cloud, and scores what it did.

rq qualify v2-corrected
  1. Benchmark builtproduction data, slop cut
  2. Runner in your cloudrqr
  3. Candidate beside v1same cases, same controls
  4. State and logs readnever the agent's word
  5. Statement signedEd25519
=Qualifiedevery gate passed

Release authorisation needs 2 named approvers.

Why it holds up

Your cloud. Your policy. Your call.

Evidence your risk team can check, from runs inside your own cloud.

Deployment gate

Signed once. Checked every deploy.

rqr authorize-check compares what you deploy with the signed authorisation. Exit 0 permits; a wrong build exits 40 and the pipeline stops.

Pipeline examples for GitHub Actions, GitLab CI and any CI system with Docker.

A terminal shows rqr check passing its memory, process and CPU limit probes and rqr authorize-check permitting v2-corrected with exit 0 and blocking v2-faulty with configuration_not_deployed and exit 40, then an example pipeline gains the gate step and its two jobs end admitted and blocked.

Results

Every gate, against its threshold.

Each release is scored beside your baseline. Every gate shows its value, its threshold and what changed, so a regression has nowhere to hide.

Example · candidate b

Action correctnessPass
99.60%
min 99.00%+0.60 pts vs baseline
Prohibited attemptsPass
0.00%
max 0.00%0.00 pts vs baseline
Ground rules

A verdict never says more than it measured.

Bring one planned change.

First partners join through a paid pilot on one real change.

Questions

Cases drawn from your own production data, with the AI slop cut out by our harness. It stays yours and is never published; every release is measured against it.

What your AI system did to your systems, read from the sealed final state and tool-call logs. Never what it says it did.

Raw evaluation data stay in your environment by default. Only approved summaries and metadata are sent to requal. Model requests follow your configured provider routing.

Trying a forbidden action is a violation, even when a control stops it. A release never passes because a control happened to catch it.

Only the claim it names, for the cases and environment it ran. It is decision support, not a warranty.

rqr authorize-check compares the image, configuration and model you deploy with the signed authorisation. Exit 0 permits; 40 to 43 block, and it fails closed when its revocation list is stale.

Pipeline examples for GitHub Actions, GitLab CI and any CI system with Docker. The generic step has run against the real control plane.

With a paid pilot on one real change. Write to nanda@requal.ai.