Skip to content
Book a proof run

Product 2

Umpire

AI agents your auditor will accept

AI systems today can approve payments, change records and write code, often without a person checking each step. Most tools only check what an AI says. Umpire checks what it does.

It sits between any AI agent, ours or anyone else’s, and the systems it touches, so nothing important happens without a check, one person can stop everything instantly, and every action leaves behind a record that can’t be quietly changed. And if a rule turns out to have been wrong, Umpire finds every decision it affected and fixes them.

Checks before it happens

Each AI agent is given a clear purpose, a time limit and a spending limit. Anything not specifically allowed is blocked before it can touch a real system. Agents never get your passwords.

One-button stop

An authorised person can pause one AI agent, or all of them, instantly. Two people are needed to restart. Records keep being kept even while paused, so the pause itself is provable.

Fixes what went wrong

If a rule was wrong for eight months, one request finds every decision it affected and drafts the fix. Work that used to take consultants months now takes an afternoon.

See Umpire work

Design simulations

Three working simulations. Run actions through the Guardian, examine an agent fleet with the Doctor, and check an AgentCert certificate yourself. Everything runs in your browser; no external system is contacted.

Open full screen

Design simulation. Synthetic actions; hashes are computed live in the page and no external system is contacted.

Umpire

Seven capabilities, one governed path

The supervisor for every AI agent.

AI agents now send messages, change records, approve requests and move money on their own, and get it wrong fast, at scale. Umpire sits between any agent, ours or a vendor’s, and the systems it touches.

What it is not

A content filter bolted onto one tool, or another AI model quietly judging the first one.

What it is

An independent layer that checks, holds or refuses every consequential action — across every agent and every system, on deterministic rules a human can read.

Firewall

The conversation

Aadhaar, PAN, GSTIN, UPI and IFSC caught by checksum, not guesswork — including inside scanned images and uploaded documents, in four languages. Prompt injection, source code, passwords and API keys, harmful replies and malicious links stopped before they travel. Adds about a tenth of a second.

Guardian

The action

An agent can only do what you signed off. Each agent gets a purpose, an expiry and a budget; anything no rule permits is refused. Big actions go to a human — and if the facts change before it runs, the approval dies rather than slipping through.

Open the Guardian demo

Receipts

The proof

Every decision comes with a parts list: which rule, which version, which model, which data, who approved it, at what time. Hash-chained, so a changed record is detected at the exact row. This is what you hand the auditor.

Open the Guardian demo

Repair

The recall

A rule was wrong for eight months. One query finds every decision it touched, works out what each should have been, drafts the corrections, sends them for approval, and fixes them — each correction linked to the decision it repairs.

Doctor

The watch

Knows what normal looks like. Learns each agent’s usual error rate, refusal rate, speed and spend, and tells the owner the moment one drifts. A misbehaving agent stops being invisible.

Open the Doctor demo

AgentCert

The test

Test it before you trust it. Runs the agent against your own data, on standards fixed before the test so nobody can move the goalposts. A scorecard and a certificate anyone can re-check from the raw evidence.

Open the AgentCert demo

Discovery: the scan

Finds the AI you didn’t know you had: coding assistants, models and tools already running, unapproved and unreported. Read-only; reads no content. Most organisations are shocked by the list.

It is the easiest place to start. Ask for a discovery scan

Nothing consequential happens without a check — no exceptions, no “trust the agent this time.” One authorised person freezes every agent in an instant; two are needed to restart, and no one can weaken a control alone. Every decision, allowed or blocked, is permanently provable, to anyone who has reason to ask.

The Guardian

Umpire · the checkpoint

Six steps, every action.

A clean actionSettle an approved claim with the registered payee.

  1. 01PrepareAgent declares the exact action and payload
  2. 02Mandate checkDeterministic evaluation — scope, value, recipient, rate, window, content
  3. 03TokenSingle-use, bound to the exact payload hash, expires in seconds
  4. 04SubmitPayload re-hashed and compared to the token binding
  5. 05Execute / blockRuns once, or stops — a timeout is never permission
  6. 06Seal receiptHash-chained to the receipt before it

Every consequential action goes through the same six steps. Press a button to run one.

Payload binding defeats tampering

If a payload changes after authorisation, including from an instruction hidden inside an incoming document, the resubmitted hash no longer matches the token. The action is blocked before it reaches the target system, and the agent is quarantined pending review.

The mandate decides, not the model

Confidence is recorded for later review but never consulted for the verdict. Green is auto-allowed, amber needs a named approver on a one-tap phone or console approval, red is human-only. Agents never hold your credentials.

Kill switch

One authorised person freezes one agent, or every agent, in an instant; two are needed to restart. Receipts continue to be written while frozen, so the freeze itself is provable too. Every control switches on or off individually, and no one can weaken a control alone: changes need a second person, and the record shows who asked and who agreed.

The receipt ledger

Hash-chained and independently verifiable by anyone, without trusting the vendor. A receipt is written for blocked actions as well as executed ones, so the ledger records what was refused, not only what ran. A changed record is detected at the exact row.

Maker-checker, at machine speed

Anything above the auto-allow threshold routes to a named approver on WhatsApp or console. They see the before-and-after, not a wall of logs, and decide in one tap.

Repair, health and certification

Umpire

When the rule was wrong, and how you prove it wasn’t.

The feature the category is missing

The repair loop

Every governance product on the market can stop a bad action. None of the platforms we have surveyed can undo one. Umpire’s dependency index and replay engine take a rule that was wrong for eight months, find every decision it touched, recompute what each should have been, draft correction packets, route them through the same approval path, and apply them, each correction permanently linked to the decision it repairs. What used to take consultants and months takes an afternoon.

The watch

The Agent Doctor

Compares today against each agent’s seven-day behavioural baseline: block rate, token mismatches, error rate, refusal rate, speed, spend, and the calibration gap between stated confidence and verified accuracy, the clearest signal of a manipulated agent. Findings are graded info, warn or critical; fixes within scope apply automatically, anything else is a proposal to a human.

Why now

Regulation

A dated forcing function, not a trend.

RBI · 13 Aug 2025

FREE-AI framework

Seven sutras, six pillars, 26 recommendations for every RBI-regulated entity. Advisory today; its Governance and Assurance pillars ask for what a content filter cannot produce — a record for every decision, independent validation, and evidence when something goes wrong.

SEBI · in force 10 Feb 2025

Regulation 16C

Binding today. Any SEBI-regulated entity is solely liable for the AI it uses — in-house or bought in — for data privacy and security, the integrity of AI output, and compliance with all applicable law. Liability has already moved to the buyer.

DPDP · 13 Nov 2026 / 13 May 2027

DPDP Rules 2025

Notified 13 Nov 2025 on an 18-month phased rollout: penalties and consent-manager registration from 13 Nov 2026, full substantive compliance by 13 May 2027, with penalties reaching ₹250 crore per violation. IRDAI guidance applies in parallel to insurers.

Penalties from 13 Nov 2026

These are India’s forcing functions, shown here because they’re the sharpest example of what’s coming everywhere. The EU AI Act, GDPR, Singapore’s MAS and PDPA guidance, and sector regulators in most major markets are converging on the same asks. Umpire’s rulebooks map to whichever framework governs where you operate; swapping in a new jurisdiction’s rulebook is a configuration change, typically days, not months.

Test it before you trust it.

Ten scored control domains, including mandate enforcement, payload binding, receipt integrity, kill-switch testing and drift monitoring, each graded from evidence, never a questionnaire, against your own data on standards fixed before the test. Anyone can verify a certificate independently, with no trust in the issuer required.

GradeWhat it means
AEvery domain scores high, none failed.
BMinor gaps, no failed critical domain.
UNo attestation. Treated as an unknown runtime.

Where it runs

Installs on your own servers in minutes, or ships on a disk and runs fully offline with no internet at all, model weights included. A developer’s laptop, hosted by us in-country, or your own cloud: same product in all three. No model sits in the enforcement path. Nothing phones home. Rulebooks map to the frameworks your sector already answers to: RBI’s FREE-AI framework, DPDP, SEBI Regulation 16C and IRDAI, or the equivalent oversight body wherever you operate.

WhoHow it is licensed
DevelopersFree
HostedPer agent, per month
Regulated entitiesPer year, deployed in-country
In one line

Umpire makes sure your AI agents can be trusted with real work, proves it to anyone who asks — and repairs the decisions that should never have happened.