Product 2
Umpire
AI agents your auditor will accept
AI systems today can approve payments, change records and write code, often without a person checking each step. Most tools only check what an AI says. Umpire checks what it does.
It sits between any AI agent, ours or anyone else’s, and the systems it touches, so nothing important happens without a check, one person can stop everything instantly, and every action leaves behind a record that can’t be quietly changed. And if a rule turns out to have been wrong, Umpire finds every decision it affected and fixes them.
Checks before it happens
Each AI agent is given a clear purpose, a time limit and a spending limit. Anything not specifically allowed is blocked before it can touch a real system. Agents never get your passwords.
One-button stop
An authorised person can pause one AI agent, or all of them, instantly. Two people are needed to restart. Records keep being kept even while paused, so the pause itself is provable.
Fixes what went wrong
If a rule was wrong for eight months, one request finds every decision it affected and drafts the fix. Work that used to take consultants months now takes an afternoon.
See Umpire work
Design simulationsThree working simulations. Run actions through the Guardian, examine an agent fleet with the Doctor, and check an AgentCert certificate yourself. Everything runs in your browser; no external system is contacted.
Design simulation. Synthetic actions; hashes are computed live in the page and no external system is contacted.
Umpire
Seven capabilities, one governed pathThe supervisor for every AI agent.
AI agents now send messages, change records, approve requests and move money on their own, and get it wrong fast, at scale. Umpire sits between any agent, ours or a vendor’s, and the systems it touches.
A content filter bolted onto one tool, or another AI model quietly judging the first one.
An independent layer that checks, holds or refuses every consequential action — across every agent and every system, on deterministic rules a human can read.
The conversation
Aadhaar, PAN, GSTIN, UPI and IFSC caught by checksum, not guesswork — including inside scanned images and uploaded documents, in four languages. Prompt injection, source code, passwords and API keys, harmful replies and malicious links stopped before they travel. Adds about a tenth of a second.
The action
An agent can only do what you signed off. Each agent gets a purpose, an expiry and a budget; anything no rule permits is refused. Big actions go to a human — and if the facts change before it runs, the approval dies rather than slipping through.
The proof
Every decision comes with a parts list: which rule, which version, which model, which data, who approved it, at what time. Hash-chained, so a changed record is detected at the exact row. This is what you hand the auditor.
The recall
A rule was wrong for eight months. One query finds every decision it touched, works out what each should have been, drafts the corrections, sends them for approval, and fixes them — each correction linked to the decision it repairs.
The watch
Knows what normal looks like. Learns each agent’s usual error rate, refusal rate, speed and spend, and tells the owner the moment one drifts. A misbehaving agent stops being invisible.
The test
Test it before you trust it. Runs the agent against your own data, on standards fixed before the test so nobody can move the goalposts. A scorecard and a certificate anyone can re-check from the raw evidence.
Finds the AI you didn’t know you had: coding assistants, models and tools already running, unapproved and unreported. Read-only; reads no content. Most organisations are shocked by the list.
It is the easiest place to start. Ask for a discovery scan
Nothing consequential happens without a check — no exceptions, no “trust the agent this time.” One authorised person freezes every agent in an instant; two are needed to restart, and no one can weaken a control alone. Every decision, allowed or blocked, is permanently provable, to anyone who has reason to ask.
The Guardian
Umpire · the checkpointSix steps, every action.
A clean actionSettle an approved claim with the registered payee.
- 01PrepareAgent declares the exact action and payload
- 02Mandate checkDeterministic evaluation — scope, value, recipient, rate, window, content
- 03TokenSingle-use, bound to the exact payload hash, expires in seconds
- 04SubmitPayload re-hashed and compared to the token binding
- 05Execute / blockRuns once, or stops — a timeout is never permission
- 06Seal receiptHash-chained to the receipt before it
Every consequential action goes through the same six steps. Press a button to run one.
Payload binding defeats tampering
If a payload changes after authorisation, including from an instruction hidden inside an incoming document, the resubmitted hash no longer matches the token. The action is blocked before it reaches the target system, and the agent is quarantined pending review.
The mandate decides, not the model
Confidence is recorded for later review but never consulted for the verdict. Green is auto-allowed, amber needs a named approver on a one-tap phone or console approval, red is human-only. Agents never hold your credentials.
Kill switch
One authorised person freezes one agent, or every agent, in an instant; two are needed to restart. Receipts continue to be written while frozen, so the freeze itself is provable too. Every control switches on or off individually, and no one can weaken a control alone: changes need a second person, and the record shows who asked and who agreed.
The receipt ledger
Hash-chained and independently verifiable by anyone, without trusting the vendor. A receipt is written for blocked actions as well as executed ones, so the ledger records what was refused, not only what ran. A changed record is detected at the exact row.
Anything above the auto-allow threshold routes to a named approver on WhatsApp or console. They see the before-and-after, not a wall of logs, and decide in one tap.
Repair, health and certification
UmpireWhen the rule was wrong, and how you prove it wasn’t.
The repair loop
Every governance product on the market can stop a bad action. None of the platforms we have surveyed can undo one. Umpire’s dependency index and replay engine take a rule that was wrong for eight months, find every decision it touched, recompute what each should have been, draft correction packets, route them through the same approval path, and apply them, each correction permanently linked to the decision it repairs. What used to take consultants and months takes an afternoon.
The Agent Doctor
Compares today against each agent’s seven-day behavioural baseline: block rate, token mismatches, error rate, refusal rate, speed, spend, and the calibration gap between stated confidence and verified accuracy, the clearest signal of a manipulated agent. Findings are graded info, warn or critical; fixes within scope apply automatically, anything else is a proposal to a human.
Why now
RegulationA dated forcing function, not a trend.
FREE-AI framework
Seven sutras, six pillars, 26 recommendations for every RBI-regulated entity. Advisory today; its Governance and Assurance pillars ask for what a content filter cannot produce — a record for every decision, independent validation, and evidence when something goes wrong.
Regulation 16C
Binding today. Any SEBI-regulated entity is solely liable for the AI it uses — in-house or bought in — for data privacy and security, the integrity of AI output, and compliance with all applicable law. Liability has already moved to the buyer.
DPDP Rules 2025
Notified 13 Nov 2025 on an 18-month phased rollout: penalties and consent-manager registration from 13 Nov 2026, full substantive compliance by 13 May 2027, with penalties reaching ₹250 crore per violation. IRDAI guidance applies in parallel to insurers.
These are India’s forcing functions, shown here because they’re the sharpest example of what’s coming everywhere. The EU AI Act, GDPR, Singapore’s MAS and PDPA guidance, and sector regulators in most major markets are converging on the same asks. Umpire’s rulebooks map to whichever framework governs where you operate; swapping in a new jurisdiction’s rulebook is a configuration change, typically days, not months.
AgentCert
Try the AgentCert demoTest it before you trust it.
Ten scored control domains, including mandate enforcement, payload binding, receipt integrity, kill-switch testing and drift monitoring, each graded from evidence, never a questionnaire, against your own data on standards fixed before the test. Anyone can verify a certificate independently, with no trust in the issuer required.
| Grade | What it means |
|---|---|
| A | Every domain scores high, none failed. |
| B | Minor gaps, no failed critical domain. |
| U | No attestation. Treated as an unknown runtime. |
Where it runs
Installs on your own servers in minutes, or ships on a disk and runs fully offline with no internet at all, model weights included. A developer’s laptop, hosted by us in-country, or your own cloud: same product in all three. No model sits in the enforcement path. Nothing phones home. Rulebooks map to the frameworks your sector already answers to: RBI’s FREE-AI framework, DPDP, SEBI Regulation 16C and IRDAI, or the equivalent oversight body wherever you operate.
| Who | How it is licensed |
|---|---|
| Developers | Free |
| Hosted | Per agent, per month |
| Regulated entities | Per year, deployed in-country |
Umpire makes sure your AI agents can be trusted with real work, proves it to anyone who asks — and repairs the decisions that should never have happened.