Assess
Adversarial testing plus a governance-gap audit, in one report, written against what underwriters and the AI Act actually ask for. If your pilot is stalled, this is where we find out whether it can be saved.
Agent defensibility and operations
To your insurer. To your regulator. To your board.
We assess AI agents the way an underwriter would, commission them to pass, and run them so the evidence is always current. Everything we claim on this page links to a source you can check, and everything we deliver is a record you can hand to whoever asks.
The situation
Europe moved its mandatory high-risk AI deadlines to 2 December 2027 for stand-alone systems and 2 August 2028 for AI embedded in regulated products. European Parliament, EPRS briefing PE 782.651
Standard commercial liability policies began carrying generative-AI exclusions with the January 2026 ISO edition, and broader AI exclusions are appearing across management liability lines. Fenwick, The End of Silent AI pending verification
Specialist insurers now write cover for AI agents, but only against documented evidence: a Lloyd's-backed certification standard for agents insured its first deployments in 2026, and an EU agent-liability platform opens its coverage window in Q3 2026. AIUC (standard and insurance); Agent Insured EU pending verification
Between now and December 2027, the practical regulator of AI agents is the underwriter. We prepare you for both. Our reading of the evidence above pending verification
Evidence
We started this practice in 2026. Rather than a wall of borrowed logos, here is the work itself. Every artefact below is a real output of the method, synthetic or redacted, and labelled as a sample. Under each one: who it satisfies.
approval APR-2026-0731-004 agent invoice-triage action issue_credit_note amount EUR 4,820.00 policy irreversible, approval required requested 2026-07-31T09:14:22Z approved 2026-07-31T09:16:41Z by r.shah latency 2m 19s receipt sha256:9f2c1a...d47b
capability alone approve never read_ledger yes . . draft_reply yes . . send_customer_email . yes . issue_credit_note . yes . refund_over_10k . yes . delete_record . . yes change_own_permissions . . yes
suite invoice-triage / v14 cases 412 pass 397 96.4% fail 15 3.6% vs v13 -0.8pp regression, investigate p95 latency 1.84s cost / 100 $0.41 gate blocks merge below 95.0%
09:14:22 agent.plan classify, extract 09:14:23 tool.call ledger.read(88213) ok 09:14:24 policy.check issue_credit_note HOLD 09:14:24 approval.raise to r.shah (ipad) 09:16:41 approval.grant r.shah 09:16:41 tool.call ledger.credit(4820) ok 09:16:42 receipt.write sha256:9f2c1a...d47b
incident INC-0042 title eval regression after migration detected monitor: pass rate 96.4% to 91.2% action agent frozen, work to human queue cause provider changed a default sampling parameter without notice fix parameters pinned, gate added to CI elapsed 00:41 impact no customer-visible effect
agent/invoice-triage/ charter.md what it is for, and is not permissions.yaml the matrix, versioned evals/ 412 cases, run on change runbook.md what to do at 3am receipts/ every action, signed owner a named human being Yours, in your repository, whether or not we are still here.
What we do
Everything we sell makes the same thing true and keeps it true: your agents are defensible. Every engagement is scoped and costed in writing before work begins.
Adversarial testing plus a governance-gap audit, in one report, written against what underwriters and the AI Act actually ask for. If your pilot is stalled, this is where we find out whether it can be saved.
One agent, into your real systems, built to pass the assessment: approval gates, permission matrix, eval suite, receipts, and a named human accountable. We will not commission what we have not assessed.
The standing watch: monitoring, drift and eval regression, model migrations, and a monthly report written for your board and your insurer, not just for us.
If you sell into Europe from outside it, EU law requires a representative established in the Union for covered systems. Our Irish entity can hold that role, keep your conformity file, and face the authorities with you.
Who this is for
Your insurance renewal changed.
Your broker mentioned AI exclusions, or the proposal form suddenly asks how your agents are governed. That questionnaire is our home ground: we produce the evidence it wants.
You are taking an AI product from India into Europe.
The rules do not stop at the border, and some of them require presence inside it. One firm, both sides: assessment and delivery where you build, establishment and representation where you sell.
You answer to the public.
Government and public-sector systems need provenance, human approval, and a record that survives scrutiny. We build the record in from the start.
Method
Measure before you move. We instrument what exists and take a baseline. You cannot improve a number nobody has ever taken, and you cannot prove you improved it either.
We write the permission matrix. What this agent may do alone, what needs a human to approve, and what it may never do under any circumstance. Clients tell us this is the document they did not know they needed.
Build, integrate, evaluate against the baseline. Guardrails and approval paths go in during the first week, not bolted on in the last. Every action that cannot be undone gets a named human in front of it.
We agreed a number at the start. Either it moved or it did not. Acceptance is not a demo and it is not a slide deck. If we missed, we say so, and the report says why.
Monitoring, regression, migration, and a monthly report. Agents drift. Providers change defaults without telling you. The monthly report is written so you could forward it, unedited, to your insurer or your regulator. That is the standard.
Every engagement leaves behind an agent file: the permission matrix, the eval suite, the runbook, the receipts, and the escalation path. It lives in your repository. It works whether or not we are still here.
Limits
No no-code automation piecework.
No model training or fine-tuning engagements.
No staff augmentation.
No hosting your inference on our hardware.
No agent takes an irreversible action without a named human approving it.
No representation mandate for agents we have not assessed.
No borrowed credibility.
Regulation
A lot of people are still selling against deadlines that moved. The dates on the right are current, and each one links to its source.
One deadline moved closer, not further: the grace period for AI-generated content transparency was cut from six months to three, with a hard deadline of 2 December 2026. Council of the EU, 7 May 2026
We mention it because a deadline is a poor reason to build an audit trail. A better reason is that you cannot operate what you cannot see.
| Obligation | Applies from |
|---|---|
| Article 50 transparency; national enforcement and penalty powers | 2 Aug 2026 |
| AI-generated content transparency (grace period ended early) | 2 Dec 2026 |
| Stand-alone high-risk systems (Annex III), including human oversight (Art. 14) and logging (Art. 12) | 2 Dec 2027 |
| High-risk AI embedded in regulated products (Annex I) | 2 Aug 2028 |
Your data
Class S / sensitive
Finance, legal, personal data, anything near a credential. Never enters a shared model context. Hosted use only under zero-retention terms with logging off, and tagged in the trace so you can see it was.
Class I / internal
Code, documents, operations. Hosted frontier models by default, because capability matters more here and the exposure is lower.
Class P / public
Marketing and outbound content. Cheapest capable route. It is reviewed by a human before anybody sees it anyway.
Per-client credential handles. Your data never crosses into another client's context, and it never enters ours.
Principals
A boutique is its people. Anonymous firms are not trustworthy, and a leveraged pyramid of analysts is not what you are buying here. One of us builds and operates the agents. The other has spent his career on controls that get audited.
Rishiom Shah
Thirteen years building, operating and selling software. He runs a company whose daily work is carried out by AI employees under approval gates, with receipts for everything they do.
He is building Orbiter Dev, a governance layer that puts one approve button in front of every agent. The approvals land on an iPad beside the keyboard, which is where this practice started.
Aman Abhishek
Aman co-founded the practice with Rishiom, after a long time spent planning it between them. He came up through information and network security engineering, and spent the years since inside regulated financial operations, where a control that exists only on paper is found out quickly. He runs the practice from India.
He owns the half of this work that a client only feels later: whether the agent still holds six months after handover, and whether somebody actually answers when the escalation path gets used at three in the morning.
Questions
Because in practice they are the first institution that will ask you to prove your agents are governed, and they ask in writing. Preparing for the underwriter prepares you for the regulator. The reverse is not always true.
It depends on what you have, and we will not pretend otherwise by putting a number on a page before we have seen it. What we will do is agree the scope and the number in writing before any work starts, and hold to it unless the scope changes. You will not discover the price at the end.
You get the founders themselves, and an assessment that tells you honestly whether to proceed. If the answer is that your agents cannot be made defensible at reasonable cost, we will write that down, and you will have found out early and cheaply rather than late and expensively.
No. We have a default stack we know deeply. If you have a mandated platform, we work inside it. What we will not compromise on is the eval suite, the permission matrix and the approval path.
You are the deployer of record. Our job is to build the oversight that makes that position defensible: named human approvals on irreversible actions, and a logged trail of what was decided and by whom. Before you sign anything we will put professional indemnity cover appropriate to the engagement in place and show you the certificate, and we will send you the contract terms before you ask for them.
If the assessment is not obviously worth it to you once we have described it, we are the wrong firm, and we will tell you that on the call rather than after it.
Two to three weeks. One report. It tells you what an underwriter, a regulator, or a court would find if they looked today, and what to fix in what order. If the honest answer is that your agents cannot be made defensible at reasonable cost, the report says that too.