Internal review copy. Not for public distribution.

Agent defensibility and operations

AI agents you can defend.

To your insurer. To your regulator. To your board.

We assess AI agents the way an underwriter would, commission them to pass, and run them so the evidence is always current. Everything we claim on this page links to a source you can check, and everything we deliver is a record you can hand to whoever asks.

Book a scoping call See the work first

Every claim on this page carries its source

No client logos, because we have no clients yet

Ireland and India


The situation

The AI Act was delayed. Your insurer was not.

Europe moved its mandatory high-risk AI deadlines to 2 December 2027 for stand-alone systems and 2 August 2028 for AI embedded in regulated products. European Parliament, EPRS briefing PE 782.651, 2026-06-16

Standard commercial liability policies began carrying generative-AI exclusions with the January 2026 ISO edition, and broader AI exclusions are appearing across management liability lines. Fenwick, The End of Silent AI, 2026-07-23 pending verification

Specialist insurers now write cover for AI agents, but only against documented evidence: a Lloyd's-backed certification standard for agents insured its first deployments in 2026, and an EU agent-liability platform opens its coverage window in Q3 2026. AIUC (standard and insurance); Agent Insured EU, 2026-05-15 pending verification

Between now and December 2027, the practical regulator of AI agents is the underwriter. We prepare you for both. Our reading of the evidence above, 2026-07-23 pending verification


Evidence

We have no client logos to show you yet.

We started this practice in 2026. Rather than a wall of borrowed logos, here is the work itself. Every artefact below is a real output of the method, synthetic or redacted, and labelled as a sample. Under each one: who it satisfies.

Sample / approval receipt APR-0731-004 what your board asks for
approval  APR-2026-0731-004
agent     invoice-triage
action    issue_credit_note
amount    EUR 4,820.00
policy    irreversible, approval required
requested 2026-07-31T09:14:22Z
approved  2026-07-31T09:16:41Z by r.shah
latency   2m 19s
receipt   sha256:9f2c1a...d47b
Sample / permission matrix PMX-invoice-triage what your underwriter prices
capability            alone  approve  never
read_ledger             yes        .      .
draft_reply             yes        .      .
send_customer_email       .      yes      .
issue_credit_note         .      yes      .
refund_over_10k           .      yes      .
delete_record             .        .    yes
change_own_permissions    .        .    yes
Sample / eval scorecard EVL-v14 what your renewal questionnaire wants
suite       invoice-triage / v14
cases       412
pass        397   96.4%
fail        15    3.6%
vs v13      -0.8pp regression, investigate
p95 latency 1.84s
cost / 100  $0.41
gate        blocks merge below 95.0%
Sample / audit trail TRC-88213 what a court admits
09:14:22  agent.plan     classify, extract
09:14:23  tool.call      ledger.read(88213)  ok
09:14:24  policy.check   issue_credit_note HOLD
09:14:24  approval.raise to r.shah (ipad)
09:16:41  approval.grant r.shah
09:16:41  tool.call      ledger.credit(4820) ok
09:16:42  receipt.write  sha256:9f2c1a...d47b
Sample / incident record INC-0042 what a regulator respects
incident  INC-0042
title     eval regression after migration
detected  monitor: pass rate 96.4% to 91.2%
action    agent frozen, work to human queue
cause     provider changed a default sampling
          parameter without notice
fix       parameters pinned, gate added to CI
elapsed   00:41
impact    no customer-visible effect
Sample / agent file AGF-invoice-triage what survives us leaving
agent/invoice-triage/
  charter.md        what it is for, and is not
  permissions.yaml  the matrix, versioned
  evals/            412 cases, run on change
  runbook.md        what to do at 3am
  receipts/         every action, signed
  owner             a named human being

Yours, in your repository, whether or
not we are still here.

What we do

One property, four stages.

Everything we sell makes the same thing true and keeps it true: your agents are defensible. Every engagement is scoped and costed in writing before work begins.

01

Assess

Adversarial testing plus a governance-gap audit, in one report, written against what underwriters and the AI Act actually ask for. If your pilot is stalled, this is where we find out whether it can be saved.

  • Adversarial test results
  • Governance-gap audit
  • Permission matrix draft
  • Remediation roadmap
02

Commission

One agent, into your real systems, built to pass the assessment: approval gates, permission matrix, eval suite, receipts, and a named human accountable. We will not commission what we have not assessed.

  • One agent in service
  • Approval paths on every irreversible action
  • Regression-gated eval suite
  • Runbook and escalation path
03

Operate and attest

The standing watch: monitoring, drift and eval regression, model migrations, and a monthly report written for your board and your insurer, not just for us.

  • Continuous monitoring and drift detection
  • Eval regression on every model change
  • Model migrations handled
  • Monthly attestation report
04

Represent

We take no representation mandate for agents we have not assessed.

If you sell into Europe from outside it, EU law requires a representative established in the Union for covered systems. Our Irish entity can hold that role, keep your conformity file, and face the authorities with you.

  • Authorised-representative mandate
  • Conformity file custody
  • Authority correspondence
  • Compliance monitoring

EU law requires providers outside the Union to appoint an authorised representative established inside it for covered AI systems; the representative verifies the conformity file, holds documentation for ten years, and faces the authorities. EU AI Act, Article 22 (AI Act Explorer), 2026-07-23 pending verification


Who this is for

Three situations we are built for.

Your insurance renewal changed.

Your broker mentioned AI exclusions, or the proposal form suddenly asks how your agents are governed. That questionnaire is our home ground: we produce the evidence it wants.

Insurance renewals increasingly arrive with AI exclusions or governance questionnaires attached. Fenwick, The End of Silent AI, 2026-07-23 pending verification

You are taking an AI product from India into Europe.

The rules do not stop at the border, and some of them require presence inside it. One firm, both sides: assessment and delivery where you build, establishment and representation where you sell.

You answer to the public.

Government and public-sector systems need provenance, human approval, and a record that survives scrutiny. We build the record in from the start.


Method

We sell rigour, so the method is the product.

I

Sounding

Measure before you move. We instrument what exists and take a baseline. You cannot improve a number nobody has ever taken, and you cannot prove you improved it either.

II

Marking the line

We write the permission matrix. What this agent may do alone, what needs a human to approve, and what it may never do under any circumstance. Clients tell us this is the document they did not know they needed.

III

Commissioning

Build, integrate, evaluate against the baseline. Guardrails and approval paths go in during the first week, not bolted on in the last. Every action that cannot be undone gets a named human in front of it.

IV

Acceptance

We agreed a number at the start. Either it moved or it did not. Acceptance is not a demo and it is not a slide deck. If we missed, we say so, and the report says why.

V

Standing watch

Monitoring, regression, migration, and a monthly report. Agents drift. Providers change defaults without telling you. The monthly report is written so you could forward it, unedited, to your insurer or your regulator. That is the standard.

Every engagement leaves behind an agent file: the permission matrix, the eval suite, the runbook, the receipts, and the escalation path. It lives in your repository. It works whether or not we are still here.


Limits

What we will not do.

No no-code automation piecework.

Wrong altitude, and the bottom of that market is saturated with people who will do it cheaper than we would.

No model training or fine-tuning engagements.

Not our stage, and mostly not your problem. The failures we see are integration and governance failures.

No staff augmentation.

We sell outcomes, not hours. If you want bodies, there are better firms and they are cheaper.

No hosting your inference on our hardware.

Your data stays in infrastructure you control or rent. We are not a cloud.

No agent takes an irreversible action without a named human approving it.

This is not negotiable, and it is the one thing we will lose an engagement over.

No representation mandate for agents we have not assessed.

A statutory role is not a rubber stamp.

No borrowed credibility.

No logo walls, no invented metrics, no case studies we did not do. When we have clients, you will see them here with their permission.


Regulation

What the EU AI Act actually requires, and when.

A lot of people are still selling against deadlines that moved. The dates on the right are current, and each one links to its source.

One deadline moved closer, not further: the grace period for AI-generated content transparency was cut from six months to three, with a hard deadline of 2 December 2026. Council of the EU, 7 May 2026, 2026-05-07

We mention it because a deadline is a poor reason to build an audit trail. A better reason is that you cannot operate what you cannot see.

We are not lawyers and this is not legal advice. Talk to your counsel.

ObligationApplies from
Article 50 transparency; national enforcement and penalty powers 2 Aug 2026
AI-generated content transparency (grace period ended early) 2 Dec 2026
Stand-alone high-risk systems (Annex III), including human oversight (Art. 14) and logging (Art. 12) 2 Dec 2027
High-risk AI embedded in regulated products (Annex I) 2 Aug 2028

Dates per the Digital Omnibus as adopted mid-2026. Cells link to sources.


Your data

Where your data goes, in three lines.

Class S / sensitive

Finance, legal, personal data, anything near a credential. Never enters a shared model context. Hosted use only under zero-retention terms with logging off, and tagged in the trace so you can see it was.

Class I / internal

Code, documents, operations. Hosted frontier models by default, because capability matters more here and the exposure is lower.

Class P / public

Marketing and outbound content. Cheapest capable route. It is reviewed by a human before anybody sees it anyway.

Per-client credential handles. Your data never crosses into another client's context, and it never enters ours.


Principals

Two people. Both of them will be on your call.

A boutique is its people. Anonymous firms are not trustworthy, and a leveraged pyramid of analysts is not what you are buying here. One of us builds and operates the agents. The other has spent his career on controls that get audited.

Rishiom Shah

Principal. Ireland.

Thirteen years building, operating and selling software. He runs a company whose daily work is carried out by AI employees under approval gates, with receipts for everything they do.

He is building Orbiter Dev, a governance layer that puts one approve button in front of every agent. The approvals land on an iPad beside the keyboard, which is where this practice started.

Aman Abhishek

Principal. India.

Aman co-founded the practice with Rishiom, after a long time spent planning it between them. He came up through information and network security engineering, and spent the years since inside regulated financial operations, where a control that exists only on paper is found out quickly. He runs the practice from India.

He owns the half of this work that a client only feels later: whether the agent still holds six months after handover, and whether somebody actually answers when the escalation path gets used at three in the morning.


Questions

The awkward ones, first.

Why do you keep talking about insurers?

Because in practice they are the first institution that will ask you to prove your agents are governed, and they ask in writing. Preparing for the underwriter prepares you for the regulator. The reverse is not always true.

What does an engagement cost?

It depends on what you have, and we will not pretend otherwise by putting a number on a page before we have seen it. What we will do is agree the scope and the number in writing before any work starts, and hold to it unless the scope changes. You will not discover the price at the end.

You have no clients. Why would I be the first?

You get the founders themselves, and an assessment that tells you honestly whether to proceed. If the answer is that your agents cannot be made defensible at reasonable cost, we will write that down, and you will have found out early and cheaply rather than late and expensively.

Do we have to use your stack?

No. We have a default stack we know deeply. If you have a mandated platform, we work inside it. What we will not compromise on is the eval suite, the permission matrix and the approval path.

Who is liable if the agent does something expensive?

You are the deployer of record. Our job is to build the oversight that makes that position defensible: named human approvals on irreversible actions, and a logged trail of what was decided and by whom. Before you sign anything we will put professional indemnity cover appropriate to the engagement in place and show you the certificate, and we will send you the contract terms before you ask for them.

How small is too small?

If the assessment is not obviously worth it to you once we have described it, we are the wrong firm, and we will tell you that on the call rather than after it.


Start with an assessment.

Two to three weeks. One report. It tells you what an underwriter, a regulator, or a court would find if they looked today, and what to fix in what order. If the honest answer is that your agents cannot be made defensible at reasonable cost, the report says that too.

Book a scoping call Read the artefacts again

Every action an agent takes, on a line you drew, in a record you can read.