Playbook

How to govern AI agents in finance

Govern an AI agent by deciding how much autonomy it has, proving it works before production and operating it after go-live. Five autonomy levels, six controls and what we do not claim.

Author
SabanaTech, Delivery team
Published
Reading time
6 min

To govern an AI agent in finance, decide how much autonomy it has, prove it works before it touches production, and keep operating it after go-live. Governance is not a committee that approves each agent. It is a set of controls the agent cannot ship without, and an operating routine that keeps them true.

Every governance conversation about an agent comes back to one question: who is accountable when it acts? The 2026 Deloitte research on AI readiness suggests few organizations can answer it yet. Only 5% say their processes are highly prepared for AI agents, and 75% of executives agree that human collaboration with agents creates more value than agent automation alone (https://www.deloitte.com/us/en/about/press-room/deloitte-survey-examines-ai-readiness-agentic-ai-success.html).

Both findings point the same way: the constraint is no longer the model. It is whether the process, the controls and the people are ready. This article sets out how we approach that for finance processes such as accounts payable, reconciliations and close.

What are the levels of autonomy for an AI agent?

Autonomy is not on or off. We use five levels, and each agent in each process is assigned one, in writing. The level can rise as the evidence supports it, and it can fall.

  1. Assist. The agent helps a person do the work, for example by summarizing a document or finding a policy. The person does everything that matters.
  2. Recommend. The agent proposes a decision with its reasoning, such as an account code for an invoice. A person decides every time.
  3. Act with approval. The agent prepares and executes the action, but only after a named person approves it. Payments and changes to master data usually start here.
  4. Act and report. The agent acts on its own and reports what it did, and people review a sample and every exception. Use it for reversible actions.
  5. Act within bounds. The agent acts alone inside limits set in advance: amount, counterparty, document type, policy. Anything outside the bounds goes to a person.

Most finance processes should not start above level 3. Move up one level at a time, and only when the evaluation results and the exception data justify it.

Which controls must be in place before an agent goes to production?

Six controls, none optional. If one is missing, the agent does not ship.

  1. Evaluation set. A frozen set of real cases with a pass mark, built before the agent. Every prompt or model change is scored against it. In our AP agent case, the golden set is 500 invoices.
  2. Least privilege. The agent receives only the tools and permissions the process requires: read this queue, post to that ledger. No open browser, no shared administrator credentials.
  3. Audit trail. Every prompt, tool call, output and approval is logged with the identity of the operator and the agent version. If you cannot replay a decision, you cannot defend it.
  4. Version pinning. Prompt and model versions are fixed and change only on a schedule, with evaluation. In our AP agent case they are pinned quarterly, so any result traces to the version that produced it.
  5. Kill switch. One configuration change pauses the agent immediately, and someone has tested it. Pausing an agent should be an ordinary operation, not an emergency.
  6. Named owner. A business owner, not the developer, is accountable for the agent's behavior. The owner approves its autonomy level and its bounds.

What does AgentOps involve after go-live?

AgentOps is the routine that keeps an agent inside its bounds once it is running. Models change, documents drift and policies are updated, so an agent that passed in March can fail in September without anyone touching it.

  • Monitor: track exception rate, escalation volume, cost per transaction and evaluation scores on live samples. Set thresholds that trigger a review.
  • Review: sample decisions on a fixed cadence, including the ones the agent handled alone. Feed errors back into the evaluation set.
  • Change with evaluation: score every prompt, model or tool change against the evaluation set before it reaches production, and roll back if it regresses.
  • Re-certify: revisit the autonomy level, bounds and owner in a quarterly review. Autonomy is earned again, not granted once.
  • Retire: remove agents and tools that nobody uses or owns.

This is where governance most often lapses. The build is funded and the operation is not. Budget AgentOps as a recurring cost from the start, the same way you would budget maintenance for any system your close depends on.

How much human review should there be?

The answer depends on consequence and reversibility, not on how impressive the model is. Money above an agreed threshold, customer-facing communication and changes to a control flagged by audit should stay with a person. The Deloitte finding points the same way: executives expect more value from people and agents working together than from replacing people outright. Put the person where accountability sits, and let the agent take the rest.

What does SabanaTech not claim?

Credibility in governance depends on being exact about what is certified. SabanaTech does not hold ISO 27001, SOC 2 or ISO 42001 certification, and we do not claim alignment with them. We also have no published outcome figures for governance as such. The evidence we can show is the AP agent case: a 500-invoice golden set and prompt and model versions pinned quarterly.

Where should you start?

Pick one finance process that already has an agent running or planned. Assign it an autonomy level in writing, name its owner and check it against the six controls. The gaps you find are your governance backlog. Our Governance & AgentOps offering covers this work, and our article on RPA vs AI agents helps you decide which steps need an agent at all.

All insights
  • Playbook

    An automation governance model that doesn't slow you down

    The wrong governance model puts every automation through a slow, multi-stage intake process. The right one writes down four rules, automates compliance with them, and gets out of the way. Here's the model we ship with new CoEs.

    Read the article
  • RPA

    RPA vs AI agents: how to decide which one a step needs

    Use an agent where a step needs judgment, RPA or an API where it needs determinism, and a person where it needs accountability. A decision guide for finance and operations, with an accounts payable example.

    Read the article

Bring one process. We will show you what changes.

Share the workflow, its monthly volume, the systems involved and where exceptions pile up. A delivery lead will come back with an assessment of fit and a practical next step.