Multi-Cloud AI

Multi-cloud agents without multi-cloud chaos: one control layer

When a multi-cloud agent programme goes wrong, it is rarely because a cloud could not run the workload. It is because nobody can answer a simple question, such as what a particular agent did last Tuesday, without opening four consoles.

Where it breaks

The pattern is familiar. Each team picks the cloud that suits its agent and sets up identity, logging and guardrails the way that cloud encourages. Six months later there are several logs in different formats, policies that disagree, evaluations nobody compares, and no single owner. Each agent works. The estate is ungovernable.

Five things to define once

Identity

Every agent runs as an identity your organisation manages, with permissions granted per task, and every human approval is tied to a named person. The same rules apply whichever cloud the agent runs on.

Policy

What each agent may do on its own, what needs a person, and what it may never do. Written once, in one place, and enforced in each environment rather than reinvented there.

Logging

One schema for agent actions, sent to one place: what the agent read, decided and changed, which rule allowed it, and who approved it.

Evaluation

One set of test cases per agent, run on every meaningful change, with thresholds agreed in advance, so results are comparable across clouds and over time.

Escalation

One route for uncertain or out-of-bounds cases to reach a person who is actually watching, with response times that someone owns.

How to draw an agent’s guardrails

Keep control and compute apart

Treat these five as a control plane that sits above the clouds, and the clouds as the compute plane where agents run. Policies and identities are defined in the control plane and pushed down; logs and evaluation results flow back up.

The benefit is that adding a cloud becomes a compute decision. The new environment receives the existing policies and starts sending the same logs. Nothing about governance is rebuilt.

The test of a good control layer

Adding a second cloud should change where agents run, and nothing about how they are governed.

What to measure across clouds

Measure the same things everywhere so a regression shows up the day it happens: escalation rate per agent, evaluation pass rate per change, the share of actions approved without edits, and the time from escalation to a human decision. A rising escalation rate on one cloud is often the first sign that something upstream has changed.

A readiness checklist before the second cloud

Before adding another environment, you should be able to say yes to each of these.

Every agent identity and its permissions are listed in one place.

There is one written policy per agent, owned by a named person.

Agent actions from every environment land in one log with one schema.

Each agent has an evaluation set that runs on every change.

Escalations reach a monitored queue with an owner.

If any of these are missing, fix them on one cloud first. They only get harder to retrofit once there are two.

Claude on Bedrock, Vertex AI or Foundry: where should each agent run?

Where Focus20 fits

We design the control layer before the second cloud arrives, so governance is set once rather than rebuilt per environment. It is the same discipline we bring to our multi-cloud architecture support for the Maharashtra Pollution Control Board.

Our AI governance model

Questions we get

Do we need a vendor’s control plane?

Not necessarily. What matters is that the five things are defined once and applied everywhere. That can be a product, your own tooling, or a mix of both.

Who should own agent policy?

Whoever answers for the outcome when an agent gets it wrong, usually a business owner rather than the platform team, with engineering responsible for enforcing it.

How do we log across providers?

Agree one schema, send every environment’s agent events to one store you control, and tie each event to the request and the person behind it.

Why agentic AI needs more than one cloud