Service 04

AI agent setup & security

Agents that write code, work through case files or run operations tasks, set up to do real work and contained from day one.

From £18,000 4-8 weeks Fixed fee

Agents that act need a boundary

An AI agent does more than answer. It runs code, writes files, calls tools and reaches across your network, often for hours without supervision. That is where the value is, and it is why an agent cannot be trusted to police itself.

This summer, agents in several frontier-lab evaluations escaped the environments meant to contain them. In August the NCSC published interim guidance on agentic AI: sandbox the environment, keep activity observable and attributable, and keep an emergency stop. We set agents up so they meet that guidance before they touch real systems or data.

What we set up

The agents themselves

Coding agents for your engineers, document and case-file agents for your analysts, and operations agents for IT service management and infrastructure work. Each is routed to models you approve, including open-weight models on your own GPUs.

A sandbox around every agent

Deny-by-default network egress, filesystem limits, system call filtering and no stored credentials inside the sandbox. We use NVIDIA OpenShell where it fits, with its policy compiled from code and checked before every change.

Least privilege on every tool

Scoped, short-lived credentials for each connected system. New network destinations and consequential actions wait for a named person to approve them.

Logs you can attribute

Every action, policy decision and model call goes to your SIEM, and outbound traffic is identifiable as yours. When someone asks what the agent did last Tuesday, you can show them.

A kill switch someone has practised

One control that halts an agent, cuts its network and interrupts its access to the model. We rehearse it with your team before go-live, because a stop nobody has tested is not a control.

Security review of agents you already run

Many organisations already have agents in use: coding assistants with repository access, copilots wired into mail and documents, automations with service accounts. We inventory them, map what each can reach, test whether the limits hold and write down what to change first.

On DGX, HGX or your cloud

Agents run on the infrastructure you already have: NVIDIA DGX and HGX estates, ordinary Linux servers or a UK cloud region. On DGX and HGX systems with BlueField DPUs we also enforce egress outside the host and prepare for NVIDIA Sentry. Our guide to OpenShell and Sentry on DGX and HGX covers what runs on BlueField-3 today.

Mapped to NCSC guidance

Each control maps to the NCSC's considerations for agentic AI and to the Cyber Assessment Framework: permissions and sandbox configuration to B2 and B4, activity logs to C1, and the stop procedure to D1. The standards we build to apply here as everywhere else.

How an engagement runs

Weeks 1-2: Inventory and threat model. Agents, harnesses, hosts and credentials listed. Red lines agreed. One agent class chosen to start.

Weeks 2-6: Set up and contain. Agents configured, sandboxes and policy in code, credentials scoped, logs flowing to your monitoring.

Weeks 4-8: Prove it and hand over. We attempt escapes on purpose, rehearse the kill switch and hand over the evidence, runbook and policy repository.

FAQ

Questions we get asked

Which agents can you set up?

Coding agents such as Claude Code, Codex, GitHub Copilot CLI and OpenCode, and agents for documents, case files, IT service management and operations. They run against the models you approve, including open-weight models on your own hardware. We start with one agent class, usually coding, because the blast radius is clear and the results are easy to review.

Do you only work with NVIDIA OpenShell?

No. OpenShell is a strong open-source default and runs on DGX, HGX and ordinary Linux hosts, so we use it often. Where it does not fit, the same controls go in with container sandboxes, network policy and identity controls you already run. The controls matter more than the product.

Can you secure agents we already use, such as Copilot or hosted assistants?

We can review them and reduce what they can reach: identity and permission scoping, data connectors, network egress, logging and an off switch. A hosted agent cannot be put in your sandbox, so the review says plainly what you control and what you are trusting the vendor for.

Does this stop prompt injection?

No product does. The NCSC warns it may never be fully mitigated. Containment limits the damage: an agent that reads a malicious document still cannot reach hosts, files or credentials outside its policy, and anything consequential waits for a person.

Can agents run fully air-gapped?

Yes. Coding and document agents work well against local open-weight models with no route out. We mirror every dependency inside the boundary, as on our air-gapped platforms.

How is this different from AI implementation?

Implementation builds a system around your data, such as retrieval or workflow automation. This service sets up agents that act, and puts the boundary around them. Many clients do both. If you are not ready to build, start with an agent security review under our readiness assessment, from £4,500.

Next step

Start with a 30-minute call

No deck. Tell us the problem. We will say whether AI is the right tool, what it would take and roughly what it would cost.