I have mounted /var/run/docker.sock into a container more times than I’d like to admit. Usually it was a CI runner or a dev container that needed to build images, and it always felt harmless because the container was “mine”. Then I read the keynote recap in Docker’s post about Cloud Sandboxes and the Sandbox Kit spec, and one example made me rethink where I let coding agents run.
The example: an agent in a container with the host Docker socket mounted found a way to read a secret on the host. Docker’s own wording is that it hadn’t discovered a new vulnerability. It was using access the configuration had given it. That sentence is the whole post for me, so I’ll spend the next 1,500 words or so on why, what I’d change in a typical setup, and where a sandbox does and doesn’t help.
I haven’t run Docker’s Cloud Sandboxes myself. Everything about them below comes from Docker’s announcement, and I’ll mark my own opinions as opinions.
The socket was never a small permission
Mounting the Docker socket gives the process inside the container control of the Docker daemon on the host. The daemon runs as root. Anything that can talk to it can start a new container, and a new container can mount any host path, including /. Docker’s security documentation says the daemon needs root and that only trusted users should be allowed to control it. A mounted socket quietly makes the container one of those users.
Here is the shape I’ve seen in plenty of compose files, mine included:
services:
agent:
image: my-coding-agent:latest
volumes:
- ./repo:/workspace
- /var/run/docker.sock:/var/run/docker.sock
environment:
- ANTHROPIC_API_KEY
- GITHUB_TOKEN
For a human running a known script, this is a convenience with a risk you’ve accepted. For an agent, it changes the math. The agent decides what to run, and it can be steered by anything it reads: a README, an issue comment, a web page, a dependency’s install script. I wrote about the identity side of that in why your bot will verify anyone who asks nicely. The socket is the infrastructure side of the same problem. If the agent gets talked into something, the socket is how “talked into something” becomes “root on the box”.
Notice the two environment variables too. The socket lets the agent reach the host, and the tokens let it reach your accounts. Either one alone is a bad day. Together they’re the whole attack in one file.
What Docker’s sandbox changes
Per Docker’s post, each Cloud Sandbox is an isolated microVM, which they describe as a small virtual machine with its own kernel and its own Docker daemon. The agent can install dependencies, build applications and run containers inside that environment. The local version of the same idea uses the same sbx command-line tool, and Docker says you can start locally, move the sandbox’s filesystem to the cloud for a longer run, and bring it back to review.
The difference from my compose file is where the boundary sits. With a socket mount, the agent shares the host’s daemon. In a microVM, the agent has its own. In the keynote demo Docker describes, the same attempt to reach the host through the Docker socket failed inside the sandbox. That is the property I care about, because it doesn’t depend on the agent behaving. Docker’s phrase is that “the infrastructure enforces the boundary, even when an agent chooses an unexpected course of action.”
I’d add my own caution. A vendor announcement is an announcement. I’d want to see the escape attempts that failed on someone else’s machine before I’d call it solved, and a VM boundary is a stronger bet than a shared kernel, not a guarantee. Still, “own kernel, own daemon” is a better starting position than “root-equivalent socket and a GitHub token”.
The part where Docker admits the limits
What I liked most about the recap is that it doesn’t oversell. It quotes the keynote saying that blocking an unapproved network destination is different from recognizing that an otherwise permitted email is going to the wrong customer. And it says there is still work to do on whether an action matches the user’s intent.
That matches what I’ve seen. A sandbox answers “where can the agent act?” It does not answer “should it have done that?” If you give an agent a token that can delete repositories, a microVM will faithfully let it delete repositories, because that’s inside the permissions you granted. The recap’s example of the right instinct is narrow permissions, such as letting an agent read and draft messages without letting it send them.
So the sandbox is the floor, and the permissions inside it are still your job. Two layers, and neither replaces the other.
What I’d change in my setup today
Even without adopting anything new, there’s a short list I’d apply to any container that runs an agent. First, delete the socket mount. If the agent needs to build images, give it something else: a remote builder, rootless Docker inside the container, or a separate daemon on a throwaway VM. Second, stop passing long-lived tokens through the environment. A scoped, short-lived token for one repository beats a personal access token every time.
Here’s the same service with the cheap hardening applied:
services:
agent:
image: my-coding-agent:latest
read_only: true
tmpfs:
- /tmp
volumes:
- ./repo:/workspace
cap_drop:
- ALL
security_opt:
- no-new-privileges:true
networks:
- agent-egress
environment:
- REPO_TOKEN_FILE=/run/secrets/repo_token
secrets:
- repo_token
None of that is exotic. A read-only root filesystem, no extra capabilities, no privilege escalation through setuid binaries, a scoped secret mounted as a file. It will stop a lot of accidents and a fair number of lazy attacks. It will not stop a determined one, because containers share the host kernel. That is exactly the gap a microVM is meant to close, and I’d treat the compose version as a stopgap rather than a destination.
If you’re already comfortable with compose healthchecks and dependency ordering, as I covered in my post on depends_on that only waited for start, the egress network is the next piece to get right. Define it as an internal network with one proxy container that allowlists the hosts the agent needs. Anything that isn’t on the list simply fails to resolve. I’d start with the package registry, your git host and the model API, and add hosts only when a task fails and I understand why.
Kits and why reviewable permissions matter
The second announcement in the recap is the Sandbox Kit specification. A Kit is an OCI image containing the agent and its tools, plus declarations of the access it needs. Docker says teams can build, push, pull, scan and pin it by digest, so changes to the requested permissions show up in review next to everything else. If a Kit asks for another network destination or credential, a reviewer sees that change.
I think this is the more interesting idea, and it applies even if you never touch Docker’s runtime. The habit it encodes is that an agent’s environment should be a reviewed artifact, not a pile of flags someone typed once. My compose file above is a weak version of that. It lives in git, so a change that adds a mount or an environment variable appears in a diff. That’s already better than most agent setups I’ve seen, which are a shell alias.
Docker published the spec under Apache 2.0 and said it will bring it to the CNCF for neutral governance. Docker Sandboxes is described as the first runtime to implement it. Whether other runtimes adopt it is the open question, and I’d wait and see before building tooling around it. Specs from one vendor have a way of staying one vendor’s spec.
What I would not do
I wouldn’t move everything to cloud sandboxes just because longer tasks can keep running after you close your laptop. That’s a real convenience, and it’s also a reason to be careful: work you can’t see is work you can’t interrupt. Before letting a long task run unattended, I’d want its credentials scoped to exactly that task, and a log I can read afterwards.
I also wouldn’t treat a sandbox as permission to grant broader tokens. The temptation is real. “It’s isolated, so I’ll give it my admin key” is the sentence that turns a good boundary into a false sense of safety.
Do this before the weekend
Run grep -r "docker.sock" . across your compose files, devcontainer configs and CI workflows. For every hit, write down what runs in that container and whether an LLM-driven tool can ever send it commands. For the ones where the answer is yes, remove the mount or move that workload somewhere disposable. Then look at which tokens those containers carry, and replace the broadest one with a scoped one.
That’s maybe an hour of work. If you want to see how I set up client infrastructure with this sort of boundary in mind, my projects are collected at abrarqasim.com.