Claude plugins bring their own runtime requirements. claude-mem and others run
their hooks under bun, which the sandbox image does not carry, so installing
plugins left every prompt in the sandbox failing with a Bun not found hook
error - after setup had reported success, since the dependency only surfaces
when a hook fires.
AI_SBX_TOOLS lists mise tools to install globally in the sandbox and defaults
to bun. mise is already present and resolves these from GitHub releases, which
the default network policy allows. Setting the variable to an empty string
installs nothing.
sbx skills import prompts before overwriting each skill already in the shared
store. setup called it with stdout and stderr redirected and stdin left
attached, so on any second run the prompt was invisible and setup hung
indefinitely partway through installing the Claude configuration.
sbx rm prompts the same way, which would have blocked --replace and remove
once a sandbox was in use.
Both now pass --force and read from /dev/null. This is the third instance of
the pattern after sbx secret set, which cancelled silently and still exited 0,
so a test now asserts at the source level that every prompting subcommand is
invoked non-interactively. The test was confirmed to fail when either --force
is removed.
Fine-grained tokens have no Checks permission. GitHub's permission reference
lists no Checks section and no check-run endpoint, and checks is absent from
the token form's pre-fill parameters, so the earlier instruction to tick
Checks: Read asked for a box that does not exist. A token created with every
listed permission still could not read check runs, which is what surfaced this.
The prompt now states the consequence rather than offering a remedy: gh pr
checks reports commit statuses only and gh run view returns no annotations.
Both degrade to empty output rather than a permission error, so without the
note they read as a broken CI integration. Job logs are unaffected; they fall
under Actions, which is granted.
secret_scanning_alerts and vulnerability_alerts move into the pre-filled URL,
leaving nothing for the operator to tick beyond repository selection.
Two of the three permissions treated as manual are in fact pre-fillable. The
earlier check scraped the rendered docs page, whose table splits those rows in
a way the parse missed; the docs source lists secret_scanning_alerts and
vulnerability_alerts as supported query parameters. Only checks is genuinely
absent, so the manual list shrinks to that one entry.
That entry now names what breaks without it. Checks: Read governs the status
rollup behind gh pr checks and the annotations behind gh run view, and both
degrade to empty results rather than permission errors, so an unticked box
reads as a broken CI integration rather than a missing scope.
The marketplace and enabled-plugin extraction move out of the install function
into host_marketplaces and host_enabled_plugins so they can be exercised
directly. The tests cover the pipe separator that keeps an absent repo from
shifting a url leftwards, rejection of marketplace sources that are neither
github nor git, disabled plugins being excluded, and the allowlist refusing to
carry credentials, transcripts or history.
The Claude configuration, plugin and skills support added for sandboxes only
applies when the agent is claude, so codex was no longer the sensible default.
Custom templates are the documented way to carry user-level configuration into a
sandbox, but sbx v0.37.x silently drops every layer stacked on the base image
(docker/sbx-releases#366), so nothing baked into an image arrives. This installs
the same material into a stock sandbox after creation instead.
setup copies an allowlist of ~/.claude into the sandbox, imports skills into the
shared store, and adds each known marketplace before installing every enabled
plugin. Enablement survives here precisely because it happens after creation:
claude plugin install writes enabledPlugins itself, whereas sbx recreates
settings.json when the sandbox is created.
Tools come from mise, copied from the host because mise.jdx.dev is outside the
default network policy. Tools resolve through shims rather than mise activate,
which only fires for interactive shells and would leave the agent silently using
system versions.
sbx exec drains stdin, which truncated both install loops to their first entry,
and tab is an IFS whitespace character, which collapsed the empty repo field and
shifted the URL into it for git-sourced marketplaces. Both are handled.
Sandboxes ignore the host ~/.claude by design: the agent runs as a separate
user with HOME elsewhere, so even a read-only mount is not picked up. Skills
can be shared with sbx skills import, but plugins carry commands, hooks and
MCP servers that only a custom image can deliver.
Adds --template, --stock-template and a repeatable --kit, persisted per
repository so run and refresh reuse them. AI_SBX_TEMPLATE supplies the default
image, so one custom template can be declared once in the user's mise config
and apply to every repository, with --template overriding it per repository
and --stock-template opting out.
save_config now packs two arrays into one argument list separated by a count,
so it ships with a round-trip test covering empty arrays, values containing
spaces, and the boundary between kits and AWS profiles.
sbx 0.29 isolated the agent with --branch, creating a host-side Git worktree.
0.37 removed that flag and reinstated --clone, which gives the agent a private
in-container clone mounted read-only and exposes its commits through a
sandbox-<name> git remote on the host. Setup fails outright against 0.37 with
'--branch is no longer supported'.
Drops the branch name plumbing entirely, since the sandbox now owns the clone
and there is no host branch to name.
Documents the two host prerequisites this surfaced: membership of the kvm
group, because sandboxes are microVMs, and the docker-sbx package rather than
docker-sandbox-bin on Arch derivatives - the latter installs only the CLI,
omitting the microVM kernel, rootfs and nerdbox shim, which makes sbx fall back
to mounting filesystems on the host and fail for any non-root user.
The App approach does not survive contact with a hundred developers and
hundreds of repositories. Minting installation tokens requires the App private
key on every developer's machine, and a key that widely distributed is a key
that grants org-wide minting to everyone holding it.
Device flow looked like the way out, since it needs no private key, but
testing showed it does not scope. A token requested with repository_id for one
repository reached a second repository in the same installation: a
permission-gated endpoint returned 200 where an installation token scoped to
one repository returned 403 for the same public repository. GitHub accepts
repository_id and silently ignores it. Per-repo scoping therefore requires
either the private key or the client secret, and neither can live on a
developer's machine.
Fine-grained PATs do scope per repository and share no secret, and GitHub
supports pre-filling the creation form via URL parameters, which removes the
toil that made them unattractive. Setup now builds that URL from the origin
remote and opens it, leaving the operator to select the repository and paste
the result.
Three permissions - checks, vulnerability_alerts and secret_scanning_alerts -
are absent from GitHub's pre-fill parameters, so they are printed as a
checklist instead of sent as parameters that would be silently dropped and
look granted. There is no parameter for repository selection either.
Tokens are no longer re-minted per launch, since a PAT outlives a session; the
new token subcommand replaces one on expiry or revocation.
A task included from the global mise config runs with the config root as its
working directory - $HOME - rather than the directory the user invoked it from.
Deriving the repository from the current directory therefore failed everywhere
except a project-level include, which defeats the point of installing the task
once and using it in every repository.
mise passes the real directory as MISE_ORIGINAL_CWD, so enter it before
resolving the repository, falling back to the current directory when the task
is run directly rather than through mise.
GitHub exposes no API to create a fine-grained PAT and no way to prefill the
creation form, so every repository meant hand-clicking a permission set and
remembering to rotate it. Installation tokens are API-mintable, so configuring
one GitHub App removes the per-repository work entirely.
A new 'app' subcommand records the App ID and private key path once. Setup then
resolves the installation for the repository, and run and refresh mint a fresh
token scoped to that single repository before every launch. Tokens expire in an
hour on their own, which retires manual rotation.
sbx secret set is invoked with --force because without it a second write prompts
for confirmation, reads the prompt from the stdin already consumed by the token,
cancels, and still exits 0 - leaving the previous, expired token in place.
The permission set is validated against GitHub's app-permissions schema. Notably
workflows has no read level, and write is required to push any commit touching
.github/workflows, which is a separate permission from actions.
Also corrects several sbx invocations that did not match the installed CLI:
--no-share-skills and --clone are not create flags, isolation is --branch; run
takes a sandbox name rather than --name; exec takes no -- separator; ls --quiet
replaces parsing tabular output; and the sandbox home is queried rather than
assumed to be /home/agent.
Adds a JWT test that verifies signatures against a generated public key and
confirms tampered input fails to verify.
Provides a shareable mise task, ai:sbx, that runs an AI coding agent in a
Docker Sandbox scoped to a single GitHub repository and a set of read-only
AWS roles.
The repository is derived from origin rather than configured, so the sandbox
identity cannot drift from the checkout in use. GitHub access is a
repository-scoped fine-grained PAT held in the sbx secret store and injected
by its host-side proxy, so the token is never exposed to the agent. The host
~/.aws directory and SSO token cache are never mounted; instead the host
exports short-lived credentials for approved read-only profiles and only
those land in the sandbox.
Host profiles are commonly suffixed to mark the grant (api-portal-readonly)
while Terraform references the account name (api-portal), so a trailing
-readonly is stripped when the profile is written into the sandbox. Two host
profiles that collapse to the same sandbox name are rejected during setup,
before any credentials are exported, since a silent overwrite would hand
Terraform the wrong identity under a plausible-looking name.
All state lives under ~/.config/ai-sbx; repositories supply nothing and need
no mise.toml.