Skip to content

The AI-control plane ​

hopbox ships an MCP control plane: the same box fleet a human reaches over SSH, an AI drives over the Model Context Protocol. It is event-driven (subscribe and react to pushed changes; never poll) and shares the daemon's engine — so an AI spawns real boxes, not a simulation.

Enable it on the daemon with --mcp-addr (the installer defaults it on, at unix:/run/hopboxd-mcp.sock).

Resources (subscribe, then react to notifications/resources/updated) ​

URIWhat
hopbox://fleetEvery box with its live phase, the agent's self-reported status, and the workspace it is in (omitted for your default one).
hopbox://box/<id>/logA delegated run's transcript — every assistant turn and tool call. ?from=<seq> reads on from a cursor. Survives the box.
hopbox://surfaces · hopbox://surface/<name>/eventsYour active rendered canvases, and one canvas's interaction events (the canvas loop).
hopbox://asksPending ask questions your boxes' agents are blocked on.
hopbox://session/<id>/eventsA conversational agent session's event stream; ?from=<seq> replays from a cursor.
hopbox://profilesYour box profiles — what each resolves to, and how an async build reported back.
hopbox://imagesThe box image catalog with a per-image skill summary — pick one for box_delegate / fleet_apply.
hopbox://guideThis plane's own how-to, also delivered as the MCP instructions.

A fleet entry looks like this — workspace is absent for your default workspace, present otherwise:

json
[
  {"id":"a1b2c3d4e5f6","name":"box1","image":"python","state":"running","updated":1755400000},
  {"id":"9f8e7d6c5b4a","name":"box1","workspace":"ws1","image":"go-dev","state":"running","updated":1755400060}
]

Two rows, same name — box names are unique only within a workspace. Group the fleet by workspace to render a project list, and there is no second call to make: membership rides the resource you are already subscribed to.

Tools ​

A box ref below is a box reference: an id from hopbox://fleet (or any unambiguous prefix of one), a bare name in your default workspace, or ws1/box1 — which is exactly workspace + name from the fleet. It is the same grammar ssh cli@host and the HTTP API take, so a box is addressed identically from every door.

Fan work out

ToolWhat
box_delegate {task, agent?, keep?, ttl?}Spawn a box, run a task in the background; the result surfaces on hopbox://fleet, and what it did on hopbox://box/<id>/log. agent:true runs a Claude Code agent (on your token). Fire-and-forget by default — the box is archived (disk freed, record kept) when done; keep:true leaves it persistent + resumable. ttl:"30m" sets a hard deadline (see below).
fleet_apply {boxes:[{key,image,task,ttl?}]}Declare a desired set of task-boxes; hopbox converges to it, idempotent per key. Prefer over many box_delegate calls.
box_spawn {name?, ttl?} / fleet_getSpawn an empty box / snapshot the fleet.

Bound work you aren't watching

ttl is a deadline from creation, enforced by the daemon. The box is torn down when it expires whatever it is doing — it beats an attached session, a keep-alive pin and the durable flag, so an agent that loops or a build that wedges stops costing you at a time you chose. It also survives the orchestrator that started the box crashing, which a timeout held in your own process cannot.

Set it on any delegation you won't be watching. Omit it and your tier decides: a fire-and-forget delegation gets its tier's default, and anything you ask for is clamped to the tier ceiling.

Drive a box (everything a human does over SSH, over MCP)

ToolWhat
box_exec {box, cmd, timeout_s?}Run a command and get the output back in the call — stdout/stderr/exit, synchronously. Auto-resumes a suspended box. The workhorse.
box_write {box, path, content} · box_read {box, path}Stage a file into a box / read one out (bytes-safe).
box_suspend · box_resume · box_rm {box}Lifecycle. An AI cleans up and checkpoints its own fleet.
session_start {box, prompt?, model?, effort?} / session_send / session_list / session_stopA persistent, conversational agent session in a box; events stream on hopbox://session/<id>/events?from=<seq> (cursor-replayable — reattach from anywhere, and both sides of the conversation are in the buffer). model / effort are fixed for the session — see model and effort.

Build an environment once — box profiles, so a delegated box already has the toolchain

ToolWhat
profile_set {name, definition}Store a profile from its JSON — a catalog base plus a build script, plus the defaults a box from it starts with. Validated now, not at build time. Storing does not build it.
profile_build {name}Compile it into an image. Asynchronous: returns the build number immediately and takes minutes.
profile_list + hopbox://profilesWhich build each profile resolves to, whether its base image has been rebuilt since (stale), and the newest attempt's status — how an asynchronous build reports back.
profile_rm {name}Delete a profile; its images are reclaimed once no box is still booting from them.
ws_profile {name, profile}Point a workspace at a profile: boxes spawned there start from it when their spec names no image. Precedence is explicit spec > workspace default > nothing. profile:"" clears it.

Spawn from one by putting its name where an image goes — box_delegate {image:"go-dev"}, fleet_apply {boxes:[{image:"go-dev"}]}, or ssh <box>:go-dev@host. That is the payoff: build the toolchain once instead of re-running the same apt-get in every box you delegate.

Do not wait on a build

profile_build is the one verb here that takes minutes. Start it, go do something else, and react to hopbox://profiles — the build box is visible on hopbox://fleet while it runs. A profile's defaults.egress is also the live allowlist for boxes spawned from it, so profile_set reaches boxes that are already running.

Share state across boxes — the shared workspace as the AI's hand-off drive

ToolWhat
wrk_write {path, content, ws?} · wrk_read {path, ws?} · wrk_list {path?, ws?}The plane's /wrk: box A writes a shard, the orchestrator reads it, box B picks it up. Pass ws to target a specific workspace's drive (omit = default). The same drive the plane's boxes in that workspace mount at /wrk.

Talk to a human

ToolWhat
surface_render {name, html}Render an interactive canvas at a URL — the canvas loop, below. Arbitrary HTML, fire-and-forget.
ask_answer {id, choice} + hopbox://asksAnswer an ask — a structured question an (often in-box) agent raised and is blocked on. hopbox://asks lists the pending ones; answering unblocks the agent. The request/response, form-shaped cousin of surface_render.

An AI can drive a box entirely without a human: box_exec to run, box_read / wrk_read to collect, surface_render to report. Everything is owner-scoped — the plane drives only its own boxes, never a human's. On connect, the server sends an instructions guide teaching all of the above.

The canvas loop ​

When an AI needs a human's decision, input, or attention, it isn't limited to chat — it can render an interactive UI and watch the human use it, live:

  1. surface_render {name:"approve", html:"<button id=ok>Approve</button>"} returns a URL. The AI gives it to the human.
  2. The AI subscribes to hopbox://surface/approve/events.
  3. Each click/input is pushed to the AI as {kind, target, value} — it reacts: re-render, branch its work, or unblock a waiting task.

Serve surfaces over HTTP with --surface-addr — on the live host they appear at box.hopbox.dev/s/…. The loop is bidirectional: the AI renders, the human acts, the AI observes.

Your own plane — ssh cli@host mcp ​

The plane is yours, by SSH key — like every other hopbox surface. Point any MCP client at hopbox using ssh as its transport; your key is your login, and the plane binds to your identity.

Claude Code — one command:

sh
claude mcp add hopbox -- ssh -T -o StrictHostKeyChecking=accept-new cli@box.hopbox.dev mcp
claude mcp list        # hopbox … ✔ Connected

Claude Desktop (and any stdio MCP client) — add to its config (claude_desktop_config.json → mcpServers):

jsonc
{
  "hopbox": {
    "command": "ssh",
    "args": ["-T", "-o", "StrictHostKeyChecking=accept-new", "cli@box.hopbox.dev", "mcp"]
  }
}

The flags matter: -T runs ssh without a pty (the plane is a byte stream), and StrictHostKeyChecking=accept-new accepts the host key on first use instead of hanging on the interactive "continue connecting?" prompt — the #1 reason a GUI MCP client silently never starts.

Prefer hopbox-mcp --ssh for anything long-running

Raw ssh as the server command works, but it never comes back from a dropped connection. Your MCP client does not restart a server that exits — it reports "Server disconnected", and the whole plane stays dead until you restart the app. One hopboxd deploy, one laptop sleep, or one wifi change is enough. An idle session is also reaped by NAT and firewalls, because ssh sends no keepalives by default: measured idle-before-death on a real setup was 2h18m, 3h05m and 11h23m.

hopbox-mcp bridges the same plane and handles both — keepalives, and reconnect with backoff:

jsonc
{
  "hopbox": {
    "command": "hopbox-mcp",
    "args": ["--ssh", "cli@box.hopbox.dev"]
  }
}

It ships in the release bundle for macOS and Linux:

sh
curl -fsSL https://hopbox.dev/dl/latest/hopbox_darwin_arm64.tar.gz | tar -xz -C /usr/local/bin hopbox-mcp

If you keep raw ssh, at least add -o ServerAliveInterval=30 -o ServerAliveCountMax=3 so an idle session isn't silently reaped. It still won't reconnect.

Subscriptions survive the reconnect too — but only through hopbox-mcp. Resource subscriptions live in the server's session, and a reconnect gets a fresh one, so the bridge remembers every resources/subscribe you send and replays it into the new connection; it then pushes you a notifications/resources/updated per restored resource, because it can't know what changed while the link was down — re-read those and you're back in sync. With raw ssh there is no reconnect at all, so there is nothing to restore: you re-subscribe by restarting the app.

"Permission denied (publickey)" — esp. Claude Desktop on macOS

Test the transport from a terminal first: ssh cli@box.hopbox.dev ls. If that lists your boxes but the app logs Permission denied (publickey), the app launched sshwithout SSH_AUTH_SOCK — a GUI app doesn't inherit your shell's environment, so ssh can't reach the agent holding your key (and if your key is agent-only, there's no file to fall back to). Recover the agent with a one-line wrapper and point the server's command at it:

sh
# ~/.hopbox/mcp-ssh   (then: chmod +x ~/.hopbox/mcp-ssh)
#!/bin/sh
# Recover the agent socket the GUI app didn't pass down, then hand off to hopbox-mcp,
# which owns the keepalives and the reconnect loop. Keep this wrapper doing ONE thing.
[ -n "$SSH_AUTH_SOCK" ] || export SSH_AUTH_SOCK="$(launchctl getenv SSH_AUTH_SOCK)"
exec hopbox-mcp "$@"
jsonc
{ "hopbox": {
    "command": "/Users/you/.hopbox/mcp-ssh",
    "args": ["--ssh", "cli@box.hopbox.dev"] } }

Without hopbox-mcp installed the wrapper can exec /usr/bin/ssh "$@" with the raw args instead — but see the warning above: that form dies permanently on the first dropped connection.

If your hopbox identity is instead a key file, skip the wrapper and add "-i", "/path/to/key" to the args (it must be the key that owns your boxes — a different key gets a different, empty fleet). Can't use SSH at all (a browser, a locked-down client)? Use the hbx_ WebSocket transport below.

hopbox://fleet then shows only your boxes; box_spawn/delegate create boxes you own; wrk.* is your /wrk; and rendered surfaces are yours. A different key gets a different, fully isolated plane. Because your boxes are yours, an agent box delegated from here runs on your stored token — box_delegate {task, agent:true} just works, no operator key.

From a browser or a third-party app — anything that can't present an SSH key — connect over a WebSocket with your hbx_ key instead: wss://host/v1/mcp?key=hbx_…. Same per-owner plane, same tools. Workspaces (ws.*) and secrets (secret.*, values never returned) are on the plane too, so a UI can manage them.

That key rides in the URL, because a browser cannot set headers on a WebSocket handshake — so give the browser a scoped key rather than your account's: apikeys create browser -w ws1 --ttl 24h. The plane behind it is that one workspace — its fleet, its boxes, its /wrk, its secrets — and the account-wide tools (ws_create/ws_rm, ws_profile, profile_*, session_*, ask_answer) are refused on it.

The client — hopbox-mcp ​

sh
hopbox-mcp --ssh cli@box.hopbox.dev            # your plane over SSH (wraps the above)
hopbox-mcp ps   --connect unix:/run/hopboxd-mcp.sock    # the operator/local plane (root socket)
hopbox-mcp watch --connect <sock> hopbox://surface/<name>/events   # print interactions live
hopbox-mcp --demo                              # self-drive a demo against box.hopbox.dev

The --connect socket is the operator/local plane (its own owner); the ssh cli@host mcp path is the per-customer one.

What a plane offers is what it can do ​

tools/list is not a fixed list. The daemon's plane — ssh cli@host mcp or wss /v1/mcp — serves every tool on this page. A plane backed by something smaller advertises only what it can actually run, and its instructions name what is absent.

hopbox-mcp --demo is the one that differs: it drives real boxes over ssh, with no daemon behind it, so it offers box_delegate, box_spawn, fleet_apply, fleet_get and the surface tools — and does not offer box_exec / box_read / box_write / box_rm / box_suspend / box_resume, wrk_*, ws_*, secret_*, ask_answer or session_*. Calling one anyway is refused with "… is not available on this backend".

An AI should therefore trust tools/list over this page: it describes the plane in front of it, not the protocol.

Humans get the same engine ​

Everything here has a human twin over SSH — ssh cli@host lists and removes boxes, key-authed, zero-install. One fleet, one engine, two front doors.

Instant isolated compute — for humans and AIs