Skip to content

Run boxes on Kubernetes as an alternative runtime - #94

Closed
kossmac wants to merge 6 commits into
splitbrain:mainfrom
kossmac:kubernetes-runtime
Closed

kossmac wants to merge 6 commits into
splitbrain:mainfrom
kossmac:kubernetes-runtime

Conversation

@kossmac

@kossmac kossmac commented Oct 4, 2026

Copy link
Copy Markdown

Adds Kubernetes as a second, experimental runtime for boxes, next to Docker. Docker stays the default; a deployment switches with RUNTIME=kubernetes.

What changed

  • A runtime seam. Everything the orchestrator did through docker.ts now goes through runtime() (orchestrator/src/runtime/), with a Docker and a Kubernetes implementation. File access for review, uploads and the workspace browser goes through fileAccess() the same way (orchestrator/src/fileaccess/). With RUNTIME=docker the behaviour is meant to be unchanged.
  • The Kubernetes runtime (orchestrator/src/kubernetes.ts, @kubernetes/client-node): a box is a pod with one PVC (workspace, home and Nix store as subPaths) and a NetworkPolicy that only lets it reach the egress proxy. File access in a pod runs a small script over exec (pod-fs-script.ts). Orphaned pods, PVCs and policies are swept like Docker's containers, networks and volumes.
  • Logins from the settings page run in a login pod on Kubernetes.
  • The box image installs tini. A box pod runs its entrypoint under it, as Docker's Init: true does, so orphaned processes are reaped and SIGTERM reaches the box.
  • Manifests for the orchestrator and the egress proxy in k8s/, and tests/smoke-test-k8s.sh, which also checks that the cluster's CNI actually enforces NetworkPolicy.
  • README: a "Kubernetes (experimental)" section with setup, settings and known limitations.

Rebase onto current main

The work was written against ee65e63. Rebasing it onto main also routes processReadings() for the new tunnel reconciler through the runtime instead of dk.

Testing

🤖 Generated with Claude Code

kossmac and others added 6 commits October 4, 2026 19:32
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Run against kind with Calico, the box pod, the review, the terminal and
the egress check each failed in a way the faked API could not show:

- The home seed ran as root without capabilities, so cp could not read
  the agent's 0700 home in the image and every pod stayed in Init. The
  setup init container now keeps CHOWN, DAC_OVERRIDE and FOWNER, as
  docker.ts's root helper does, seeds a home once rather than at every
  start, and gives the workspace and Nix store roots to the agent.
- The orchestrator's Role lacked get on pods/exec, which the client
  library's WebSocket exec needs, so review and threads got a 403.
- A terminal fed the pty's output back to the shell as input: one
  stream was both stdin and stdout. It now has two, and resizes the pty
  through the library's terminal-size protocol.
- A stop returned while the pod was still terminating, so a start right
  after it found the old pod and was left with none. A stop now waits
  for the pod to go, with Docker's 10 second grace.
- ensureProxyAttached reported an existing NetworkPolicy as a missing
  attachment, which logged every running box every minute.
- A path with .. in it answered 500 rather than 404.

The smoke test runs under macOS bash and checks the Nix store too.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A Kubernetes runtime knew no image at all, so a box stayed on the image it
was created with however often BOX_IMAGE moved. Every start makes a new pod
anyway, so the pod is now made from BOX_IMAGE and the row records it, and
the runtime knows an image by its reference: a box whose running pod is on
another reference moves at its next start, as a Docker box rolls onto a new
image id. That takes a tag per build; a moving latest is still not seen.

A pod gone after a stop is the normal case here, so it is logged as a
start rather than warned about as a lost container.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A box had a claim each for its workspace, home and Nix store, so every
running box attached three block volumes to its node. Hetzner attaches at
most 16 to a server, which stopped a node at five running boxes. The
three are now subPaths of one claim, boxes-data-<id>, named once among the
pod's volumes, and K8S_VOLUME_SIZE replaces the three sizes.

ensureAddedVolumes no longer makes a claim: every Kubernetes box has had
its one claim since creation, and an empty stand-in for a lost one would
only hide the loss. Boxes made with three claims are not carried over;
none exist yet.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019GcXSze2tsHcn3F59k9VeY
A login always ran in a Docker container, so on Kubernetes the settings
page failed with connect ENOENT /var/run/docker.sock. It now runs in a
throwaway pod from the box image, hardened like a box pod, with its home
and /tmp in memory and no service account token. The pod carries only the
login label, so no box NetworkPolicy selects it and it reaches the
internet directly, as the Docker login container does on the bridge.

A pod is Pending until scheduled and pulled, so the login waits for it to
run, and gives up with the reason when its image cannot be pulled. The
orphan sweep lists login pods too, so a restart mid-login no longer leaves
one behind.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A box pod ran the entrypoint as PID 1, so its `sleep infinity` became the
parent of every process orphaned in the box and never waited on one. Each
command whose shell went first stayed a zombie, held a pid against the
limit, and read as work that kept the box awake. The kernel also discarded
SIGTERM for it, so every stop waited out the grace period.

The image now installs tini, and the box pod runs the entrypoint under it,
as Docker's `Init: true` does. The Docker runtime and the login pod are
unchanged. A box pod needs an image built from this commit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@splitbrain-bot

Copy link
Copy Markdown

Summary

The missing features block a merge. The review with git status and diff, attachments and agent configuration are what make Boxes more than a container with an agent in it. A Kubernetes runtime without them is not a usable alternative.

The gaps are not separate bugs. They come from the runtime layer, which does not cover the thing these features depend on. Before more code is written, there needs to be a clear concept of what this layer has to be able to do, and how each runtime does it.

Why the features are missing

Large parts of the orchestrator assume that a box's files are a directory on the orchestrator's own filesystem. They also assume this directory can be read and written whether or not the box runs. Docker makes this true through bind mounts from DATA_DIR. Code that relies on it:

  • Attachment uploads write into workspacePathOf(id) directly (app.ts). The route's own comment says it needs no container for this reason.
  • Repository discovery for the review walks the workspace with readdirSync (review/repos.ts). For a Kubernetes box it gets an empty path and finds nothing.
  • Agent configuration is materialized into a host directory and bind-mounted read-only (agents.ts). Kubernetes gets an empty emptyDir instead.
  • Box sizes are measured on those directories, while the box is stopped as well.

The PR adds fileAccess() for some of these call sites and leaves the others on host paths. It also changes what a stopped box is: on Kubernetes, opening the review now starts the box (ensureFilesReachable). The Runtime interface itself mirrors docker.ts, with networks, attaching the proxy, seeding a home and copying volumes. On Kubernetes, several of these methods are stubs or throw an error. boxes.ts, app.ts and egress.ts also check cfg.RUNTIME === 'kubernetes' directly in several places.

Logins are outside the Runtime interface entirely. They have their own LoginRuntime with a Docker and a Kubernetes implementation. login.ts imports docker.ts and kubernetes.ts directly, and app.ts chooses between them with a second cfg.RUNTIME check. Starting a short-lived container from the box image and running an exec in it is a runtime operation, but the new layer does not cover it.

What a concept needs to answer

  1. How does the orchestrator reach a box's storage?
  2. What can be done with a stopped box? Today, review, uploads and size measurement work without a running container.
  3. How does read-only configuration, such as the agent configuration, get into a box?
  4. What does network isolation mean as an operation, independent of Docker networks and NetworkPolicies? This includes how the orchestrator reaches the proxy's control channel.
  5. What does the image lifecycle need: pulling, detecting a new image, pruning old ones? Which parts does a runtime have to offer, and which can it leave out?
  6. Which operations are box operations, and which are details of one runtime?
  7. What happens when a long-lived exec connection ends while its process still runs? The terminal and the adapter each hold one exec open for as long as they are used. With Docker, these connections have no time limit. On Kubernetes, they run through the API server and the kubelet, which closes streaming connections after 4 hours by default (streamingConnectionIdleTimeout). An open terminal also keeps the idle reaper from stopping the box (reaper.ts), so a terminal tab left open for hours is a realistic case.
  8. Where an operation can work the same way on both runtimes, does it? Today the PR adds a second way of doing things in several places. For example, Docker reads a box's process list from outside with docker top, while Kubernetes runs ps inside the box. Both runtimes could run ps inside the box. Every difference between the runtimes is a place where a feature can work on one and break on the other.
  9. How are both runtimes tested on GitHub, so that a change cannot break one of them unnoticed? Today CI runs the unit tests and tests/smoke-test.sh against a real Docker deployment for every pull request. The Kubernetes runtime is tested only with fake API clients. tests/smoke-test-k8s.sh needs a cluster and is not part of CI.

Findings independent of the concept

  • Box pods get a service account token. The login pod sets automountServiceAccountToken: false, but the box pod in createPod does not. On a cluster whose API server has a public address, the egress proxy lets a box reach it with the namespace's default token.
  • Service links leak into every box. enableServiceLinks is not set to false, so every box and login pod gets environment variables for every Service in the namespace.
  • The orchestrator API is open to the whole cluster. The API has no authentication. With Docker, compose.yaml publishes it on loopback only. k8s/orchestrator.yaml has no NetworkPolicy for it, so every pod in the cluster can call it, including login pods, whose egress is not restricted.
  • Nothing waits for a box pod to run. podState reports Pending as running, and startPod only checks that the pod exists. Since a stop deletes the pod, every start goes through scheduling, attaching the volume and the setup init container again. The adapter spawn retries 3 times over about 4 seconds. The terminal and the review do not retry. Pulling an image or attaching a cloud block volume takes longer than that. Login pods wait in podRunning, but box pods do not.
  • The QoS comment in createPod is wrong. The init container has no resources, and init containers count for the QoS class, so the pod is Burstable, not Guaranteed.
  • Every file operation starts a new node process in the pod. Opening one file in the review does this several times.
  • Many comments are out of date. egress-proxy.yaml says there is no orchestrator Deployment and no RBAC yet, but orchestrator.yaml has both. The RBAC comment mentions three PVCs, but there is one now. Several comments refer to "the plan" or "a later phase".

tini

Docker's Init: true runs docker-init, which is tini. Instead of installing tini and overriding the command of box pods (BOX_COMMAND), the image can start tini in its own entrypoint:

ENTRYPOINT ["/usr/bin/tini", "--", "/usr/local/bin/entrypoint.sh"]

Init: true can then be removed for the box and login containers. The root helper in oneShot replaces the entrypoint, so it keeps Init: true. Both runtimes then behave the same, and the override is no longer needed. gateway/background.ts already recognises tini as the box's init process. This has to be done in one step: with Init: true and tini in the entrypoint, tini is not PID 1 and does not reap orphans without -s.

Smaller PRs

Once there is a concept, smaller PRs towards it may be easier to get merged, each one useful on its own. The tini change is one of them.

Out of scope: multiple users

This is not part of this PR and does not need to be solved here. Boxes is a tool for one person: it has no accounts, and credentials, settings and the box list belong to the whole deployment. On a shared cluster, supporting several users becomes more reasonable. The user would likely be passed in as a request header by the reverse proxy that handles authentication. The runtime concept should not rule this out.

@kossmac

kossmac commented Oct 5, 2026

Copy link
Copy Markdown
Author

Thanks for the thorough review. Agreed: the runtime layer needs a concept before more code goes in, and this PR is not mergeable as it is.

  • tini is split out as Start the box image under tini #97, done the way you suggested: tini in the image's ENTRYPOINT, Init: true dropped for the box and login containers, and kept for oneShot.
  • Service account token, service links, orchestrator NetworkPolicy: all three are fixed on the branch we run internally. Box pods set automountServiceAccountToken: false. Box and login pods set enableServiceLinks: false. k8s/orchestrator.yaml ships a NetworkPolicy that admits only the ingress controller. These will be part of whatever comes out of the concept, not pushed onto this PR.

🤖 Generated with Claude Code

@kossmac kossmac closed this Oct 5, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants