Conversation
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Run against kind with Calico, the box pod, the review, the terminal and the egress check each failed in a way the faked API could not show: - The home seed ran as root without capabilities, so cp could not read the agent's 0700 home in the image and every pod stayed in Init. The setup init container now keeps CHOWN, DAC_OVERRIDE and FOWNER, as docker.ts's root helper does, seeds a home once rather than at every start, and gives the workspace and Nix store roots to the agent. - The orchestrator's Role lacked get on pods/exec, which the client library's WebSocket exec needs, so review and threads got a 403. - A terminal fed the pty's output back to the shell as input: one stream was both stdin and stdout. It now has two, and resizes the pty through the library's terminal-size protocol. - A stop returned while the pod was still terminating, so a start right after it found the old pod and was left with none. A stop now waits for the pod to go, with Docker's 10 second grace. - ensureProxyAttached reported an existing NetworkPolicy as a missing attachment, which logged every running box every minute. - A path with .. in it answered 500 rather than 404. The smoke test runs under macOS bash and checks the Nix store too. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A Kubernetes runtime knew no image at all, so a box stayed on the image it was created with however often BOX_IMAGE moved. Every start makes a new pod anyway, so the pod is now made from BOX_IMAGE and the row records it, and the runtime knows an image by its reference: a box whose running pod is on another reference moves at its next start, as a Docker box rolls onto a new image id. That takes a tag per build; a moving latest is still not seen. A pod gone after a stop is the normal case here, so it is logged as a start rather than warned about as a lost container. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A box had a claim each for its workspace, home and Nix store, so every running box attached three block volumes to its node. Hetzner attaches at most 16 to a server, which stopped a node at five running boxes. The three are now subPaths of one claim, boxes-data-<id>, named once among the pod's volumes, and K8S_VOLUME_SIZE replaces the three sizes. ensureAddedVolumes no longer makes a claim: every Kubernetes box has had its one claim since creation, and an empty stand-in for a lost one would only hide the loss. Boxes made with three claims are not carried over; none exist yet. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019GcXSze2tsHcn3F59k9VeY
A login always ran in a Docker container, so on Kubernetes the settings page failed with connect ENOENT /var/run/docker.sock. It now runs in a throwaway pod from the box image, hardened like a box pod, with its home and /tmp in memory and no service account token. The pod carries only the login label, so no box NetworkPolicy selects it and it reaches the internet directly, as the Docker login container does on the bridge. A pod is Pending until scheduled and pulled, so the login waits for it to run, and gives up with the reason when its image cannot be pulled. The orphan sweep lists login pods too, so a restart mid-login no longer leaves one behind. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A box pod ran the entrypoint as PID 1, so its `sleep infinity` became the parent of every process orphaned in the box and never waited on one. Each command whose shell went first stayed a zombie, held a pid against the limit, and read as work that kept the box awake. The kernel also discarded SIGTERM for it, so every stop waited out the grace period. The image now installs tini, and the box pod runs the entrypoint under it, as Docker's `Init: true` does. The Docker runtime and the login pod are unchanged. A box pod needs an image built from this commit. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
SummaryThe missing features block a merge. The review with git status and diff, attachments and agent configuration are what make Boxes more than a container with an agent in it. A Kubernetes runtime without them is not a usable alternative. The gaps are not separate bugs. They come from the runtime layer, which does not cover the thing these features depend on. Before more code is written, there needs to be a clear concept of what this layer has to be able to do, and how each runtime does it. Why the features are missingLarge parts of the orchestrator assume that a box's files are a directory on the orchestrator's own filesystem. They also assume this directory can be read and written whether or not the box runs. Docker makes this true through bind mounts from
The PR adds Logins are outside the What a concept needs to answer
Findings independent of the concept
tiniDocker's ENTRYPOINT ["/usr/bin/tini", "--", "/usr/local/bin/entrypoint.sh"]
Smaller PRsOnce there is a concept, smaller PRs towards it may be easier to get merged, each one useful on its own. The tini change is one of them. Out of scope: multiple usersThis is not part of this PR and does not need to be solved here. Boxes is a tool for one person: it has no accounts, and credentials, settings and the box list belong to the whole deployment. On a shared cluster, supporting several users becomes more reasonable. The user would likely be passed in as a request header by the reverse proxy that handles authentication. The runtime concept should not rule this out. |
|
Thanks for the thorough review. Agreed: the runtime layer needs a concept before more code goes in, and this PR is not mergeable as it is.
🤖 Generated with Claude Code |
Adds Kubernetes as a second, experimental runtime for boxes, next to Docker. Docker stays the default; a deployment switches with
RUNTIME=kubernetes.What changed
docker.tsnow goes throughruntime()(orchestrator/src/runtime/), with a Docker and a Kubernetes implementation. File access for review, uploads and the workspace browser goes throughfileAccess()the same way (orchestrator/src/fileaccess/). WithRUNTIME=dockerthe behaviour is meant to be unchanged.orchestrator/src/kubernetes.ts,@kubernetes/client-node): a box is a pod with one PVC (workspace, home and Nix store as subPaths) and aNetworkPolicythat only lets it reach the egress proxy. File access in a pod runs a small script over exec (pod-fs-script.ts). Orphaned pods, PVCs and policies are swept like Docker's containers, networks and volumes.Init: truedoes, so orphaned processes are reaped and SIGTERM reaches the box.k8s/, andtests/smoke-test-k8s.sh, which also checks that the cluster's CNI actually enforcesNetworkPolicy.Rebase onto current main
The work was written against
ee65e63. Rebasing it onto main also routesprocessReadings()for the new tunnel reconciler through the runtime instead ofdk.Testing
npm run checkandnpm testinorchestrator,proxyanddashboard. One test,review/fs.test.ts > a file that need not exist yet resolves…, fails on macOS because of the/var→/private/varsymlink. It fails the same way onmainand is unrelated to this change.tests/smoke-test-k8s.sh) and on an internal cluster.🤖 Generated with Claude Code