Move a running LLM generation from one GPU to another without losing it.
WarmSwap moves a running generation from one GPU worker to another on the same host, without losing the session. This page gets you from nothing to a session that survives a drain, and then to a certificate for your machine.
Read it in order. Step 4 is the one people skip and then wonder why migration refuses.
Status. Steps 4 and 5 are what every result in
doc/internal/work-log/was produced by, and theevidence/bundles are their output. The container image in step 1 has not yet been built and run on a GPU host — the Dockerfile pins the evaluated tuple and passesdocker build --check, and that is all it has been through so far. Until someone runs it end to end, treat steps 1–3 as the install path they are meant to be, not as one this repository has evidence for.
doc/internal/work-log/T23.md). A machine you have to yourself,
or a cpuset, or Kubernetes’ static CPU manager.docker build -t warmswap:0.1.0 .
The image pins the tuple the evidence was measured on: vLLM 0.11.0, torch 2.8.0+cu128, transformers 4.57.6, CPython 3.12. Changing any of them means re-running step 4.
docker run -d --name worker-a --gpus '"device=0"' --ipc=host \
-v /srv/hf:/hf -e HF_HOME=/hf -v /run/warmswap:/run/warmswap \
warmswap:0.1.0 worker \
--worker-id worker-a --socket-dir /run/warmswap/sockets \
--adapter vllm --model Qwen/Qwen2.5-7B-Instruct \
--revision a09a35458c702b33eeacc393d103063234e8bc28 \
--gpu-memory-utilization 0.88 --max-model-len 16384 \
--certificate-id "$CERTIFICATE_ID"
Repeat per GPU — worker-b on device=1, worker-c on device=2, and so on. Each prints
one JSON line when it is ready, carrying the incarnation and socket paths the next step
needs. Every worker takes the same model, revision and settings: they are hashed into the
fingerprint, and a worker that differs is one no certificate covers.
--ipc=host is required: both workers and the coordinator map one page-locked transfer
arena through /dev/shm, and Docker’s default 64 MB is far too small.
Write deployment.json as docs/production.md
describes — it is the only file the coordinator reads — then:
docker run -d --name warmswap --gpus none --ipc=host --network host \
-v /run/warmswap:/run/warmswap -v /etc/warmswap:/etc/warmswap \
warmswap:0.1.0 serve --deployment /etc/warmswap/deployment.json
GET /health/ready answers 200 once it is up. GET /v1/plans tells you which worker pairs
may migrate natively — and, when none may, exactly why.
Until this passes, native migration is unavailable. That is deliberate: a certificate names the hardware and software its evidence was earned on, and WarmSwap refuses to move a session between workers nothing has measured. Sessions still serve; they just will not move.
# on the GPU host: stage src/, tests/, doc/, pyproject.toml, experiments/gpu/*.py
# and experiments/gpu/host/* into /workspace/ws2 (experiments/gpu/host/README.md
# says exactly what and why), then:
cd /workspace/ws2
WS_BUILD=$(git -C /path/to/repo rev-parse --short HEAD) \
nohup setsid ./chain.sh > chain.log 2>&1 < /dev/null & # ~2 hours, unattended
# back on your machine, once it finishes:
experiments/evidence/collect_host_evidence.sh "ssh <host>" evidence/mybox
WS_BUNDLE=evidence/mybox bash experiments/evidence/render_t19.sh
WS_BUNDLE=evidence/mybox WS_ADAPTER_BUILD=$(git rev-parse HEAD) \
python experiments/evidence/certify.py # issues, or refuses with reasons
It runs the protocol suite against real workers, then every gate: exact continuation,
migration pause against a replay baseline, serving overhead, leak drill. It stops at the
first failure and says which. A pass writes CERTIFICATE.json; point certificate_path at
it and restart the coordinator.
If the serving-overhead gate (G4) fails, check the run-queue log written beside it
(g4_*_sched.log) before suspecting the code: on every host where it failed so far, the
coordinator was waiting for a CPU it had to share.
curl -XPOST localhost:8080/v1/workers/worker-a/drains \
-H "Authorization: Bearer $OPERATOR_TOKEN" -H 'content-type: application/json' \
-d '{"target_worker_id":"worker-b","deadline":"'"$(date -u -d '+2 min' +%FT%TZ)"'"}'
The drain report names every session it moved, the pause each one took, and anything it refused to move, with the reason. A client streaming tokens sees the stream continue.
max_active_sessions in deployment.json on a host with more GPUs — one migration at a
time regardless, one coordinator, SQLite. Enough for a pilot, not for a public service.