WarmSwap

Move a running LLM generation from one GPU to another without losing it.

View My GitHub Profile

Quickstart

WarmSwap moves a running generation from one GPU worker to another on the same host, without losing the session. This page gets you from nothing to a session that survives a drain, and then to a certificate for your machine.

Read it in order. Step 4 is the one people skip and then wonder why migration refuses.

Status. Steps 4 and 5 are what every result in doc/internal/work-log/ was produced by, and the evidence/ bundles are their output. The container image in step 1 has not yet been built and run on a GPU host — the Dockerfile pins the evaluated tuple and passes docker build --check, and that is all it has been through so far. Until someone runs it end to end, treat steps 1–3 as the install path they are meant to be, not as one this repository has evidence for.

What you need

1. Build the image

docker build -t warmswap:0.1.0 .

The image pins the tuple the evidence was measured on: vLLM 0.11.0, torch 2.8.0+cu128, transformers 4.57.6, CPython 3.12. Changing any of them means re-running step 4.

2. Start one worker per GPU

docker run -d --name worker-a --gpus '"device=0"' --ipc=host \
  -v /srv/hf:/hf -e HF_HOME=/hf -v /run/warmswap:/run/warmswap \
  warmswap:0.1.0 worker \
  --worker-id worker-a --socket-dir /run/warmswap/sockets \
  --adapter vllm --model Qwen/Qwen2.5-7B-Instruct \
  --revision a09a35458c702b33eeacc393d103063234e8bc28 \
  --gpu-memory-utilization 0.88 --max-model-len 16384 \
  --certificate-id "$CERTIFICATE_ID"

Repeat per GPU — worker-b on device=1, worker-c on device=2, and so on. Each prints one JSON line when it is ready, carrying the incarnation and socket paths the next step needs. Every worker takes the same model, revision and settings: they are hashed into the fingerprint, and a worker that differs is one no certificate covers.

--ipc=host is required: both workers and the coordinator map one page-locked transfer arena through /dev/shm, and Docker’s default 64 MB is far too small.

3. Start the coordinator

Write deployment.json as docs/production.md describes — it is the only file the coordinator reads — then:

docker run -d --name warmswap --gpus none --ipc=host --network host \
  -v /run/warmswap:/run/warmswap -v /etc/warmswap:/etc/warmswap \
  warmswap:0.1.0 serve --deployment /etc/warmswap/deployment.json

GET /health/ready answers 200 once it is up. GET /v1/plans tells you which worker pairs may migrate natively — and, when none may, exactly why.

4. Certify your machine

Until this passes, native migration is unavailable. That is deliberate: a certificate names the hardware and software its evidence was earned on, and WarmSwap refuses to move a session between workers nothing has measured. Sessions still serve; they just will not move.

# on the GPU host: stage src/, tests/, doc/, pyproject.toml, experiments/gpu/*.py
# and experiments/gpu/host/* into /workspace/ws2 (experiments/gpu/host/README.md
# says exactly what and why), then:
cd /workspace/ws2
WS_BUILD=$(git -C /path/to/repo rev-parse --short HEAD) \
  nohup setsid ./chain.sh > chain.log 2>&1 < /dev/null &   # ~2 hours, unattended

# back on your machine, once it finishes:
experiments/evidence/collect_host_evidence.sh "ssh <host>" evidence/mybox
WS_BUNDLE=evidence/mybox bash experiments/evidence/render_t19.sh
WS_BUNDLE=evidence/mybox WS_ADAPTER_BUILD=$(git rev-parse HEAD) \
  python experiments/evidence/certify.py                   # issues, or refuses with reasons

It runs the protocol suite against real workers, then every gate: exact continuation, migration pause against a replay baseline, serving overhead, leak drill. It stops at the first failure and says which. A pass writes CERTIFICATE.json; point certificate_path at it and restart the coordinator.

If the serving-overhead gate (G4) fails, check the run-queue log written beside it (g4_*_sched.log) before suspecting the code: on every host where it failed so far, the coordinator was waiting for a CPU it had to share.

5. Watch a session survive a drain

curl -XPOST localhost:8080/v1/workers/worker-a/drains \
  -H "Authorization: Bearer $OPERATOR_TOKEN" -H 'content-type: application/json' \
  -d '{"target_worker_id":"worker-b","deadline":"'"$(date -u -d '+2 min' +%FT%TZ)"'"}'

The drain report names every session it moved, the pause each one took, and anything it refused to move, with the reason. A client streaming tokens sees the stream continue.

What this does not do yet