WarmSwap

Move a running LLM generation from one GPU to another without losing it.

View My GitHub Profile

Certifying your machine

WarmSwap will not move a session between any two workers unless something has measured that moving it is exact on that hardware, with that build. The artefact that records the measurement is a certificate, and producing one for your machine is a one-off, unattended job of roughly two hours.

Until it passes, the service serves normally and refuses to migrate. That is the intended failure mode.

Why it is per-machine

The certificate names an execution fingerprint: model revision, weights and tokenizer digests, dtype, KV dtype, attention backend, block size, GPU class, and the engine settings that change what a forward pass computes (prefix caching, eager execution, quantization, LoRA). Change any of them and you have a different plan, and no evidence about it.

Two findings from our own runs explain why this is strict:

More than two GPUs

One certificate covers a whole host when its GPUs are identical: the evaluation records a fingerprint per worker, and the certificate names every ordered pair between them — four GPUs is twelve directed edges. If one worker reports a different fingerprint from the rest, certification is refused and the message says which worker and what it reported, because a certificate here names one source and one target fingerprint and cannot describe a mixed host.

Running it

On the GPU host, with the repository staged (see experiments/gpu/host/README.md for exactly what to copy and why):

cd /workspace/ws2
WS_BUILD=$(git -C /path/to/repo rev-parse --short HEAD) \
  nohup setsid ./chain.sh > chain.log 2>&1 < /dev/null &

Launch it detached. A stage killed by a dropped SSH session leaves a half-written result.

The chain installs the pinned runtime, fetches and SHA-verifies the model, asks the engine to describe itself, runs the protocol suite against real workers, then every gate — stopping at the first thing that does not pass and saying which.

Then, on your machine:

experiments/evidence/collect_host_evidence.sh "ssh <host>" evidence/mybox
WS_BUNDLE=evidence/mybox bash experiments/evidence/render_t19.sh     # the matrix
WS_BUNDLE=evidence/mybox WS_ADAPTER_BUILD=$(git rev-parse HEAD) \
  python experiments/evidence/certify.py                             # issues, or refuses

A pass writes CERTIFICATE.json. Point certificate_path at it in deployment.json and restart the coordinator; GET /v1/plans should now list an evaluated plan.

What the gates check

Gate Question Threshold
G1 Does the protocol hold against real workers, including every failure-matrix row? all pass
G2 Do migrations actually happen natively, reusing the cache? native rate and reuse floors
G3 Is the pause cheaper than recomputing the prompt? p95 pause ≤ 50% of replay
G4 What does running WarmSwap cost a normal request? ≤10% throughput, ≤+10% p95 latency
G5 Is the continuation exact? identical logits and KV; task degradation bounded
G6 Does repeated migration leak? 1000 cycles, no residue
G7 Has a real user’s workflow been validated? needs a partner — unrun

not_run, inconclusive and failed are first-class outcomes. Nothing is ever promoted to a pass for want of evidence, and certify.py lists precisely what blocked it.

Reading a failure

Expiry

Certificates expire — 30 days by default. Native migration stops when they do, quietly, while serving continues. Watch warmswap_certificate_expires_seconds; the coordinator also warns in its logs from a week out.

What a certificate is, and is not

A certificate is a record of measurements taken under stated conditions on a stated date, by you, on your own hardware. It is not a warranty, a guarantee, or a service-level commitment, and it does not transfer operational risk to anyone. Nobody countersigns it and nothing phones home: certify.py writes it from your evidence bundle, the coordinator reads it from the path in your deployment file, and the whole check is local.

It covers same-host GPU-to-GPU migration on the tuple measured. See limits.md for what the software does not do.

Publishing what you measure — including results that are unfavourable — is welcome. State the deployment tuple and the version alongside the numbers, so a reader knows what they describe.