Move a running LLM generation from one GPU to another without losing it.
WarmSwap will not move a session between any two workers unless something has measured that moving it is exact on that hardware, with that build. The artefact that records the measurement is a certificate, and producing one for your machine is a one-off, unattended job of roughly two hours.
Until it passes, the service serves normally and refuses to migrate. That is the intended failure mode.
The certificate names an execution fingerprint: model revision, weights and tokenizer digests, dtype, KV dtype, attention backend, block size, GPU class, and the engine settings that change what a forward pass computes (prefix caching, eager execution, quantization, LoRA). Change any of them and you have a different plan, and no evidence about it.
Two findings from our own runs explain why this is strict:
max_model_len rather than the model’s own
positional limit produced a certificate naming a configuration nobody runs. It fails
closed, so nothing unsafe happened — but nothing would have migrated either.One certificate covers a whole host when its GPUs are identical: the evaluation records a fingerprint per worker, and the certificate names every ordered pair between them — four GPUs is twelve directed edges. If one worker reports a different fingerprint from the rest, certification is refused and the message says which worker and what it reported, because a certificate here names one source and one target fingerprint and cannot describe a mixed host.
On the GPU host, with the repository staged (see experiments/gpu/host/README.md for exactly
what to copy and why):
cd /workspace/ws2
WS_BUILD=$(git -C /path/to/repo rev-parse --short HEAD) \
nohup setsid ./chain.sh > chain.log 2>&1 < /dev/null &
Launch it detached. A stage killed by a dropped SSH session leaves a half-written result.
The chain installs the pinned runtime, fetches and SHA-verifies the model, asks the engine to describe itself, runs the protocol suite against real workers, then every gate — stopping at the first thing that does not pass and saying which.
Then, on your machine:
experiments/evidence/collect_host_evidence.sh "ssh <host>" evidence/mybox
WS_BUNDLE=evidence/mybox bash experiments/evidence/render_t19.sh # the matrix
WS_BUNDLE=evidence/mybox WS_ADAPTER_BUILD=$(git rev-parse HEAD) \
python experiments/evidence/certify.py # issues, or refuses
A pass writes CERTIFICATE.json. Point certificate_path at it in deployment.json and
restart the coordinator; GET /v1/plans should now list an evaluated plan.
| Gate | Question | Threshold |
|---|---|---|
| G1 | Does the protocol hold against real workers, including every failure-matrix row? | all pass |
| G2 | Do migrations actually happen natively, reusing the cache? | native rate and reuse floors |
| G3 | Is the pause cheaper than recomputing the prompt? | p95 pause ≤ 50% of replay |
| G4 | What does running WarmSwap cost a normal request? | ≤10% throughput, ≤+10% p95 latency |
| G5 | Is the continuation exact? | identical logits and KV; task degradation bounded |
| G6 | Does repeated migration leak? | 1000 cycles, no residue |
| G7 | Has a real user’s workflow been validated? | needs a partner — unrun |
not_run, inconclusive and failed are first-class outcomes. Nothing is ever promoted to
a pass for want of evidence, and certify.py lists precisely what blocked it.
g4_*_sched.log). On every host
where it failed for us, the coordinator was waiting for a CPU it had to share.Certificates expire — 30 days by default. Native migration stops when they do, quietly, while
serving continues. Watch warmswap_certificate_expires_seconds; the coordinator also warns
in its logs from a week out.
A certificate is a record of measurements taken under stated conditions on a stated date, by
you, on your own hardware. It is not a warranty, a guarantee, or a service-level commitment,
and it does not transfer operational risk to anyone. Nobody countersigns it and nothing
phones home: certify.py writes it from your evidence bundle, the coordinator reads it from
the path in your deployment file, and the whole check is local.
It covers same-host GPU-to-GPU migration on the tuple measured. See limits.md for what the software does not do.
Publishing what you measure — including results that are unfavourable — is welcome. State the deployment tuple and the version alongside the numbers, so a reader knows what they describe.