A live generation's real state is the KV cache sitting in one GPU's memory. Move the request and you leave it behind — the model re-reads a prompt the user already paid for. WarmSwap moves the cache, and the session continues on the next GPU at the next token. Press drain and watch it happen.
Two GPUs is the smallest deployment that can migrate anything. Four or eight is one worker per card and a single certificate covering every pair between them.
So the protocol takes it to a point where its state is exactly describable, and only then moves it. Everything before the commit can be abandoned, and abandoning it leaves the session generating on the source as though nothing happened.
| 01 | quiesce | Stop issuing decode permits, wait for the one in flight. Between tokens is the only clean boundary. |
| 02 | reserve | Ask the target for KV blocks, completion budget, staging. A refusal costs nothing: the session never stopped. |
| 03 | freeze | Take the canonical boundary into a manifest naming every chunk and its checksum. |
| 04 | transfer | The target re-checksums each chunk and rebuilds the boundary before it owns anything. |
| 05 | commit | One atomic write moves ownership to a new epoch. A single row, and the only step that cannot be undone. |
| 06 | activate | The target resumes at the next position; the source drops its state and returns its blocks. |
2× RTX 3090, vLLM 0.11.0, Qwen2.5-7B-Instruct. The evidence bundles and the certificate they produced are in the repository — including the runs that failed, and why.
A migration is only exact if both GPUs run the same execution plan. WarmSwap hashes that plan into a fingerprint and moves a session only when a certificate says that exact pair was measured and found exact — on this hardware, with this build.
{
"plans": [],
"certification_refusals": [
"evaluation did not issue cuda_release:b3:1:…:
G4: failed — p95 latency +26% (max 10%)"
]
}
Serving carries on. Nothing migrates, and the reason is a sentence rather than a silent fallback.
# build the image, pinned to the evaluated tuple docker build -t warmswap:0.1.0 . # one worker per GPU (--ipc=host: both map one shared arena) docker run -d --gpus '"device=0"' --ipc=host warmswap:0.1.0 worker \ --worker-id worker-a --socket-dir /run/warmswap/sockets \ --adapter vllm --model Qwen/Qwen2.5-7B-Instruct # the coordinator holds no GPU at all docker run -d --ipc=host --network host warmswap:0.1.0 serve \ --deployment /etc/warmswap/deployment.json # move every session off worker-a, before a deadline curl -XPOST localhost:8080/v1/workers/worker-a/drains \ -d '{"target_worker_id":"worker-b","deadline":"…"}'
The quickstart goes from a bare two-GPU host to a session surviving a drain. The white paper is the evaluation: method, results, and what the gates caught.
the licence
the certificate
producing one
A deployment that has drifted from the evidence justifying it is one the system refuses to migrate. That check is the product, and it is yours to run.