Move a running LLM generation from one GPU to another without losing it.
Verified state fidelity, and fail-closed compatibility checks.
A technical report. Version 1.0 — September 2026.
Build 198f336, certificate cuda_release:b3:1:5639164045f.
Serving a long-context LLM session builds a large, per-session KV cache inside one GPU’s memory. Any operation that removes that GPU from service — patching, rebalancing, scaling down, replacing a failing card — destroys it. The request can be retried, but the prompt is re-read from scratch, and at 16k context that is seconds of GPU time per session, paid twice.
WarmSwap moves the cache instead. A coordinator quiesces the session at a token boundary, transfers the KV state to a second worker, and commits ownership in one atomic step. The session continues at the next token and the client sees no interruption.
The contribution is not that a cache can be copied. It is the conditions under which copying it is safe, and a mechanism that refuses when they do not hold. We define an execution fingerprint over everything that changes what a forward pass computes, issue certificates only from measured evidence, and fail closed when a deployment drifts from the evidence that justified it.
On the evaluated configuration (2× RTX 3090, vLLM 0.11.0, Qwen2.5-7B-Instruct), a resumed session compared against one left alone at the same batch schedule is bitwise identical — zero differing KV positions, maximum logit delta 0.0 across 80 concurrent probes. Under differing schedules the runtime’s own batching moves logits whether or not anything migrated (§7), so the claim is state fidelity, not schedule-independent output.
The migration pause is 0.14–0.32× the cost of recomputing the same prompt, and running the layer costs a normal request +4.0% p95 token latency and 1.1% throughput.
A session’s state is not in the application. It is in the accelerator: for a 16k-token session on a 7B model, roughly 0.9 GB of key/value tensors, distributed across paged blocks in one GPU’s memory. Three properties make this awkward.
Existing options each give up one of these. Draining and retrying pays (1). Draining slowly (“cordon and wait”) pays availability instead, and cannot bound how long a session runs. Checkpoint/restore of the whole process (CRIU-style) moves far more than the session and still cannot span two live engines. Live VM migration (Clark et al., NSDI 2005) solved the analogous problem for virtual machines by moving memory underneath a running guest; the GPU case differs in that the state lives in an accelerator managed by a userspace engine, and correctness is judged by generated tokens rather than by the guest’s own consistency.
The evaluated system moves a session between GPUs on one host, under greedy decoding, on
one certified deployment tuple at a time. The coordinator holds a worker per GPU and
certifies every ordered pair among them, so a four-GPU host is twelve directed edges under
one certificate; the measurements in §5 were taken on two, which is the smallest deployment
that can migrate at all. Two defaults bound a larger host and are configuration rather than
protocol: max_active_sessions (four, service-wide) and max_concurrent_migrations (one,
because the shared arena holds a single snapshot).
What it does not do is enumerated in limits.md; cross-host migration is designed and unbuilt, and the reasons are architectural rather than incidental (§8).
A coordinator owns every session: its transcript, its durable state, and the right to say which worker owns it. It holds a worker per GPU — two in the evaluated configuration, more on a larger host — and a drain names the target it wants among them. Workers own GPUs and speak only to the coordinator, over private Unix-domain sockets in a service-owned directory. Clients speak only to the coordinator.
The coordinator never touches a GPU. KV bytes move through a page-locked shared arena that every worker on the host maps at startup, so a migration copies them once, through host memory, never through the coordinator and never over a socket.
A session cannot be captured mid-token. Between tokens it can: at that moment its state is
exactly the accepted transcript T of length N, a pending token T[K] where K = N-1,
and the KV for positions [0, K).
Two rules make that boundary transferable:
Seven phases, exactly one of them irreversible:
| Phase | Property | |
|---|---|---|
| 1 | Quiesce | Stop issuing decode permits; wait for the in-flight token |
| 2 | Reserve | Target commits KV blocks, completion budget, staging — or refuses |
| 3 | Freeze & export | Boundary taken; manifest names every chunk and checksum |
| 4 | Transfer & validate | Target re-checksums and rebuilds the boundary |
| 5 | Commit | One atomic write moves ownership to a new epoch |
| 6 | Activate | Target resumes at the next position |
| 7 | Release | Source drops state; blocks return to the pool |
Everything before (5) is abandonable, and abandoning it leaves the session generating on the source as though nothing happened. After (5), the system enters recovery rather than rewriting the committed owner backwards. Ownership is an integer epoch: a proposal from a stale epoch is rejected, which is what makes a partially-observed handoff safe.
The design target is that a migration either completes or leaves no trace. Faults were injected at each of the twelve named points in the protocol and the response asserted specifically — not that “an error happened”, but that the source still owns the session, the target holds no reservation, and generation continued.
Restoring KV state computed under one execution plan into an engine running a different plan is undefined behaviour dressed as an optimisation. WarmSwap hashes the plan into an execution fingerprint: weights and tokenizer digests, model configuration, RoPE configuration, dtype and KV dtype, attention backend and mask semantics, tensor and pipeline parallel sizes, block size, layer and head geometry, GPU class, adapter build, runtime version — and the engine settings that change what a step computes: prefix caching, eager execution, quantization, speculative decoding, and whether the engine was started with LoRA.
Those last five are read from the live engine, not from the configuration we asked for, and the adapter refuses to start on an engine it cannot interrogate. An engine whose settings cannot be read is an engine nobody has measured.
A certificate states: these two fingerprints, this build, this evidence profile, measured on
this date, expiring on that one. At startup the coordinator registers a migration path only
where a certificate covers the pair. With none, /v1/plans returns an empty list and a
refusal in words, and the service keeps serving.
Certificates are issued only from evaluation output, never from passing unit tests, and
the issuing tool reports not_run, inconclusive and failed as first-class outcomes. A
gate that did not run never becomes a gate that passed.
Three arms, matched: U (WarmSwap, native migration), B0 (the engine alone, no coordinator — the overhead baseline), and B2 (replay: recompute the prompt on the target — the pause baseline). Gates G1–G6 are defined over these arms with preregistered thresholds; G7 requires a partner workflow and is unrun.
Two instrument corrections were necessary and are worth recording, because both initially produced false results:
Certified run, build 198f336, 2× RTX 3090, vLLM 0.11.0, Qwen2.5-7B-Instruct, whole machine
(48/48 CPUs), PCIe 4.0 ×16:
| Gate | Measurement | Threshold | Result |
|---|---|---|---|
| G1 | Protocol suite on real workers: 31 passed, 2 skipped | all pass | pass |
| G2 | Native migration rate per cohort: 105/105, 108/108, 105/105 | floors met | pass |
| G3 | p95 pause ÷ replay: 0.142 / 0.178 / 0.226 / 0.317 | ≤ 0.5 | pass |
| G4 | Throughput loss 0.7–1.1%; p95 token latency +4.0% / +2.8% | ≤10%, ≤+10% | pass |
| G5 | 0 position mismatches, 0 differing KV positions, max logit delta 0.0; task degradation upper bound 0.0084 | exact | pass |
| G6 | 1000 migrate/cancel cycles | no residue | pass |
| G7 | — | partner workflow | not run |
Requirements R01–R12 were established on hardware, with a 52-check drill covering identity, retention, snapshot access and resource limits.
For scale: at 15,872 tokens the native pause was ~0.6 s against ~4.2 s to replay the same prompt, moving ~0.9 GB through the arena.
The gates earned their cost by failing.
The coordinator’s CPU is part of the deployment. On a shared host (32 of 64 CPUs, ~45%
of the machine busy with other tenants) G4 failed at +22–26% p95 — while the same build
passed at +2–4% on machines with their own cores. Ten candidate causes were eliminated one at
a time (storage fsync, CFS throttling, core pinning, Python GC, the GIL, logging, timers, GPU
clocks, prestep, the code itself — by running the previously-passing build on the failing
host). The cause was scheduler run-queue delay: measured from /proc/<pid>/schedstat, the
coordinator waited 1.09 ms for a core per 1 ms it ran, against 0.001 on a dedicated
machine. It sits on every token’s path, so that delay is per-token latency. This is now a
deployment requirement, and G4 records run-queue delay beside its result.
A fixed timeout in quiesce. One migration in sixteen aborted at 15,872 tokens with “cannot freeze with a decode permit outstanding”. The producer’s stand-down allowed the in-flight token a fixed one second; a token queued behind another request’s 16k prefill legitimately takes longer. The abort was safe — the source kept the session — but it cost the drain. Quiesce now waits for the preparation phase’s budget.
A fingerprint that named the wrong thing. The fingerprint’s max_position_embeddings
carried the deployment’s max_model_len rather than the model’s own limit. Since each
evaluation harness chose a different context length, the gates were measured across five
different fingerprints, and the certificate named a 2048-token configuration that no
performance gate measured and no deployment runs. It failed closed — nothing unsafe occurred
— but nothing would have migrated either. The field now carries the model’s limit, verified
identical at 2048 and 17408.
Workers expose no public generation or snapshot endpoint. Control sockets live in a 0700 service-owned directory whose ownership and mode are verified before binding, with a shared secret in a 0600 file; every reply is bound to the request id, the operation id and the worker’s incarnation. Sessions are namespaced per credential, transcripts are retained under a stated policy, and a restart reconciles every non-terminal session against what the workers actually hold.
Operationally the requirements are few and specific: reserved CPU for the coordinator, a
pre-populated model cache with the Hub unreachable at run time (weights are half the
fingerprint), --ipc=host or an equivalently sized /dev/shm for the arena, and monitoring
of warmswap_certificate_expires_seconds — migration stops when a certificate lapses, while
serving continues.
Not the transport interface: the coordinator already runs a migration to completion over a transport whose receiver shares no memory with the sender. Three things block it:
A cross-host design must also answer an economic question, not just a technical one. A 0.9 GB transfer is ~0.7 s at 10 Gbps line rate before overhead, while replay onto a target whose prefix cache already holds the prompt measured 112 ms against native transfer’s 764 ms at 16k. Native migration earns its place when the target is cold and the link is fast. A layer that cannot tell the difference will lose to replay in production, which is why a policy engine is the next design step rather than a bigger pipe.
The evaluation harness is in the repository and runs unattended in about two hours on a
two-GPU host: experiments/gpu/host/chain.sh. It installs the pinned runtime, SHA-verifies
the model, asks the engine to describe itself, runs the protocol suite against real workers,
then every gate — stopping at the first failure and naming it. experiments/evidence/ builds
the matrix and issues or refuses the certificate.
The raw bundles behind §5.2 — including the runs that failed — are retained per host, and the
development history that produced them is in doc/internal/work-log/.
WarmSwap is licensed under Apache 2.0; see /LICENSE. Certificates are produced by the
operator, on the operator’s own hardware; see /docs/certification.md.