WarmSwap

Move a running LLM generation from one GPU to another without losing it.

View My GitHub Profile

WarmSwap: safe live KV-cache migration across GPU worker drains

Verified state fidelity, and fail-closed compatibility checks.

A technical report. Version 1.0 — September 2026. Build 198f336, certificate cuda_release:b3:1:5639164045f.


Abstract

Serving a long-context LLM session builds a large, per-session KV cache inside one GPU’s memory. Any operation that removes that GPU from service — patching, rebalancing, scaling down, replacing a failing card — destroys it. The request can be retried, but the prompt is re-read from scratch, and at 16k context that is seconds of GPU time per session, paid twice.

WarmSwap moves the cache instead. A coordinator quiesces the session at a token boundary, transfers the KV state to a second worker, and commits ownership in one atomic step. The session continues at the next token and the client sees no interruption.

The contribution is not that a cache can be copied. It is the conditions under which copying it is safe, and a mechanism that refuses when they do not hold. We define an execution fingerprint over everything that changes what a forward pass computes, issue certificates only from measured evidence, and fail closed when a deployment drifts from the evidence that justified it.

On the evaluated configuration (2× RTX 3090, vLLM 0.11.0, Qwen2.5-7B-Instruct), a resumed session compared against one left alone at the same batch schedule is bitwise identical — zero differing KV positions, maximum logit delta 0.0 across 80 concurrent probes. Under differing schedules the runtime’s own batching moves logits whether or not anything migrated (§7), so the claim is state fidelity, not schedule-independent output.

The migration pause is 0.14–0.32× the cost of recomputing the same prompt, and running the layer costs a normal request +4.0% p95 token latency and 1.1% throughput.


1. The problem

A session’s state is not in the application. It is in the accelerator: for a 16k-token session on a 7B model, roughly 0.9 GB of key/value tensors, distributed across paged blocks in one GPU’s memory. Three properties make this awkward.

  1. It is expensive to rebuild. Re-prefilling 16k tokens took ~4.2 s of GPU time in our measurements. That cost lands on a user who has already waited once.
  2. It is invisible to the scheduler. Kubernetes can drain a node and a load balancer can shift traffic, but neither can move what lives in device memory.
  3. It is exact or it is wrong. An LLM’s output is a deterministic function of its state under greedy decoding. Restore that state approximately and the continuation diverges — quietly, and in a way no status code reports.

Existing options each give up one of these. Draining and retrying pays (1). Draining slowly (“cordon and wait”) pays availability instead, and cannot bound how long a session runs. Checkpoint/restore of the whole process (CRIU-style) moves far more than the session and still cannot span two live engines. Live VM migration (Clark et al., NSDI 2005) solved the analogous problem for virtual machines by moving memory underneath a running guest; the GPU case differs in that the state lives in an accelerator managed by a userspace engine, and correctness is judged by generated tokens rather than by the guest’s own consistency.

2. Scope

The evaluated system moves a session between GPUs on one host, under greedy decoding, on one certified deployment tuple at a time. The coordinator holds a worker per GPU and certifies every ordered pair among them, so a four-GPU host is twelve directed edges under one certificate; the measurements in §5 were taken on two, which is the smallest deployment that can migrate at all. Two defaults bound a larger host and are configuration rather than protocol: max_active_sessions (four, service-wide) and max_concurrent_migrations (one, because the shared arena holds a single snapshot).

What it does not do is enumerated in limits.md; cross-host migration is designed and unbuilt, and the reasons are architectural rather than incidental (§8).

3. Design

3.1 One coordinator, one worker per GPU

A coordinator owns every session: its transcript, its durable state, and the right to say which worker owns it. It holds a worker per GPU — two in the evaluated configuration, more on a larger host — and a drain names the target it wants among them. Workers own GPUs and speak only to the coordinator, over private Unix-domain sockets in a service-owned directory. Clients speak only to the coordinator.

The coordinator never touches a GPU. KV bytes move through a page-locked shared arena that every worker on the host maps at startup, so a migration copies them once, through host memory, never through the coordinator and never over a socket.

3.2 The canonical boundary

A session cannot be captured mid-token. Between tokens it can: at that moment its state is exactly the accepted transcript T of length N, a pending token T[K] where K = N-1, and the KV for positions [0, K).

Two rules make that boundary transferable:

3.3 The transaction

Seven phases, exactly one of them irreversible:

  Phase Property
1 Quiesce Stop issuing decode permits; wait for the in-flight token
2 Reserve Target commits KV blocks, completion budget, staging — or refuses
3 Freeze & export Boundary taken; manifest names every chunk and checksum
4 Transfer & validate Target re-checksums and rebuilds the boundary
5 Commit One atomic write moves ownership to a new epoch
6 Activate Target resumes at the next position
7 Release Source drops state; blocks return to the pool

Everything before (5) is abandonable, and abandoning it leaves the session generating on the source as though nothing happened. After (5), the system enters recovery rather than rewriting the committed owner backwards. Ownership is an integer epoch: a proposal from a stale epoch is rejected, which is what makes a partially-observed handoff safe.

3.4 Failure semantics

The design target is that a migration either completes or leaves no trace. Faults were injected at each of the twelve named points in the protocol and the response asserted specifically — not that “an error happened”, but that the source still owns the session, the target holds no reservation, and generation continued.

4. Compatibility: fingerprints and certificates

4.1 The fingerprint

Restoring KV state computed under one execution plan into an engine running a different plan is undefined behaviour dressed as an optimisation. WarmSwap hashes the plan into an execution fingerprint: weights and tokenizer digests, model configuration, RoPE configuration, dtype and KV dtype, attention backend and mask semantics, tensor and pipeline parallel sizes, block size, layer and head geometry, GPU class, adapter build, runtime version — and the engine settings that change what a step computes: prefix caching, eager execution, quantization, speculative decoding, and whether the engine was started with LoRA.

Those last five are read from the live engine, not from the configuration we asked for, and the adapter refuses to start on an engine it cannot interrogate. An engine whose settings cannot be read is an engine nobody has measured.

4.2 The certificate

A certificate states: these two fingerprints, this build, this evidence profile, measured on this date, expiring on that one. At startup the coordinator registers a migration path only where a certificate covers the pair. With none, /v1/plans returns an empty list and a refusal in words, and the service keeps serving.

Certificates are issued only from evaluation output, never from passing unit tests, and the issuing tool reports not_run, inconclusive and failed as first-class outcomes. A gate that did not run never becomes a gate that passed.

5. Evaluation

5.1 Method

Three arms, matched: U (WarmSwap, native migration), B0 (the engine alone, no coordinator — the overhead baseline), and B2 (replay: recompute the prompt on the target — the pause baseline). Gates G1–G6 are defined over these arms with preregistered thresholds; G7 requires a partner workflow and is unrun.

Two instrument corrections were necessary and are worth recording, because both initially produced false results:

5.2 Results

Certified run, build 198f336, 2× RTX 3090, vLLM 0.11.0, Qwen2.5-7B-Instruct, whole machine (48/48 CPUs), PCIe 4.0 ×16:

Gate Measurement Threshold Result
G1 Protocol suite on real workers: 31 passed, 2 skipped all pass pass
G2 Native migration rate per cohort: 105/105, 108/108, 105/105 floors met pass
G3 p95 pause ÷ replay: 0.142 / 0.178 / 0.226 / 0.317 ≤ 0.5 pass
G4 Throughput loss 0.7–1.1%; p95 token latency +4.0% / +2.8% ≤10%, ≤+10% pass
G5 0 position mismatches, 0 differing KV positions, max logit delta 0.0; task degradation upper bound 0.0084 exact pass
G6 1000 migrate/cancel cycles no residue pass
G7 — partner workflow not run

Requirements R01–R12 were established on hardware, with a 52-check drill covering identity, retention, snapshot access and resource limits.

For scale: at 15,872 tokens the native pause was ~0.6 s against ~4.2 s to replay the same prompt, moving ~0.9 GB through the arena.

5.3 What the evaluation caught

The gates earned their cost by failing.

The coordinator’s CPU is part of the deployment. On a shared host (32 of 64 CPUs, ~45% of the machine busy with other tenants) G4 failed at +22–26% p95 — while the same build passed at +2–4% on machines with their own cores. Ten candidate causes were eliminated one at a time (storage fsync, CFS throttling, core pinning, Python GC, the GIL, logging, timers, GPU clocks, prestep, the code itself — by running the previously-passing build on the failing host). The cause was scheduler run-queue delay: measured from /proc/<pid>/schedstat, the coordinator waited 1.09 ms for a core per 1 ms it ran, against 0.001 on a dedicated machine. It sits on every token’s path, so that delay is per-token latency. This is now a deployment requirement, and G4 records run-queue delay beside its result.

A fixed timeout in quiesce. One migration in sixteen aborted at 15,872 tokens with “cannot freeze with a decode permit outstanding”. The producer’s stand-down allowed the in-flight token a fixed one second; a token queued behind another request’s 16k prefill legitimately takes longer. The abort was safe — the source kept the session — but it cost the drain. Quiesce now waits for the preparation phase’s budget.

A fingerprint that named the wrong thing. The fingerprint’s max_position_embeddings carried the deployment’s max_model_len rather than the model’s own limit. Since each evaluation harness chose a different context length, the gates were measured across five different fingerprints, and the certificate named a 2048-token configuration that no performance gate measured and no deployment runs. It failed closed — nothing unsafe occurred — but nothing would have migrated either. The field now carries the model’s limit, verified identical at 2048 and 17408.

6. Security and operations

Workers expose no public generation or snapshot endpoint. Control sockets live in a 0700 service-owned directory whose ownership and mode are verified before binding, with a shared secret in a 0600 file; every reply is bound to the request id, the operation id and the worker’s incarnation. Sessions are namespaced per credential, transcripts are retained under a stated policy, and a restart reconciles every non-terminal session against what the workers actually hold.

Operationally the requirements are few and specific: reserved CPU for the coordinator, a pre-populated model cache with the Hub unreachable at run time (weights are half the fingerprint), --ipc=host or an equivalently sized /dev/shm for the arena, and monitoring of warmswap_certificate_expires_seconds — migration stops when a certificate lapses, while serving continues.

7. Limitations and threats to validity

8. What blocks cross-host migration

Not the transport interface: the coordinator already runs a migration to completion over a transport whose receiver shares no memory with the sender. Three things block it:

  1. Export writes into a host-shared arena, and import re-checksums by reading the same arena.
  2. Import installs state through in-process engine hooks — the block pool and prefix cache of a vLLM engine in this process.
  3. Nothing authenticates a transfer between hosts; the payload socket’s security is the directory’s mode.

A cross-host design must also answer an economic question, not just a technical one. A 0.9 GB transfer is ~0.7 s at 10 Gbps line rate before overhead, while replay onto a target whose prefix cache already holds the prompt measured 112 ms against native transfer’s 764 ms at 16k. Native migration earns its place when the target is cold and the link is fast. A layer that cannot tell the difference will lose to replay in production, which is why a policy engine is the next design step rather than a bigger pipe.

9. Reproducing this

The evaluation harness is in the repository and runs unattended in about two hours on a two-GPU host: experiments/gpu/host/chain.sh. It installs the pinned runtime, SHA-verifies the model, asks the engine to describe itself, runs the protocol suite against real workers, then every gate — stopping at the first failure and naming it. experiments/evidence/ builds the matrix and issues or refuses the certificate.

The raw bundles behind §5.2 — including the runs that failed — are retained per host, and the development history that produced them is in doc/internal/work-log/.


WarmSwap is licensed under Apache 2.0; see /LICENSE. Certificates are produced by the operator, on the operator’s own hardware; see /docs/certification.md.