WarmSwap 0.1.0
Attaching workers · mapping transfer arena…
worker-a: vllm 0.11.0
readme — WarmSwap

Drain a GPU.
Keep the sessions running on it.

A live generation's real state is the KV cache sitting in one GPU's memory. Move the request and you leave it behind — the model re-reads a prompt the user already paid for. WarmSwap moves the cache, and the session continues on the next GPU at the next token. Press drain and watch it happen.

Two GPUs is the smallest deployment that can migrate anything. Four or eight is one worker per card and a single certificate covering every pair between them.

worker-a · GPU 0
RTX 3090 · 24 GBowner
KV blocks96 / 96
worker-b · GPU 1
RTX 3090 · 24 GBidle
KV blocks0 / 96
client — /v1/sessions/sess_4f2a/events
server-sent events
telemetry
staterunning
owner / epochworker-a / 0
pause—
recomputed0
bytes moved0

A session cannot be
photographed mid-token.

So the protocol takes it to a point where its state is exactly describable, and only then moves it. Everything before the commit can be abandoned, and abandoning it leaves the session generating on the source as though nothing happened.

protocol — seven phases, one irreversible
01quiesceStop issuing decode permits, wait for the one in flight. Between tokens is the only clean boundary.
02reserveAsk the target for KV blocks, completion budget, staging. A refusal costs nothing: the session never stopped.
03freezeTake the canonical boundary into a manifest naming every chunk and its checksum.
04transferThe target re-checksums each chunk and rebuilds the boundary before it owns anything.
05commitOne atomic write moves ownership to a new epoch. A single row, and the only step that cannot be undone.
06activateThe target resumes at the next position; the source drops its state and returns its blocks.
why it is safe
  • Whole blocks only. A boundary mid-block leaves a tail to recompute, and one recomputed position changed the output in 14 of 24 early trials.
  • One execution plan. Prefix caching, eager execution, quantization, LoRA, dtype, attention backend, model revision — hashed into a single fingerprint.
  • Report, never force. A drain that cannot move a session says which one and why, and leaves it running.

Every number here
came off real hardware.

2× RTX 3090, vLLM 0.11.0, Qwen2.5-7B-Instruct. The evidence bundles and the certificate they produced are in the repository — including the runs that failed, and why.

continuation
0.0max logit delta Against a session left alone on the same batch schedule, a resumed session is bitwise identical: zero differing KV positions across 80 concurrent probes.
pause
0.14–0.32×vs. re-prefill Migration pause measured against recomputing the same prompt. The gate's limit is 0.5×.
overhead
+4.0%p95 token latency What running WarmSwap costs a normal request, throughput down 1.1%. Budget: 10%.
leak drill
1000cycles, no residue Migrate, cancel, repeat. No residual references, no drift in host or device memory.
native rate
16/16per evaluation cell Every migration completed natively, with zero positions recomputed.
protocol suite
31passed on real workers Every row of the failure matrix, fault-injected at each phase of the transaction.

Its default answer
is no.

A migration is only exact if both GPUs run the same execution plan. WarmSwap hashes that plan into a fingerprint and moves a session only when a certificate says that exact pair was measured and found exact — on this hardware, with this build.

GET /v1/plans — uncertified deployment
{
  "plans": [],
  "certification_refusals": [
    "evaluation did not issue cuda_release:b3:1:…:
     G4: failed — p95 latency +26% (max 10%)"
  ]
}

Serving carries on. Nothing migrates, and the reason is a sentence rather than a silent fallback.

what that buys
  • No silent “close enough”. A pair nobody measured cannot migrate.
  • Certificates expire; migration stops with them while serving continues.
  • A drift in driver, engine or model settings changes the fingerprint — and the system notices before a user does.
  • Certifying your own machine is an unattended two-hour job, and the first thing the quickstart walks you through.

Two workers,
one coordinator.

terminal — install
# build the image, pinned to the evaluated tuple
docker build -t warmswap:0.1.0 .

# one worker per GPU (--ipc=host: both map one shared arena)
docker run -d --gpus '"device=0"' --ipc=host warmswap:0.1.0 worker \
  --worker-id worker-a --socket-dir /run/warmswap/sockets \
  --adapter vllm --model Qwen/Qwen2.5-7B-Instruct

# the coordinator holds no GPU at all
docker run -d --ipc=host --network host warmswap:0.1.0 serve \
  --deployment /etc/warmswap/deployment.json

# move every session off worker-a, before a deadline
curl -XPOST localhost:8080/v1/workers/worker-a/drains \
  -d '{"target_worker_id":"worker-b","deadline":"…"}'

The quickstart goes from a bare two-GPU host to a session surviving a drain. The white paper is the evaluation: method, results, and what the gates caught.

scope
  • One host, any number of GPUs on it. One worker per card, every pair certified. Cross-host is designed, not built.
  • One certified tuple. Another model, GPU or engine version needs its own run.
  • Four sessions by default, one migration at a time. Raise the session limit deliberately; nothing above four is measured.
  • No user has run it yet. The customer gate is unrun, and every number here is our own measurement.
licence — Apache 2.0

the licence

  • Apache 2.0. Use it, modify it, ship it, build a service on it.
  • Patent grant included. No agreement to sign, nobody to ask.
  • Publish what you measure — including results that are unfavourable.

the certificate

  • An interlock, not a gate: it decides whether this machine may migrate.
  • You produce it, from your own evidence, on your own hardware.
  • Nothing is countersigned and nothing phones home.

producing one

  • An unattended run of about two hours against your GPUs.
  • Certificates name one deployment tuple and expire.
  • Until one passes, serving carries on and migration refuses.

A deployment that has drifted from the evidence justifying it is one the system refuses to migrate. That check is the product, and it is yours to run.