WarmSwap

Move a running LLM generation from one GPU to another without losing it.

View My GitHub Profile

What WarmSwap does not do

Stated plainly, because a system whose selling point is refusing to guess should not be vague about its own edges.

   
One host, any number of GPUs on it A coordinator takes as many workers as the host has GPUs — two, four, eight — and certifies every ordered pair between them. Moving a session between hosts is designed and not built: the export writes into a host-shared arena, and import installs state through in-process engine hooks.
One migration at a time The shared transfer arena holds one snapshot, so max_concurrent_migrations is 1 and draining an eight-GPU host moves sessions serially. Lifting it needs per-migration arena regions, which is unbuilt.
One certified tuple at a time A certificate covers exactly the hardware, driver, vLLM, model and settings it was measured on. Anything else needs its own evaluation run.
A small default envelope Four active sessions service-wide (raise it with max_active_sessions in deployment.json), one coordinator, SQLite. Sized for a pilot, not a public service, and untested above that.
Greedy decoding Stochastic sampling is out of scope for the evaluated build: exactness is the claim, and it is only meaningful for a deterministic decode path.
vLLM The adapter is written and verified against vLLM 0.11.0 specifically, by walking its live object graph. Other engines need their own adapter and their own evidence.
No multi-tenancy beyond namespaces Tokens map to namespaces and namespaces isolate sessions. There is no quota system, no billing, no per-tenant scheduling.
Not proven with users Nobody has run this on their own workload. The customer gate (G7) is unrun, and no performance claim here comes from anything but our own measurements on our own hardware.

Things that are true but easy to misread