Move a running LLM generation from one GPU to another without losing it.
Stated plainly, because a system whose selling point is refusing to guess should not be vague about its own edges.
| One host, any number of GPUs on it | A coordinator takes as many workers as the host has GPUs — two, four, eight — and certifies every ordered pair between them. Moving a session between hosts is designed and not built: the export writes into a host-shared arena, and import installs state through in-process engine hooks. |
| One migration at a time | The shared transfer arena holds one snapshot, so max_concurrent_migrations is 1 and draining an eight-GPU host moves sessions serially. Lifting it needs per-migration arena regions, which is unbuilt. |
| One certified tuple at a time | A certificate covers exactly the hardware, driver, vLLM, model and settings it was measured on. Anything else needs its own evaluation run. |
| A small default envelope | Four active sessions service-wide (raise it with max_active_sessions in deployment.json), one coordinator, SQLite. Sized for a pilot, not a public service, and untested above that. |
| Greedy decoding | Stochastic sampling is out of scope for the evaluated build: exactness is the claim, and it is only meaningful for a deterministic decode path. |
| vLLM | The adapter is written and verified against vLLM 0.11.0 specifically, by walking its live object graph. Other engines need their own adapter and their own evidence. |
| No multi-tenancy beyond namespaces | Tokens map to namespaces and namespaces isolate sessions. There is no quota system, no billing, no per-tenant scheduling. |
| Not proven with users | Nobody has run this on their own workload. The customer gate (G7) is unrun, and no performance claim here comes from anything but our own measurements on our own hardware. |
deployment.json and another certification run. Adding a second
machine is the cross-host work that does not exist yet.