Move a running LLM generation from one GPU to another without losing it.
Status: the deployment sequence below has been run end to end on two RTX 3090s with the pinned 7B — two worker processes, one per GPU, a coordinator connected to both, and a live generation drained between them through the HTTP API. Where a step has not been exercised on hardware, it says so.
| Capability | State |
|---|---|
| Session lifecycle, durable transcript, SSE replay | Implemented; exercised on the release pair |
| Generation | Implemented; one producer task per session, gated on a subscriber |
| Drains, admission closure, capacity refusal | Implemented; live sessions drained between GPUs on one host, four sessions at a time by default |
| Migration transaction, abort, post-commit recovery | Implemented; native handoff and abort both exercised |
| Native KV transfer against vLLM | Implemented; reuse 1.000000, recomputed 0 across 300+ migrations at 8k and 16k |
| Replay fallback | Implemented, default off; recomputation counters instrumented. Drilled both ways on the release tuple: with the owning worker killed, a session that permitted replay resumed on the survivor and ran to completion, and one that did not ended failed with recovery_unavailable |
| Restart reconciliation | Implemented; three crash points drilled on real engines, plus a worker loss resolved across two nonterminal sessions in one pass |
| Transfer through shared page-locked memory | Implemented, and required with vLLM workers: that adapter exports and imports only through the arena. The payload-socket path exists for the mock adapter; a vLLM worker without an arena refuses migration and says so |
Measured on the release tuple — two RTX 3090s, one worker process per GPU, one host, the pinned 7B in FP16, every drain requested through the operator API:
| Gate | Result |
|---|---|
| G1 protocol suite on the release tuple | passed |
| G3 frozen pause ≤ 0.5 × B2 | passed — worst p95 ratio 0.331 across all four required cells (8k and 16k, concurrency 1 and 4) |
| G2 native rate ≥ 99%, reuse ≥ 95% | passed — native rate and reuse both 1.000 across all four cells, every cohort past its 100-sample floor (105, 120, 105, 108 records) |
| G5 continuation and task quality | passed. Continuation: 20/20 identical over 256 tokens against a 20/20 control. Held-out tasks: degradation bound 0.0084, within 0.01. Migrated against uninterrupted logits at the same batch schedule (D18): bitwise identical at concurrency 4 on 80 distinct prompts, 3,020 positions, and across GPUs on 12. The served drain probe, whose reference could not share the candidate’s batch or prefix-cache state, read 4 of 162 over the per-logit tolerance; it is reported in the verdict, not used for it |
| G4 ordinary serving overhead | passes on the current build — 0.5% and 1.3% throughput loss and +2.7% and +5.6% p95 token latency at one session, 1.8% and 1.8% with +3.4% and +6.0% at four, against 10% budgets, on the second host (driver 570.211.01), prestep on and off interleaved. The build measured on the release pair, before the worker computed ahead and batched its sessions’ round trips, failed at 18.6% / +28.4% |
| G6 leak drill | passed on the release tuple — 1,000 alternating migrate/abort cycles across two 7B workers, one per GPU. Zero residual KV, sessions or descriptors; worker RSS drifted 53 and 27 MiB, reported rather than counted |
| G7 commercial | not run; needs an operator, not a benchmark |
| Certificate | not issued — experiments/evidence/certify.py writes evidence/T19/CERTIFICATE.json from the bundle and names G4 as blocking. Both workers’ fingerprints are recorded and identical |
| R01–R12 on hardware | all established on the release tuple. R01–R03, R06, R07, R09, R10 and R12 by requirements_drill.py, 52 of 52 checks against the served product; R04, R05, R08 and R11 by the G2/G3 cells and the release demo checklist |
Per-cell frozen pause, p95: 493 ms at 8k/c1, 597 ms at 8k/c4, 864 ms at 16k/c1, 1,121 ms at 16k/c4, against a replay baseline of 1.9 s and 4.1 s respectively.
vLLM’s own prefix cache beats this product when the target is warm. Measured at concurrency 1, replaying a session onto a target whose cache already holds the prompt:
| cell | native transfer | cold replay | warm-cache replay |
|---|---|---|---|
| N=8192 | 448 ms | 1,878 ms | 105 ms |
| N=15872 | 764 ms | 4,068 ms | 112 ms |
Warm-cache replay is nearly flat in context length, because the cache supplies the whole prompt whatever its size. So this build earns its place only on contexts the target cannot already have: long and unique. If your sessions share a long system prompt that the target sees often, prefix caching is the better choice at any context length and this is overhead. If your sessions carry long per-user context the other worker has never seen, cold replay costs 1.9 s where a transfer costs 448 ms.
Serving overhead is within G4’s budget on the current build, and it depends on two things
staying on. Measured with durability unchanged — a synchronous=FULL commit for every token —
the service costs about 1–2% of throughput and adds 3–6% to p95 token latency at one and four
sessions. That rests on the worker computing each session’s next token while the coordinator
makes the last one durable (--prestep, default on) and on a worker’s sessions sharing one
control round trip. With --prestep off the same build loses 15–25% of throughput and adds
45–65% to p95: the durable write is back in series with every engine step. Use off only to
diagnose, never to serve.
Storage still matters: the serial path’s cost follows the commit time, which was 3.0 ms per token on the release pair and 1.0 ms on the second host. Put the database on the fastest durable volume you have. And treat any single run as approximate: G4 here is two runs a cell plus a third in the bundle, interleaved with the other arm on one host.
The handoff itself does not change numerics; the runtime’s batching does. On the release pair a session resumed at an aligned boundary produces bitwise the same logits and KV as the same session left alone at the same batch schedule, on the same GPU or the other one. But vLLM’s own output depends on what else is on the worker: a prompt that hits the prefix cache, a prefill that shares a batch, or neighbours that arrive mid-decode each move logits by up to 0.18 nats, and changed greedy output on one prompt of four with no migration anywhere. A migrated session lands in a different batch, so it sees that variance — as does every session whose neighbours come and go. If something downstream consumes log-probabilities, measure its sensitivity to batching, not to migration (doc/internal/work-log/T19 §25).
Treat any single number here as approximate. The one-session cell on fast storage measured 6.0%, 6.6%, 11.9% and 12.2% across four runs of one build, so plan against the upper end of that range rather than the figure that happens to be quoted. Latency is the criterion that fails by the wider margin and is the one to design around: it is roughly twice its budget and no configuration measured here brings it inside.
Follows doc/internal/09 §”Deployment sequence”.
Install the runtime at exact versions, including the ones nothing pins for you. The release evidence was measured on:
| Python | 3.12.14 |
| CUDA | 12.8 |
| torch | 2.8.0+cu128 (pulled in by vLLM) |
| vLLM | 0.11.0 |
| transformers | 4.57.6 |
| model | Qwen/Qwen2.5-7B-Instruct @ a09a35458c702b33eeacc393d103063234e8bc28 |
| driver | 580.173.02 |
Install vllm==0.11.0 and then transformers==4.57.6, in that order. vLLM 0.11.0
declares only transformers>=4.55.2, so a resolver on a fresh image takes the newest
release — 5.x, which vLLM 0.11.0 predates, and whose tokenizer API it crashes on
(Qwen2Tokenizer has no attribute all_special_tokens_extended). This has happened
twice: on the first instance and again on a rebuilt one, where the image also arrived
with torch 2.11.0 and no vLLM. Check the versions after installing, not before.
One conflict is expected and harmless: the project’s FastAPI range holds starlette at
0.48, below what vLLM’s prometheus-fastapi-instrumentator wants. That package serves
vLLM’s own HTTP server, which this service does not run.
Verify the release tuple, and point the service at the certificate evaluation
issued. experiments/evidence/certify.py writes CERTIFICATE.json from the evidence
bundle. Set certificate_path to it in the deployment file, set certificate_id to the
id inside it, and start both workers with --certificate-id set to that same id — the
target refuses any snapshot whose manifest names a different certificate. At startup the
service registers an edge only for worker pairs whose live fingerprints are the ones the
certificate names, and only if the certificate was issued, is unexpired, and matches this
deployment’s evidence profile and snapshot schema. Every other case leaves native
migration disabled and says why in GET /v1/plans under certification_refusals;
serving continues. A configured certificate file that cannot be read stops startup.
Without certificate_path, native migration is disabled. self_certify: true makes the
service certify fingerprint-matched pairs itself, with no evaluation behind them — what
an evaluation run needs, because it has to migrate before any certificate exists, and
nothing else. Those plans show "basis": "self_issued" and startup warns about them.
StartupValidator still refuses a mock certificate outright.
flock on coordinator.lock beside
the database. A second process fails immediately with coordinator_lock_held. Run
exactly one uvicorn worker: two workers are two coordinators against one store.AF_UNIX caps it at 104 bytes on
macOS and 108 on Linux, and SocketDirectory.socket_for refuses a longer one.GET /v1/plans — each with "basis":
"evaluated" — with migration_enabled still false.One coordinator and one worker per GPU, all on one host. Two is the smallest deployment that can migrate anything; four or eight is the same arrangement with more workers in the deployment file. The workers are started first — the coordinator connects to the sockets they report — and each prints one JSON line when it is ready, carrying the incarnation and socket paths its entry in the deployment file needs.
CUDA_VISIBLE_DEVICES=0 VLLM_ENABLE_V1_MULTIPROCESSING=0 \
python -m warmswap.worker_control.worker_main \
--worker-id worker-a --socket-dir /run/warmswap/sockets \
--adapter vllm --model Qwen/Qwen2.5-7B-Instruct \
--revision a09a35458c702b33eeacc393d103063234e8bc28 \
--gpu-memory-utilization 0.88 --max-model-len 16384 \
--certificate-id "$CERTIFICATE_ID"
--certificate-id must be the id inside the certificate the coordinator loads: the target
refuses any snapshot whose manifest names a different one. --prestep defaults to on and
belongs on; off is a diagnostic (see G4 above). Leave --observed-logprobs unset in
production — it is evaluation instrumentation and costs engine time. Repeat per GPU:
worker-b on CUDA_VISIBLE_DEVICES=1, worker-c on 2, and so on. Every worker on the
host must be started with the same model, revision and settings — they are hashed into the
fingerprint, and a worker that differs is one the certificate does not cover.
Two things scale with the number of GPUs. max_active_sessions in the deployment file
is counted service-wide and defaults to four, so a host with more GPUs than that will refuse
admissions long before its GPUs are busy — raise it deliberately, and measure, because
nothing above four has been evaluated. And max_concurrent_migrations stays at one: the
shared transfer arena holds a single snapshot, so draining an eight-GPU host moves sessions
one after another rather than in parallel.
Give the coordinator CPU it does not have to queue for. It sits on the path of every
token — each grant and each accept goes through it — and it needs 2.5–3 ms of CPU a
token to do that. On T23 (a shared machine, ~45% of its 64 CPUs busy with other
tenants) it spent as long again runnable but waiting for a core, measured from
/proc/<pid>/schedstat: 1.09 ms waiting per 1 ms running. The median token was
unaffected and the p95 rose 26–31%, which failed G4; the same build had passed at +3–6%
on a host where it was not queued. Throttling, storage, CPU placement, GC and the code
were each ruled out first (doc/internal/work-log/T23). Reserve cores for the coordinator — a
dedicated cpuset, or Kubernetes’ static CPU manager with Guaranteed QoS — and run G4 on
the host that will serve, because its budget does not carry from one host to another.
A coordinator on a machine it shares with busy neighbours will serve correctly and
slowly; the certificate will refuse it, which is the point.
Set HF_HOME to a cache you have already populated and HF_HUB_OFFLINE=1 alongside it.
The weights are half the fingerprint; a worker that can reach the Hub at startup is a
worker whose weights can change without the certificate noticing, and one that cannot
reach it is a worker that will not start at all when the Hub is slow or blocked. Both
failures are avoidable by pinning the cache. Pre-populate it once per host and verify the
blobs (experiments/gpu/host/setup_model.sh does this against a recorded SHA-256 list).
The coordinator reads one file and nothing else:
{
"database_path": "/var/lib/warmswap/warmswap.db",
"socket_dir": "/run/warmswap/sockets",
"certificate_path": "/etc/warmswap/CERTIFICATE.json",
"certificate_id": "cuda_release:b3:1:9858f98f526",
"evidence_profile": "cuda_release",
"transfer_arena": true,
"host": "127.0.0.1",
"port": 8080,
"max_active_sessions": 8,
"workers": [
{"worker_id": "worker-a", "control_socket": "/run/warmswap/sockets/worker-a.sock",
"payload_socket": "/run/warmswap/sockets/worker-a.payload.sock", "incarnation": "inc_..."},
{"worker_id": "worker-b", "control_socket": "/run/warmswap/sockets/worker-b.sock",
"payload_socket": "/run/warmswap/sockets/worker-b.payload.sock", "incarnation": "inc_..."},
{"worker_id": "worker-c", "control_socket": "/run/warmswap/sockets/worker-c.sock",
"payload_socket": "/run/warmswap/sockets/worker-c.payload.sock", "incarnation": "inc_..."},
{"worker_id": "worker-d", "control_socket": "/run/warmswap/sockets/worker-d.sock",
"payload_socket": "/run/warmswap/sockets/worker-d.payload.sock", "incarnation": "inc_..."}
],
"credentials": [
{"token": "...", "namespace": "tenant-a"},
{"token": "...", "namespace": "operators", "is_operator": true}
],
"models": [
{"alias": "Qwen/Qwen2.5-7B-Instruct", "weights_digest": "Qwen/Qwen2.5-7B-Instruct@a09a3545...",
"tokenizer_digest": "Qwen/Qwen2.5-7B-Instruct@a09a3545...", "eos_token_ids": [151643, 151645]}
],
"tokenizer": {"directory": "/var/lib/warmswap/tokenizer", "model_id": "Qwen/Qwen2.5-7B-Instruct",
"revision": "a09a35458c702b33eeacc393d103063234e8bc28"}
}
python -m warmswap serve --config /etc/warmswap/deployment.json
certificate_path is what makes native migration available at all, and certificate_id
must match the id inside it (see step 1). self_certify: true exists for evaluation runs,
which have to migrate before any certificate can exist; a serving deployment leaves it out.
allow_mock_certificates is refused outright in production. Run one coordinator process:
two are two coordinators against one store, which is why the service takes an exclusive lock
on the database and workers=1 is not configurable.
A unit apiece, workers before the coordinator:
# /etc/systemd/system/warmswap-worker@.service
[Unit]
Description=warmswap worker %i
After=network.target
[Service]
Environment=VLLM_ENABLE_V1_MULTIPROCESSING=0
EnvironmentFile=/etc/warmswap/worker-%i.env
ExecStart=/opt/warmswap/.venv/bin/python -m warmswap.worker_control.worker_main $WORKER_ARGS
Restart=on-failure
[Install]
WantedBy=multi-user.target
# /etc/systemd/system/warmswap.service
[Unit]
Description=warmswap coordinator
After=warmswap-worker@a.service warmswap-worker@b.service
Requires=warmswap-worker@a.service warmswap-worker@b.service
[Service]
ExecStart=/opt/warmswap/.venv/bin/python -m warmswap serve --config /etc/warmswap/deployment.json
Restart=on-failure
TimeoutStopSec=120
[Install]
WantedBy=multi-user.target
A restarted worker comes up with a new incarnation, which fences everything the old one
owned, so the deployment file’s incarnation entries have to be updated from its new ready
line before the coordinator is started against it — an incarnation that does not match is
refused at connect, deliberately.
What a restart costs. Accepted tokens are durable, so a coordinator restart loses no
output: reconciliation resumes the producers for sessions that were still running, and a
client reconnects to /v1/sessions/{id}/events with Last-Event-ID and continues from its
cursor. What breaks is the stream itself, and any migration in flight resolves through the
failure matrix rather than continuing. Before a planned restart, follow the rollback order in
step 1 of doc/internal/09: stop admissions, drain each worker, then stop the service.
| Check | How | Expected |
|---|---|---|
| Short session | POST /v1/sessions, subscribe to events |
created then token events; terminal completed |
| Reconnect | Disconnect mid-stream, reconnect with Last-Event-ID |
Events strictly after the cursor, no gap |
| Native drain | POST /v1/workers/{a}/drains |
drained, each session completed_native |
| Aborted import | Same, with the target’s import path failing | aborted_source_resumed, source still owns the session |
| Safe shutdown | Stop the coordinator, restart | Reconciliation resolves every nonterminal session; history identical |
There are no automatic backups (doc/internal/09), and restart durability is not backup recovery: the store survives a restart because it is fsynced, not because a copy exists somewhere. If you want one, take it while the service runs with a consistent read rather than by copying files:
sqlite3 /var/lib/warmswap/warmswap.db "VACUUM INTO '/backups/warmswap-$(date -u +%FT%H%M).db'"
Never copy warmswap.db alongside its -wal and -shm by hand while the service is
writing: the three move independently and the result is a store nobody can vouch for.
A backup contains transcripts — prompts and generated tokens, per tenant. It inherits every obligation the live store has: the same encryption at rest, the same access control, and a retention policy of its own, because a copy does not expire when retention sweeps the original (doc/internal/09’s retention applies to the service, not to your backup tool).
To restore: stop the service, put the file in place with no -wal or -shm beside it from
a different generation, and start again. Reconciliation runs before readiness and resolves
what it finds. Sessions whose owning worker incarnation no longer exists are not resumed —
each is replayed where the session permitted replay, and otherwise ended explicitly — so a
restored store returns a consistent service, not the sessions that were live when it was taken.
A certificate is valid for 30 days from issue. When it lapses, native migration stops and
serving carries on: drains report blocked, and GET /v1/plans shows the expiry that
passed. Nothing else announces it, so watch warmswap_certificate_expires_seconds — it is
read when you scrape it, counts down, goes negative once expired, and reads 0 when no
certificate is registered at all. Alert below seven days. Startup warns inside the same
window, and again when every certificate has already expired.
Renewing means re-running evaluation on the build and tuple you serve, and pointing
certificate_path at the certificate it issues: a certificate is evidence with a date on it,
not a licence to extend.
Taken from doc/internal/09 and made concrete against this implementation.
| Situation | Response |
|---|---|
| Native transfer failed, source healthy | The source resumes automatically and the migration records aborted_source_resumed. Inspect GET /v1/drains/{id} for the reason. Keep the failed target quarantined until verified — abort_prepare has already discarded its import, but confirm warmswap_open_reservations returned to its baseline. |
| Drain blocked at the deadline | The source stays closed to new admissions and its remaining sessions keep running. The deadline never terminates a session. Inspect the per-session reasons; either wait, or cancel the drain and start a new one with a later deadline. |
| Source dies | Sessions on it become unprovable at the next reconciliation. With allow_replay_fallback=true they replay onto the other worker; otherwise they terminate with recovery_unavailable and the last accepted cursor. Do not describe either as recovered KV. |
| Coordinator restart | Readiness stays false until reconciliation finishes. Clients reconnect by cursor. Check the reconciliation report for sessions resolved as failed or replayed. |
| Disk full or fsync error | The store stops acknowledging output; readiness fails. Restore storage headroom, restart, and let reconciliation run before serving. Do not delete WAL files by hand. |
| Suspected state corruption | Revoke the affected certificate (PlanRegistry.revoke) — every edge it covers is denied immediately. Set migration_enabled=false. Existing sessions continue on their current worker. |
| Rollback | Stop admissions, drain safely, stop the service, deploy the prior build with a compatible schema, rerun the smoke checks. A runtime rollback changes the fingerprint and needs its own still-valid certificate. |
| Limit | Value | Where |
|---|---|---|
| Prompt + reserved completion | 16,384 tokens | SessionService._validate_budget |
max_new_tokens |
1–512 | request schema |
| Active sessions | 4 | AdmissionService.may_admit_new_session |
| Concurrent migrations | 1 | coordinator semaphore |
| Request body | 1 MiB | Limits.max_request_body_bytes |
| Control frame | 1 MiB | worker_control.messages.MAX_FRAME_BYTES |
| Chunk frame | 8 MiB | transport.framing.MAX_CHUNK_FRAME_BYTES |
| Transfer staging | 256 MiB | Limits.max_staging_bytes |
| Shared transfer arena | one worker’s longest context, page-locked once at deployment | TransferArena; GET /v1/plans reports its size and whether the lock was taken |
| Snapshot payload | 8 GiB hard ceiling | state.boundary.MAX_SNAPSHOT_BYTES |
| Active session lifetime | 10 minutes | Timeouts.active_session_lifetime |
| Disconnect grace | 30 seconds | Timeouts.disconnect_grace |
| Terminal retention | 1 hour | Timeouts.terminal_retention |
GET /v1/plans answers two questions an operator needs before a drain: which edges are
certified, and how state crosses them. transfer.kind is arena when both workers share one
page-locked buffer and payload_socket when the bytes are copied worker to coordinator to
worker; transfer.page_locked false on an arena means transfers run at roughly half speed and
the host could not take the lock.
GET /metrics requires an operator credential. Every metric label is drawn from a closed
set (outcome, phase, a stable reason code) — session and migration ids never appear in
labels.
Logs are JSON, one object per line, through an allowlist: a field not on the list is dropped and counted. Prompts, token ids, rendered text and tensor payloads can never reach a log line.
Once ownership has moved to the target, there is no rollback. If the target then fails to activate the session, the service now decides immediately — the same decision a coordinator restart would make:
allow_replay_fallback: true): the session is rebuilt from its accepted
tokens on the other worker and keeps generating; the migration reports completed_replay;failed with code recovery_unavailable, and the
migration reports failed.Before this build neither happened: the session sat in recovering with its stream stalled
until the coordinator restarted. If you see a session in recovering for more than a few
seconds on this build, that is a fault to report, not a state to wait out. A failure after
the target was already running the session is treated as the successful handoff it was.
Earlier builds could report target already holds the maximum active sessions for a few
seconds after sessions were cancelled, naming sessions that had already ended. A cancel that
landed on a control connection closed moments earlier was lost, and the slot came back only when
the 5-second sweep reclaimed it. Cancels are now retried across the reconnect. If you still see
that refusal naming sessions the API reports as terminal, it is a regression.
Before freezing a session, the coordinator advances it to a boundary the target can restore whole. Only complete allocator blocks carry transferred state, so a session frozen mid-block leaves a tail for the target to recompute — and a recomputed tail changes the continuation.
Two things follow for you. The migration generates up to block_size - 1 extra tokens before
the freeze: they are real output the session wanted anyway, they count against its
max_new_tokens budget, and they land before the frozen phase, so they do not widen the
pause G3 measures. And a session close to its budget may finish during that advance — that is a
natural completion, reported as cancelled, not a fault.
Alignment is best effort and never fails a migration. A session that cannot reach an aligned boundary migrates from where it is, exactly as it did before.
A drain reports one outcome per session, and two of them are not faults. cancelled with
reason request has already finished generating means the request ended between the drain’s
precheck and the freeze — the session completes where it is, and nothing was lost. Until this
build that case read failed; if you are comparing against older drain reports, expect the
failed count to drop and the cancelled count to rise by the same amount, with no change in
behaviour. blocked with session has not reached an eligible boundary means the request is
still in prefill: it is worth retrying the drain, and the other two are not.
A boundary refusal now names its state — unknown to this worker, already finished, still
in prefill, no accepted tokens. Only the last two are worth waiting on.
A drain waits for a session whose migration came back cancelled, up to the drain’s deadline,
and reports drained once it has completed. Earlier builds finalized straight away: the drain
read blocked and the worker stayed closed to admission, although the session finished a moment
later, and a blocked drain is never re-evaluated. On the release pair that happened on every
drain whose session stopped just after the precheck. If you are still running such a build and see
a blocked drain whose only blocker is cancelled, cancel the drain to reopen the worker.
experiments/gpu/gpu_watch.py is a watchdog that runs on the GPU host itself, under
supervisord, so it outlives any operator’s session and is restarted if it dies. It uses the
standard library and the system interpreter, so a broken experiment environment cannot take
it down.
# /etc/supervisor/conf.d/gpu_watch.conf
[program:gpu_watch]
command=/usr/bin/python3 /workspace/ws2/gpu_watch.py --root /workspace/ws2 --out /workspace/ws2/watch
autostart=true
autorestart=true
startretries=1000
stdout_logfile=/workspace/ws2/watch/supervisor_stdout.txt
stderr_logfile=/workspace/ws2/watch/supervisor_stderr.txt
It keeps two files: watch/events.jsonl, one JSON object per alert or notable event, and
watch/status.json, the latest snapshot. Tail the first; read the second. What pages, and why
each rule is shaped the way it is:
| Alert | Fires when |
|---|---|
LOG_FAILURE |
a new line in any *.log shows a Traceback, OOM, kill, nonzero exit, position or transcript mismatch, or a failed drill check ([FAIL]). History present at startup never alerts, and a line split across two writes is still one match |
STALL |
an experiment is alive, every GPU is idle, and no log has grown for 7 minutes — measured on log progress, because a worker loading weights sits at 0% utilisation for about a minute and a half |
GPU_QUERY_FAILED |
nvidia-smi fails two ticks running |
GPU_HOT / GPU_THROTTLE |
85 °C, or a hardware-slowdown, thermal or power-brake throttle reason starts |
RESOURCE_LOW |
disk, /dev/shm (the transfer arena) or available RAM crosses its floor |
CANARY_FAILED |
after 30 idle minutes, each GPU multiplies two float64 matrices and the result disagrees with the CPU. It catches a card that answers nvidia-smi but computes wrongly |
Every alert fires when its condition starts, not on every tick it persists. Jobs are identified by script name only: command lines on these hosts can carry secrets, including other services’ access tokens, and none of that is ever copied into a record.
Verified on the release pair: an injected Traceback alerted within one tick; a silent fake
job raised one STALL; kill -9 of the watchdog was restarted by supervisord without
re-alerting on old failures; and the canary matched the CPU to 1.8e-15 on both GPUs.
These are properties of the design, not defects:
/dev/shm. A container
started with Docker’s default --shm-size of 64 MiB cannot hold it; the deployment says so
in its log and migrates over the payload socket instead, which works everywhere and is
slower. Give the container --shm-size at least the value GET /v1/plans reports.