WarmSwap

Move a running LLM generation from one GPU to another without losing it.

View My GitHub Profile

Operating warmswap

Status: the deployment sequence below has been run end to end on two RTX 3090s with the pinned 7B — two worker processes, one per GPU, a coordinator connected to both, and a live generation drained between them through the HTTP API. Where a step has not been exercised on hardware, it says so.

What this build can and cannot do

Capability State
Session lifecycle, durable transcript, SSE replay Implemented; exercised on the release pair
Generation Implemented; one producer task per session, gated on a subscriber
Drains, admission closure, capacity refusal Implemented; live sessions drained between GPUs on one host, four sessions at a time by default
Migration transaction, abort, post-commit recovery Implemented; native handoff and abort both exercised
Native KV transfer against vLLM Implemented; reuse 1.000000, recomputed 0 across 300+ migrations at 8k and 16k
Replay fallback Implemented, default off; recomputation counters instrumented. Drilled both ways on the release tuple: with the owning worker killed, a session that permitted replay resumed on the survivor and ran to completion, and one that did not ended failed with recovery_unavailable
Restart reconciliation Implemented; three crash points drilled on real engines, plus a worker loss resolved across two nonterminal sessions in one pass
Transfer through shared page-locked memory Implemented, and required with vLLM workers: that adapter exports and imports only through the arena. The payload-socket path exists for the mock adapter; a vLLM worker without an arena refuses migration and says so

What is established, and at what grade

Measured on the release tuple — two RTX 3090s, one worker process per GPU, one host, the pinned 7B in FP16, every drain requested through the operator API:

Gate Result
G1 protocol suite on the release tuple passed
G3 frozen pause ≤ 0.5 × B2 passed — worst p95 ratio 0.331 across all four required cells (8k and 16k, concurrency 1 and 4)
G2 native rate ≥ 99%, reuse ≥ 95% passed — native rate and reuse both 1.000 across all four cells, every cohort past its 100-sample floor (105, 120, 105, 108 records)
G5 continuation and task quality passed. Continuation: 20/20 identical over 256 tokens against a 20/20 control. Held-out tasks: degradation bound 0.0084, within 0.01. Migrated against uninterrupted logits at the same batch schedule (D18): bitwise identical at concurrency 4 on 80 distinct prompts, 3,020 positions, and across GPUs on 12. The served drain probe, whose reference could not share the candidate’s batch or prefix-cache state, read 4 of 162 over the per-logit tolerance; it is reported in the verdict, not used for it
G4 ordinary serving overhead passes on the current build — 0.5% and 1.3% throughput loss and +2.7% and +5.6% p95 token latency at one session, 1.8% and 1.8% with +3.4% and +6.0% at four, against 10% budgets, on the second host (driver 570.211.01), prestep on and off interleaved. The build measured on the release pair, before the worker computed ahead and batched its sessions’ round trips, failed at 18.6% / +28.4%
G6 leak drill passed on the release tuple — 1,000 alternating migrate/abort cycles across two 7B workers, one per GPU. Zero residual KV, sessions or descriptors; worker RSS drifted 53 and 27 MiB, reported rather than counted
G7 commercial not run; needs an operator, not a benchmark
Certificate not issued — experiments/evidence/certify.py writes evidence/T19/CERTIFICATE.json from the bundle and names G4 as blocking. Both workers’ fingerprints are recorded and identical
R01–R12 on hardware all established on the release tuple. R01–R03, R06, R07, R09, R10 and R12 by requirements_drill.py, 52 of 52 checks against the served product; R04, R05, R08 and R11 by the G2/G3 cells and the release demo checklist

Per-cell frozen pause, p95: 493 ms at 8k/c1, 597 ms at 8k/c4, 864 ms at 16k/c1, 1,121 ms at 16k/c4, against a replay baseline of 1.9 s and 4.1 s respectively.

Before you deploy this, read this

vLLM’s own prefix cache beats this product when the target is warm. Measured at concurrency 1, replaying a session onto a target whose cache already holds the prompt:

cell native transfer cold replay warm-cache replay
N=8192 448 ms 1,878 ms 105 ms
N=15872 764 ms 4,068 ms 112 ms

Warm-cache replay is nearly flat in context length, because the cache supplies the whole prompt whatever its size. So this build earns its place only on contexts the target cannot already have: long and unique. If your sessions share a long system prompt that the target sees often, prefix caching is the better choice at any context length and this is overhead. If your sessions carry long per-user context the other worker has never seen, cold replay costs 1.9 s where a transfer costs 448 ms.

Serving overhead is within G4’s budget on the current build, and it depends on two things staying on. Measured with durability unchanged — a synchronous=FULL commit for every token — the service costs about 1–2% of throughput and adds 3–6% to p95 token latency at one and four sessions. That rests on the worker computing each session’s next token while the coordinator makes the last one durable (--prestep, default on) and on a worker’s sessions sharing one control round trip. With --prestep off the same build loses 15–25% of throughput and adds 45–65% to p95: the durable write is back in series with every engine step. Use off only to diagnose, never to serve.

Storage still matters: the serial path’s cost follows the commit time, which was 3.0 ms per token on the release pair and 1.0 ms on the second host. Put the database on the fastest durable volume you have. And treat any single run as approximate: G4 here is two runs a cell plus a third in the bundle, interleaved with the other arm on one host.

The handoff itself does not change numerics; the runtime’s batching does. On the release pair a session resumed at an aligned boundary produces bitwise the same logits and KV as the same session left alone at the same batch schedule, on the same GPU or the other one. But vLLM’s own output depends on what else is on the worker: a prompt that hits the prefix cache, a prefill that shares a batch, or neighbours that arrive mid-decode each move logits by up to 0.18 nats, and changed greedy output on one prompt of four with no migration anywhere. A migrated session lands in a different batch, so it sees that variance — as does every session whose neighbours come and go. If something downstream consumes log-probabilities, measure its sensitivity to batching, not to migration (doc/internal/work-log/T19 §25).

Treat any single number here as approximate. The one-session cell on fast storage measured 6.0%, 6.6%, 11.9% and 12.2% across four runs of one build, so plan against the upper end of that range rather than the figure that happens to be quoted. Latency is the criterion that fails by the wider margin and is the one to design around: it is roughly twice its budget and no configuration measured here brings it inside.

Deployment sequence

Follows doc/internal/09 §”Deployment sequence”.

  1. Install the runtime at exact versions, including the ones nothing pins for you. The release evidence was measured on:

       
    Python 3.12.14
    CUDA 12.8
    torch 2.8.0+cu128 (pulled in by vLLM)
    vLLM 0.11.0
    transformers 4.57.6
    model Qwen/Qwen2.5-7B-Instruct @ a09a35458c702b33eeacc393d103063234e8bc28
    driver 580.173.02

    Install vllm==0.11.0 and then transformers==4.57.6, in that order. vLLM 0.11.0 declares only transformers>=4.55.2, so a resolver on a fresh image takes the newest release — 5.x, which vLLM 0.11.0 predates, and whose tokenizer API it crashes on (Qwen2Tokenizer has no attribute all_special_tokens_extended). This has happened twice: on the first instance and again on a rebuilt one, where the image also arrived with torch 2.11.0 and no vLLM. Check the versions after installing, not before.

    One conflict is expected and harmless: the project’s FastAPI range holds starlette at 0.48, below what vLLM’s prometheus-fastapi-instrumentator wants. That package serves vLLM’s own HTTP server, which this service does not run.

  2. Verify the release tuple, and point the service at the certificate evaluation issued. experiments/evidence/certify.py writes CERTIFICATE.json from the evidence bundle. Set certificate_path to it in the deployment file, set certificate_id to the id inside it, and start both workers with --certificate-id set to that same id — the target refuses any snapshot whose manifest names a different certificate. At startup the service registers an edge only for worker pairs whose live fingerprints are the ones the certificate names, and only if the certificate was issued, is unexpired, and matches this deployment’s evidence profile and snapshot schema. Every other case leaves native migration disabled and says why in GET /v1/plans under certification_refusals; serving continues. A configured certificate file that cannot be read stops startup.

    Without certificate_path, native migration is disabled. self_certify: true makes the service certify fingerprint-matched pairs itself, with no evaluation behind them — what an evaluation run needs, because it has to migrate before any certificate exists, and nothing else. Those plans show "basis": "self_issued" and startup warns about them. StartupValidator still refuses a mock certificate outright.

  3. Start the coordinator. It acquires an exclusive flock on coordinator.lock beside the database. A second process fails immediately with coordinator_lock_held. Run exactly one uvicorn worker: two workers are two coordinators against one store.
  4. Start both workers with distinct incarnations and the same execution fingerprint. Each listens on a private socket in a 0700 directory; the control token lives in a 0600 file there. Keep the socket directory path short — AF_UNIX caps it at 104 bytes on macOS and 108 on Linux, and SocketDirectory.socket_for refuses a longer one.
  5. Confirm the certified A→B and B→A edges in GET /v1/plans — each with "basis": "evaluated" — with migration_enabled still false.
  6. Run the smoke checks below.
  7. Enable migration for allowlisted applications.

Running it

One coordinator and one worker per GPU, all on one host. Two is the smallest deployment that can migrate anything; four or eight is the same arrangement with more workers in the deployment file. The workers are started first — the coordinator connects to the sockets they report — and each prints one JSON line when it is ready, carrying the incarnation and socket paths its entry in the deployment file needs.

CUDA_VISIBLE_DEVICES=0 VLLM_ENABLE_V1_MULTIPROCESSING=0 \
  python -m warmswap.worker_control.worker_main \
  --worker-id worker-a --socket-dir /run/warmswap/sockets \
  --adapter vllm --model Qwen/Qwen2.5-7B-Instruct \
  --revision a09a35458c702b33eeacc393d103063234e8bc28 \
  --gpu-memory-utilization 0.88 --max-model-len 16384 \
  --certificate-id "$CERTIFICATE_ID"

--certificate-id must be the id inside the certificate the coordinator loads: the target refuses any snapshot whose manifest names a different one. --prestep defaults to on and belongs on; off is a diagnostic (see G4 above). Leave --observed-logprobs unset in production — it is evaluation instrumentation and costs engine time. Repeat per GPU: worker-b on CUDA_VISIBLE_DEVICES=1, worker-c on 2, and so on. Every worker on the host must be started with the same model, revision and settings — they are hashed into the fingerprint, and a worker that differs is one the certificate does not cover.

Two things scale with the number of GPUs. max_active_sessions in the deployment file is counted service-wide and defaults to four, so a host with more GPUs than that will refuse admissions long before its GPUs are busy — raise it deliberately, and measure, because nothing above four has been evaluated. And max_concurrent_migrations stays at one: the shared transfer arena holds a single snapshot, so draining an eight-GPU host moves sessions one after another rather than in parallel.

Give the coordinator CPU it does not have to queue for. It sits on the path of every token — each grant and each accept goes through it — and it needs 2.5–3 ms of CPU a token to do that. On T23 (a shared machine, ~45% of its 64 CPUs busy with other tenants) it spent as long again runnable but waiting for a core, measured from /proc/<pid>/schedstat: 1.09 ms waiting per 1 ms running. The median token was unaffected and the p95 rose 26–31%, which failed G4; the same build had passed at +3–6% on a host where it was not queued. Throttling, storage, CPU placement, GC and the code were each ruled out first (doc/internal/work-log/T23). Reserve cores for the coordinator — a dedicated cpuset, or Kubernetes’ static CPU manager with Guaranteed QoS — and run G4 on the host that will serve, because its budget does not carry from one host to another. A coordinator on a machine it shares with busy neighbours will serve correctly and slowly; the certificate will refuse it, which is the point.

Set HF_HOME to a cache you have already populated and HF_HUB_OFFLINE=1 alongside it. The weights are half the fingerprint; a worker that can reach the Hub at startup is a worker whose weights can change without the certificate noticing, and one that cannot reach it is a worker that will not start at all when the Hub is slow or blocked. Both failures are avoidable by pinning the cache. Pre-populate it once per host and verify the blobs (experiments/gpu/host/setup_model.sh does this against a recorded SHA-256 list).

The coordinator reads one file and nothing else:

{
  "database_path": "/var/lib/warmswap/warmswap.db",
  "socket_dir": "/run/warmswap/sockets",
  "certificate_path": "/etc/warmswap/CERTIFICATE.json",
  "certificate_id": "cuda_release:b3:1:9858f98f526",
  "evidence_profile": "cuda_release",
  "transfer_arena": true,
  "host": "127.0.0.1",
  "port": 8080,
  "max_active_sessions": 8,
  "workers": [
    {"worker_id": "worker-a", "control_socket": "/run/warmswap/sockets/worker-a.sock",
     "payload_socket": "/run/warmswap/sockets/worker-a.payload.sock", "incarnation": "inc_..."},
    {"worker_id": "worker-b", "control_socket": "/run/warmswap/sockets/worker-b.sock",
     "payload_socket": "/run/warmswap/sockets/worker-b.payload.sock", "incarnation": "inc_..."},
    {"worker_id": "worker-c", "control_socket": "/run/warmswap/sockets/worker-c.sock",
     "payload_socket": "/run/warmswap/sockets/worker-c.payload.sock", "incarnation": "inc_..."},
    {"worker_id": "worker-d", "control_socket": "/run/warmswap/sockets/worker-d.sock",
     "payload_socket": "/run/warmswap/sockets/worker-d.payload.sock", "incarnation": "inc_..."}
  ],
  "credentials": [
    {"token": "...", "namespace": "tenant-a"},
    {"token": "...", "namespace": "operators", "is_operator": true}
  ],
  "models": [
    {"alias": "Qwen/Qwen2.5-7B-Instruct", "weights_digest": "Qwen/Qwen2.5-7B-Instruct@a09a3545...",
     "tokenizer_digest": "Qwen/Qwen2.5-7B-Instruct@a09a3545...", "eos_token_ids": [151643, 151645]}
  ],
  "tokenizer": {"directory": "/var/lib/warmswap/tokenizer", "model_id": "Qwen/Qwen2.5-7B-Instruct",
                "revision": "a09a35458c702b33eeacc393d103063234e8bc28"}
}
python -m warmswap serve --config /etc/warmswap/deployment.json

certificate_path is what makes native migration available at all, and certificate_id must match the id inside it (see step 1). self_certify: true exists for evaluation runs, which have to migrate before any certificate can exist; a serving deployment leaves it out. allow_mock_certificates is refused outright in production. Run one coordinator process: two are two coordinators against one store, which is why the service takes an exclusive lock on the database and workers=1 is not configurable.

A unit apiece, workers before the coordinator:

# /etc/systemd/system/warmswap-worker@.service
[Unit]
Description=warmswap worker %i
After=network.target
[Service]
Environment=VLLM_ENABLE_V1_MULTIPROCESSING=0
EnvironmentFile=/etc/warmswap/worker-%i.env
ExecStart=/opt/warmswap/.venv/bin/python -m warmswap.worker_control.worker_main $WORKER_ARGS
Restart=on-failure
[Install]
WantedBy=multi-user.target

# /etc/systemd/system/warmswap.service
[Unit]
Description=warmswap coordinator
After=warmswap-worker@a.service warmswap-worker@b.service
Requires=warmswap-worker@a.service warmswap-worker@b.service
[Service]
ExecStart=/opt/warmswap/.venv/bin/python -m warmswap serve --config /etc/warmswap/deployment.json
Restart=on-failure
TimeoutStopSec=120
[Install]
WantedBy=multi-user.target

A restarted worker comes up with a new incarnation, which fences everything the old one owned, so the deployment file’s incarnation entries have to be updated from its new ready line before the coordinator is started against it — an incarnation that does not match is refused at connect, deliberately.

What a restart costs. Accepted tokens are durable, so a coordinator restart loses no output: reconciliation resumes the producers for sessions that were still running, and a client reconnects to /v1/sessions/{id}/events with Last-Event-ID and continues from its cursor. What breaks is the stream itself, and any migration in flight resolves through the failure matrix rather than continuing. Before a planned restart, follow the rollback order in step 1 of doc/internal/09: stop admissions, drain each worker, then stop the service.

Smoke checks before enabling migration

Check How Expected
Short session POST /v1/sessions, subscribe to events created then token events; terminal completed
Reconnect Disconnect mid-stream, reconnect with Last-Event-ID Events strictly after the cursor, no gap
Native drain POST /v1/workers/{a}/drains drained, each session completed_native
Aborted import Same, with the target’s import path failing aborted_source_resumed, source still owns the session
Safe shutdown Stop the coordinator, restart Reconciliation resolves every nonterminal session; history identical

Backing up the store

There are no automatic backups (doc/internal/09), and restart durability is not backup recovery: the store survives a restart because it is fsynced, not because a copy exists somewhere. If you want one, take it while the service runs with a consistent read rather than by copying files:

sqlite3 /var/lib/warmswap/warmswap.db "VACUUM INTO '/backups/warmswap-$(date -u +%FT%H%M).db'"

Never copy warmswap.db alongside its -wal and -shm by hand while the service is writing: the three move independently and the result is a store nobody can vouch for.

A backup contains transcripts — prompts and generated tokens, per tenant. It inherits every obligation the live store has: the same encryption at rest, the same access control, and a retention policy of its own, because a copy does not expire when retention sweeps the original (doc/internal/09’s retention applies to the service, not to your backup tool).

To restore: stop the service, put the file in place with no -wal or -shm beside it from a different generation, and start again. Reconciliation runs before readiness and resolves what it finds. Sessions whose owning worker incarnation no longer exists are not resumed — each is replayed where the session permitted replay, and otherwise ended explicitly — so a restored store returns a consistent service, not the sessions that were live when it was taken.

The certificate expires

A certificate is valid for 30 days from issue. When it lapses, native migration stops and serving carries on: drains report blocked, and GET /v1/plans shows the expiry that passed. Nothing else announces it, so watch warmswap_certificate_expires_seconds — it is read when you scrape it, counts down, goes negative once expired, and reads 0 when no certificate is registered at all. Alert below seven days. Startup warns inside the same window, and again when every certificate has already expired.

Renewing means re-running evaluation on the build and tuple you serve, and pointing certificate_path at the certificate it issues: a certificate is evidence with a date on it, not a licence to extend.

Runbook

Taken from doc/internal/09 and made concrete against this implementation.

Situation Response
Native transfer failed, source healthy The source resumes automatically and the migration records aborted_source_resumed. Inspect GET /v1/drains/{id} for the reason. Keep the failed target quarantined until verified — abort_prepare has already discarded its import, but confirm warmswap_open_reservations returned to its baseline.
Drain blocked at the deadline The source stays closed to new admissions and its remaining sessions keep running. The deadline never terminates a session. Inspect the per-session reasons; either wait, or cancel the drain and start a new one with a later deadline.
Source dies Sessions on it become unprovable at the next reconciliation. With allow_replay_fallback=true they replay onto the other worker; otherwise they terminate with recovery_unavailable and the last accepted cursor. Do not describe either as recovered KV.
Coordinator restart Readiness stays false until reconciliation finishes. Clients reconnect by cursor. Check the reconciliation report for sessions resolved as failed or replayed.
Disk full or fsync error The store stops acknowledging output; readiness fails. Restore storage headroom, restart, and let reconciliation run before serving. Do not delete WAL files by hand.
Suspected state corruption Revoke the affected certificate (PlanRegistry.revoke) — every edge it covers is denied immediately. Set migration_enabled=false. Existing sessions continue on their current worker.
Rollback Stop admissions, drain safely, stop the service, deploy the prior build with a compatible schema, rerun the smoke checks. A runtime rollback changes the fingerprint and needs its own still-valid certificate.

Limits enforced by this build

Limit Value Where
Prompt + reserved completion 16,384 tokens SessionService._validate_budget
max_new_tokens 1–512 request schema
Active sessions 4 AdmissionService.may_admit_new_session
Concurrent migrations 1 coordinator semaphore
Request body 1 MiB Limits.max_request_body_bytes
Control frame 1 MiB worker_control.messages.MAX_FRAME_BYTES
Chunk frame 8 MiB transport.framing.MAX_CHUNK_FRAME_BYTES
Transfer staging 256 MiB Limits.max_staging_bytes
Shared transfer arena one worker’s longest context, page-locked once at deployment TransferArena; GET /v1/plans reports its size and whether the lock was taken
Snapshot payload 8 GiB hard ceiling state.boundary.MAX_SNAPSHOT_BYTES
Active session lifetime 10 minutes Timeouts.active_session_lifetime
Disconnect grace 30 seconds Timeouts.disconnect_grace
Terminal retention 1 hour Timeouts.terminal_retention

Observability

GET /v1/plans answers two questions an operator needs before a drain: which edges are certified, and how state crosses them. transfer.kind is arena when both workers share one page-locked buffer and payload_socket when the bytes are copied worker to coordinator to worker; transfer.page_locked false on an arena means transfers run at roughly half speed and the host could not take the lock.

GET /metrics requires an operator credential. Every metric label is drawn from a closed set (outcome, phase, a stable reason code) — session and migration ids never appear in labels.

Logs are JSON, one object per line, through an allowlist: a field not on the list is dropped and counted. Prompts, token ids, rendered text and tensor payloads can never reach a log line.

When a handoff fails after it has committed

Once ownership has moved to the target, there is no rollback. If the target then fails to activate the session, the service now decides immediately — the same decision a coordinator restart would make:

Before this build neither happened: the session sat in recovering with its stream stalled until the coordinator restarted. If you see a session in recovering for more than a few seconds on this build, that is a fault to report, not a state to wait out. A failure after the target was already running the session is treated as the successful handoff it was.

Capacity refusals right after a cancel

Earlier builds could report target already holds the maximum active sessions for a few seconds after sessions were cancelled, naming sessions that had already ended. A cancel that landed on a control connection closed moments earlier was lost, and the slot came back only when the 5-second sweep reclaimed it. Cancels are now retried across the reconnect. If you still see that refusal naming sessions the API reports as terminal, it is a regression.

Boundary alignment

Before freezing a session, the coordinator advances it to a boundary the target can restore whole. Only complete allocator blocks carry transferred state, so a session frozen mid-block leaves a tail for the target to recompute — and a recomputed tail changes the continuation.

Two things follow for you. The migration generates up to block_size - 1 extra tokens before the freeze: they are real output the session wanted anyway, they count against its max_new_tokens budget, and they land before the frozen phase, so they do not widen the pause G3 measures. And a session close to its budget may finish during that advance — that is a natural completion, reported as cancelled, not a fault.

Alignment is best effort and never fails a migration. A session that cannot reach an aligned boundary migrates from where it is, exactly as it did before.

Reading a drain’s per-session outcomes

A drain reports one outcome per session, and two of them are not faults. cancelled with reason request has already finished generating means the request ended between the drain’s precheck and the freeze — the session completes where it is, and nothing was lost. Until this build that case read failed; if you are comparing against older drain reports, expect the failed count to drop and the cancelled count to rise by the same amount, with no change in behaviour. blocked with session has not reached an eligible boundary means the request is still in prefill: it is worth retrying the drain, and the other two are not.

A boundary refusal now names its state — unknown to this worker, already finished, still in prefill, no accepted tokens. Only the last two are worth waiting on.

A drain waits for a session whose migration came back cancelled, up to the drain’s deadline, and reports drained once it has completed. Earlier builds finalized straight away: the drain read blocked and the worker stayed closed to admission, although the session finished a moment later, and a blocked drain is never re-evaluated. On the release pair that happened on every drain whose session stopped just after the precheck. If you are still running such a build and see a blocked drain whose only blocker is cancelled, cancel the drain to reopen the worker.

Watching a GPU host

experiments/gpu/gpu_watch.py is a watchdog that runs on the GPU host itself, under supervisord, so it outlives any operator’s session and is restarted if it dies. It uses the standard library and the system interpreter, so a broken experiment environment cannot take it down.

# /etc/supervisor/conf.d/gpu_watch.conf
[program:gpu_watch]
command=/usr/bin/python3 /workspace/ws2/gpu_watch.py --root /workspace/ws2 --out /workspace/ws2/watch
autostart=true
autorestart=true
startretries=1000
stdout_logfile=/workspace/ws2/watch/supervisor_stdout.txt
stderr_logfile=/workspace/ws2/watch/supervisor_stderr.txt

It keeps two files: watch/events.jsonl, one JSON object per alert or notable event, and watch/status.json, the latest snapshot. Tail the first; read the second. What pages, and why each rule is shaped the way it is:

Alert Fires when
LOG_FAILURE a new line in any *.log shows a Traceback, OOM, kill, nonzero exit, position or transcript mismatch, or a failed drill check ([FAIL]). History present at startup never alerts, and a line split across two writes is still one match
STALL an experiment is alive, every GPU is idle, and no log has grown for 7 minutes — measured on log progress, because a worker loading weights sits at 0% utilisation for about a minute and a half
GPU_QUERY_FAILED nvidia-smi fails two ticks running
GPU_HOT / GPU_THROTTLE 85 °C, or a hardware-slowdown, thermal or power-brake throttle reason starts
RESOURCE_LOW disk, /dev/shm (the transfer arena) or available RAM crosses its floor
CANARY_FAILED after 30 idle minutes, each GPU multiplies two float64 matrices and the result disagrees with the CPU. It catches a card that answers nvidia-smi but computes wrongly

Every alert fires when its condition starts, not on every tick it persists. Jobs are identified by script name only: command lines on these hosts can carry secrets, including other services’ access tokens, and none of that is ever copied into a record.

Verified on the release pair: an injected Traceback alerted within one tick; a silent fake job raised one STALL; kill -9 of the watchdog was restarted by supervisord without re-alerting on old failures; and the canary matched the CPU to 1.8e-15 on both GPUs.

Known limitations

These are properties of the design, not defects: