Operations
What to watch, what to alert on, how far it goes, and what to do when it misbehaves.
Health
| Endpoint | Says | Use as |
|---|---|---|
/live | The process is up. | Liveness probe. |
/health | It can reach the database. | Readiness probe. |
/metrics | Prometheus text format. | Scrape target. |
Protect /metrics with QAWK_METRICS_TOKEN unless it is
only reachable from inside — it tells a reader how many devices you have, how many
are failing, and what your channels are called.
Metrics
Prometheus' text format, with no library behind it.
| Metric | Kind | What |
|---|---|---|
qawk_http_requests_total{surface,method,code} | counter | Requests served. surface is ddi (devices), mgmt, qawk or other. |
qawk_http_request_duration_seconds{surface} | histogram | Latency, 5 ms to 10 s. |
qawk_targets{status} | gauge | Devices by update status. |
qawk_actions_active, qawk_rollouts_running | gauge | What is in flight. |
qawk_fleet_devices{fleet} and friends | gauge | Each channel's members, on-release, updating, failed. |
qawk_fleet_releases_pending | gauge | Approvals waiting. |
qawk_fleet_releases_halted | gauge | Releases stopped by their threshold. |
qawk_leader | gauge | 1 on the instance running the background jobs. |
qawk_db_connections{state} | gauge | This instance's pool: acquired, idle, total. |
qawk_uptime_seconds, qawk_goroutines, qawk_heap_bytes | gauge | The process. |
The counters are per instance — sum them. The database gauges are the same on
every instance — take one, or max.
The four alerts worth having
| Alert on | Means |
|---|---|
qawk_fleet_releases_halted > 0 | A release stopped on failures. Someone has to look. Nothing else will move it. |
qawk_fleet_releases_pending > 0 for hours | A release is waiting for an approval nobody has noticed. |
sum(qawk_leader) != 1 | Nobody is running the background jobs (nothing progresses), or two instances think they are. |
qawk_targets{status="error"} rising | Devices failing outside any threshold you set. |
A halted release is the one alert not to route to a dashboard nobody reads. Halting is Qawk deciding it should not carry on without a human. If nobody is told, the fleet simply stops updating, quietly, for as long as it takes someone to notice.
Tracing
The standard OTEL_* environment variables turn on OTLP export —
traces and metrics. Nothing else to configure.
Scaling
The target is ten thousand devices polling every thirty seconds: about 330 polls a second, plus downloads.
What makes it work
- The HTTP side is stateless. Any number of instances behind a load balancer on one database; a device can poll one and report to another.
- Background jobs run once. Whichever instance holds a PostgreSQL advisory lock. The lock belongs to a connection, so an instance that dies — or loses its database — loses it, and another takes over within ten seconds. No extra component, no leader-election sidecar.
- The polling path is short. One transaction on the device's row and its open actions. Tenant settings are cached five seconds per instance instead of read on every request.
- A poll does not wait for the disk. What a poll writes — when the device
was last seen, from where — commits with
synchronous_commit = off: lost in a crash, it is written again at the next poll thirty seconds later. Everything that matters — assignments, reports, releases — commits synchronously.
Measured
Ten thousand simulated devices from qawk-load, against one instance
with the default pool of 20 connections and PostgreSQL 16 in a container, on a lab
machine whose disk was busy with other builds (17–21% of the time stalled on
I/O during the runs):
| Devices | Polling every | Requests | p50 | p99 | Max | Errors |
|---|---|---|---|---|---|---|
| 10,000 | 60 s | 167/s | 1 ms | 2 ms | 1.1 s | 0 of 39,995 |
| 10,000 | 30 s | 333/s | 1 ms | 2 ms | 1.2 s | 0 of 69,994 |
| 10,000 | 60 s, before the asynchronous poll commit | 162/s | 841 ms | 23.4 s | 25.6 s | 0 of 38,941 |
The process sat below 200 MiB. The third row is why the second exists: without the asynchronous poll commit, every poll waited for its own fsync, and on a busy disk ten thousand devices meant twenty-second stalls.
These are polls, not downloads. Ten thousand devices each pulling a
200 MB artifact is a bandwidth and storage problem, not a request-rate one,
and it is bounded by your disk and your network. Use waves
(wavePercent) to spread it, which is what they are for.
Running more than one
- Artifacts must be shared: a volume every instance mounts (ReadWriteMany).
- Keep the sum of
QAWK_DB_MAX_CONNSacross instances under PostgreSQL'smax_connections. - Roll one instance at a time. Devices retry.
Troubleshooting
Nothing is progressing
No waves go out, no system deployment moves, no channel adopts anything. Check
sum(qawk_leader): if it is 0, no instance is running the background
jobs — usually because every instance has lost the database. Check
/health.
A device is not being updated
Go down this list in order; it is ordered by how often each turns out to be the answer:
- Is it in a channel? A device in none gets nothing. Check its attributes against the channel's rule.
- Is it part of a system? Then the orchestrator updates it, never the channel's set. Look at the system deployment, not the channel.
- Is its system
mixed? The orchestrator leaves a system alone while its devices are in different channels. - Is the channel frozen? Or its release halted?
- Does it have an action open already? A device with one is waited for, never overtaken — the channel reaches it on the next pass.
- Is it calling in at all?
lastControllerRequestAton the device. If it is old, this is a network or a device problem, not a Qawk one. - Is the set assignable to it? Incomplete, invalid, or a type its target type does not accept.
A device downloads and downloads
Check /qawk/v1/downloads. A download that restarts from zero
repeatedly usually means a proxy in front of Qawk is not passing
Range through, so a device that loses its connection starts over
instead of resuming.
Links point at the wrong host
Devices get download URLs they cannot reach. Set QAWK_PUBLIC_URL to
the address devices actually use.
Uploads fail at a certain size
The proxy, not Qawk. nginx's client_max_body_size defaults to
1 MB. See TLS and reverse proxies.
Everything is slow after an upgrade
Check qawk_db_connections. If acquired sits at the
pool limit, raise QAWK_DB_MAX_CONNS — or find the query holding
them, which the debug log level will show.
Reading more
QAWK_LOG_LEVEL=debug logs every request. It is loud; turn it off
again.