Operations

What to watch, what to alert on, how far it goes, and what to do when it misbehaves.

Health

EndpointSaysUse as
/liveThe process is up.Liveness probe.
/healthIt can reach the database.Readiness probe.
/metricsPrometheus text format.Scrape target.

Protect /metrics with QAWK_METRICS_TOKEN unless it is only reachable from inside — it tells a reader how many devices you have, how many are failing, and what your channels are called.

Metrics

Prometheus' text format, with no library behind it.

MetricKindWhat
qawk_http_requests_total{surface,method,code}counterRequests served. surface is ddi (devices), mgmt, qawk or other.
qawk_http_request_duration_seconds{surface}histogramLatency, 5 ms to 10 s.
qawk_targets{status}gaugeDevices by update status.
qawk_actions_active, qawk_rollouts_runninggaugeWhat is in flight.
qawk_fleet_devices{fleet} and friendsgaugeEach channel's members, on-release, updating, failed.
qawk_fleet_releases_pendinggaugeApprovals waiting.
qawk_fleet_releases_haltedgaugeReleases stopped by their threshold.
qawk_leadergauge1 on the instance running the background jobs.
qawk_db_connections{state}gaugeThis instance's pool: acquired, idle, total.
qawk_uptime_seconds, qawk_goroutines, qawk_heap_bytesgaugeThe process.

The counters are per instance — sum them. The database gauges are the same on every instance — take one, or max.

The four alerts worth having

Alert onMeans
qawk_fleet_releases_halted > 0A release stopped on failures. Someone has to look. Nothing else will move it.
qawk_fleet_releases_pending > 0 for hoursA release is waiting for an approval nobody has noticed.
sum(qawk_leader) != 1Nobody is running the background jobs (nothing progresses), or two instances think they are.
qawk_targets{status="error"} risingDevices failing outside any threshold you set.
Note

A halted release is the one alert not to route to a dashboard nobody reads. Halting is Qawk deciding it should not carry on without a human. If nobody is told, the fleet simply stops updating, quietly, for as long as it takes someone to notice.

Tracing

The standard OTEL_* environment variables turn on OTLP export — traces and metrics. Nothing else to configure.

Scaling

The target is ten thousand devices polling every thirty seconds: about 330 polls a second, plus downloads.

What makes it work

Measured

Ten thousand simulated devices from qawk-load, against one instance with the default pool of 20 connections and PostgreSQL 16 in a container, on a lab machine whose disk was busy with other builds (17–21% of the time stalled on I/O during the runs):

DevicesPolling everyRequestsp50p99MaxErrors
10,00060 s167/s1 ms2 ms1.1 s0 of 39,995
10,00030 s333/s1 ms2 ms1.2 s0 of 69,994
10,00060 s, before the asynchronous poll commit162/s841 ms23.4 s25.6 s0 of 38,941

The process sat below 200 MiB. The third row is why the second exists: without the asynchronous poll commit, every poll waited for its own fsync, and on a busy disk ten thousand devices meant twenty-second stalls.

Careful

These are polls, not downloads. Ten thousand devices each pulling a 200 MB artifact is a bandwidth and storage problem, not a request-rate one, and it is bounded by your disk and your network. Use waves (wavePercent) to spread it, which is what they are for.

Running more than one

Troubleshooting

Nothing is progressing

No waves go out, no system deployment moves, no channel adopts anything. Check sum(qawk_leader): if it is 0, no instance is running the background jobs — usually because every instance has lost the database. Check /health.

A device is not being updated

Go down this list in order; it is ordered by how often each turns out to be the answer:

  1. Is it in a channel? A device in none gets nothing. Check its attributes against the channel's rule.
  2. Is it part of a system? Then the orchestrator updates it, never the channel's set. Look at the system deployment, not the channel.
  3. Is its system mixed? The orchestrator leaves a system alone while its devices are in different channels.
  4. Is the channel frozen? Or its release halted?
  5. Does it have an action open already? A device with one is waited for, never overtaken — the channel reaches it on the next pass.
  6. Is it calling in at all? lastControllerRequestAt on the device. If it is old, this is a network or a device problem, not a Qawk one.
  7. Is the set assignable to it? Incomplete, invalid, or a type its target type does not accept.

A device downloads and downloads

Check /qawk/v1/downloads. A download that restarts from zero repeatedly usually means a proxy in front of Qawk is not passing Range through, so a device that loses its connection starts over instead of resuming.

Links point at the wrong host

Devices get download URLs they cannot reach. Set QAWK_PUBLIC_URL to the address devices actually use.

Uploads fail at a certain size

The proxy, not Qawk. nginx's client_max_body_size defaults to 1 MB. See TLS and reverse proxies.

Everything is slow after an upgrade

Check qawk_db_connections. If acquired sits at the pool limit, raise QAWK_DB_MAX_CONNS — or find the query holding them, which the debug log level will show.

Reading more

QAWK_LOG_LEVEL=debug logs every request. It is loud; turn it off again.