chmonitorchmonitor
Advanced

PeerDB Monitoring

Read-only PeerDB Mirrors and Peers section — configure PEERDB_API_URL to surface replication status, throughput, and lag without mutating PeerDB.

CHM includes an optional, view-only PeerDB section (Mirrors and Peers) that surfaces replication status, throughput, lag, and per-mirror detail from a PeerDB deployment. CHM never mutates PeerDB — it proxies a read-only allowlist of the PeerDB REST API.

Further reading: see the blog post for why CDC pipelines need this kind of monitoring — lag, replication-slot growth, and batch failures — and a walkthrough of the views below.

Open it at /peerdb once configured. The section is hidden until PEERDB_API_URL is set.

What you can see

PeerDB mirror detail in chmonitor: throughput, replication lag, rows synced, partition sync history and QRep partition progress

ViewWhereHighlights
Mirrors fleet/peerdbStatus KPIs (count-up while still aggregating, last-known numbers cached), per-mirror throughput, prefix groups (qrep_sg_fleetreporting1_*) you can collapse, a lag-triage strip (worst-lag mirrors), and a collapsible logs & alerts feed aggregated across every mirror with error/warn/info filters.
Snapshot / QRep progressmirror detailPer-table initial-load progress from initial_load — partitions completed, rows synced, avg time per partition, and fetch/consolidate phase badges. QRep jobs add searchable, paged partition progress (UUID, duration, start/end, rows in partition vs synced) and a partition-sync history chart.
CDC batch historymirror detailRecent CDC batches (id, LSN range, rows, duration) plus a rows-per-batch chart.
Operation mixmirror detailPer-table insert / update / delete split from table_total_counts.
Slot health/peerdb/peersReplication slots across Postgres peers classified ok / warn / critical by lag, active state, and WAL status; worst-first.
Peer infopeer detailRedacted peer config and server version from peers/info, alongside slots, slot-lag history, and active queries.

Configure

Set the API URL and password

Set PEERDB_API_URL (and PEERDB_PASSWORD if your PeerDB API requires auth), then restart the app.

For the PeerDB UI behind NextAuth, include the /api suffix:

PEERDB_API_URL=https://peerdb.example.com/api
PEERDB_PASSWORD=your-peerdb-ui-password

For a raw flow-api with no auth, use the bare origin:

PEERDB_API_URL=http://localhost:8113

Tune caching and timeouts

VariableDefaultDescription
PEERDB_API_URL—Base URL of the PeerDB REST API. For the PeerDB UI (NextAuth) include the /api suffix; for a raw flow-api use the bare origin (e.g. http://host:8113).
PEERDB_PASSWORD—Sent as HTTP Basic with an empty username (base64(":" + password)). Leave empty if the API has no auth. Server-side only — never sent to the browser.
PEERDB_CACHE_TTL_MS10000TTL for the server-side response cache (set 0 to disable).
PEERDB_CACHE_MAX_ENTRIES500Max cached responses before oldest entries are evicted.
PEERDB_FETCH_TIMEOUT_MS10000Upstream request timeout.
PEERDB_SWEEP_CONCURRENCY8Max in-flight reads per alert-sweep tick. Lower it if PeerDB struggles under the sweep.
PEERDB_SWEEP_BUDGET_MS60000Wall-clock budget for one alert-sweep collection. Mirrors left unread when it elapses are deferred to the next tick.
PEERDB_SWEEP_MAX_MIRRORSunset (no limit)Optional cap on mirrors read per alert-sweep tick. Leave unset so every listed mirror is read; see Partial coverage.

See the full list in Environment Variables.

Connection status

The header shows a status pill that distinguishes:

  • Connected — API reachable and authenticated.
  • Auth failed — credentials rejected (check PEERDB_PASSWORD). For the PeerDB UI this is the UI login password.
  • Unreachable — wrong PEERDB_API_URL or a network/DNS issue.
  • Not configured — PEERDB_API_URL is unset.

Security

Read-only proxy

CHM proxies only a read-only allowlist of PeerDB endpoints (app/api/v1/peerdb/[...slug]). Mutating endpoints (create/drop/pause, alert config, maintenance) are rejected with 403. The PeerDB credential is attached server-side and never reaches the browser bundle, and secret-shaped peer config fields are masked in the UI.

The section also respects Feature Permissions — gate it with CHM_FEATURE_PEERDB_ACCESS=authenticated or disable it with CHM_FEATURE_PEERDB_ENABLED=false. The proxy enforces the same gate, so it cannot be reached directly when the feature is disabled or restricted.

Local development (mock)

To preview the full UI without a real PeerDB instance, run the bundled mock server:

pnpm run peerdb:mock                                  # serves :8113
PEERDB_API_URL=http://localhost:8113 pnpm run dev     # → /peerdb

Large mirror fleets

PeerDB's only index on flow_errors is flow_name, so the mirror-logs endpoint filtered by level=ERROR scans every error row for that mirror. On a fleet with a multi-million-row flow_errors (QRep writes most of it at info) that read is expensive and frequently returns nothing useful.

To make it cheap, add a composite index to the PeerDB catalog Postgres:

CREATE INDEX CONCURRENTLY flow_errors_flow_name_error_type_idx
  ON flow_errors (flow_name, error_type);

Without it, per-mirror ERROR-log reads degrade from a few hundred milliseconds to seconds, and enough of them in parallel will time out.

The alert sweep bounds its own load regardless of the index:

  • it reads ERROR logs only for mirrors whose status is not STATUS_RUNNING, or that report an errorMessage — a healthy running mirror costs no log read;
  • it runs at most PEERDB_SWEEP_CONCURRENCY reads at a time;
  • it stops starting new work after PEERDB_SWEEP_BUDGET_MS, deferring the remaining mirrors to the next tick rather than overrunning the sweep.

On a 72-mirror fleet this turns ~100 log reads per tick into zero when the fleet is healthy.

Partial coverage

The alert sweep reads every mirror PeerDB lists. There is no hidden ceiling on fleet size: concurrency and wall-clock budget limit how fast it reads, not which mirrors it is allowed to read, and the mirrors most likely to alert are the ones it reads first.

A tick that could not read the whole fleet — because PEERDB_SWEEP_BUDGET_MS elapsed, or because PEERDB_SWEEP_MAX_MIRRORS is set — says so rather than reporting a clean result over a fleet it never looked at:

  • the sweep summary reports the mirrors checked out of the mirrors listed, and flags the run as partial;
  • an alert_events audit row with decision peerdb-coverage:partial records the counts, e.g. 50 of 72 mirrors checked — partial result: 22 unchecked.

A partial run is a coverage fact, not an incident: nothing is dispatched for it and no mirror's alert state is changed. The mirrors that were read are still classified normally. Mirrors left unread are retried on the next tick, so a fleet too large for one tick needs either a larger PEERDB_SWEEP_BUDGET_MS or a smaller fleet — not a cap.

Troubleshooting

SymptomLikely cause
Every PeerDB check reports errored on a large fleetThe sweep is timing out against the PeerDB catalog. Add the flow_errors (flow_name, error_type) index above, and lower PEERDB_SWEEP_CONCURRENCY.
Pill shows Auth failedWrong PEERDB_PASSWORD. For the PeerDB UI, use the UI login password; for a raw flow-api, match its configured password (or leave empty).
Pill shows UnreachablePEERDB_API_URL host/port wrong, or the API is not reachable from the CHM server.
Mirrors load but charts are emptyThe mirror has no recent CDC graph/batch data yet, or the PeerDB version doesn't expose those endpoints.
Large fleets show partial KPI totalsPer-row metrics load lazily above 24 mirrors; the Throughput/Rows-synced cards label how many mirrors are loaded. Expand a row to load its metrics.
Alert sweep reports peerdb-coverage:partialThe tick ran out of budget (or PEERDB_SWEEP_MAX_MIRRORS is set) before reading every mirror. Raise PEERDB_SWEEP_BUDGET_MS, lower PEERDB_SWEEP_CONCURRENCY if reads are timing out, or unset PEERDB_SWEEP_MAX_MIRRORS. See Partial coverage.

Blog: Monitoring PeerDB: snapshot progress, batch history, fleet lag, and slot health

On this page