PeerDB Monitoring
Read-only PeerDB Mirrors and Peers section — configure PEERDB_API_URL to surface replication status, throughput, and lag without mutating PeerDB.
CHM includes an optional, view-only PeerDB section (Mirrors and Peers) that surfaces replication status, throughput, lag, and per-mirror detail from a PeerDB deployment. CHM never mutates PeerDB — it proxies a read-only allowlist of the PeerDB REST API.
Further reading: see the blog post for why CDC pipelines need this kind of monitoring — lag, replication-slot growth, and batch failures — and a walkthrough of the views below.
Open it at /peerdb once configured. The section is hidden until
PEERDB_API_URL is set.
What you can see

| View | Where | Highlights |
|---|---|---|
| Mirrors fleet | /peerdb | Status KPIs (count-up while still aggregating, last-known numbers cached), per-mirror throughput, prefix groups (qrep_sg_fleetreporting1_*) you can collapse, a lag-triage strip (worst-lag mirrors), and a collapsible logs & alerts feed aggregated across every mirror with error/warn/info filters. |
| Snapshot / QRep progress | mirror detail | Per-table initial-load progress from initial_load — partitions completed, rows synced, avg time per partition, and fetch/consolidate phase badges. QRep jobs add searchable, paged partition progress (UUID, duration, start/end, rows in partition vs synced) and a partition-sync history chart. |
| CDC batch history | mirror detail | Recent CDC batches (id, LSN range, rows, duration) plus a rows-per-batch chart. |
| Operation mix | mirror detail | Per-table insert / update / delete split from table_total_counts. |
| Slot health | /peerdb/peers | Replication slots across Postgres peers classified ok / warn / critical by lag, active state, and WAL status; worst-first. |
| Peer info | peer detail | Redacted peer config and server version from peers/info, alongside slots, slot-lag history, and active queries. |
Configure
Set the API URL and password
Set PEERDB_API_URL (and PEERDB_PASSWORD if your PeerDB API requires auth),
then restart the app.
For the PeerDB UI behind NextAuth, include the /api suffix:
PEERDB_API_URL=https://peerdb.example.com/api
PEERDB_PASSWORD=your-peerdb-ui-passwordFor a raw flow-api with no auth, use the bare origin:
PEERDB_API_URL=http://localhost:8113Tune caching and timeouts
| Variable | Default | Description |
|---|---|---|
PEERDB_API_URL | — | Base URL of the PeerDB REST API. For the PeerDB UI (NextAuth) include the /api suffix; for a raw flow-api use the bare origin (e.g. http://host:8113). |
PEERDB_PASSWORD | — | Sent as HTTP Basic with an empty username (base64(":" + password)). Leave empty if the API has no auth. Server-side only — never sent to the browser. |
PEERDB_CACHE_TTL_MS | 10000 | TTL for the server-side response cache (set 0 to disable). |
PEERDB_CACHE_MAX_ENTRIES | 500 | Max cached responses before oldest entries are evicted. |
PEERDB_FETCH_TIMEOUT_MS | 10000 | Upstream request timeout. |
PEERDB_SWEEP_CONCURRENCY | 8 | Max in-flight reads per alert-sweep tick. Lower it if PeerDB struggles under the sweep. |
PEERDB_SWEEP_BUDGET_MS | 60000 | Wall-clock budget for one alert-sweep collection. Mirrors left unread when it elapses are deferred to the next tick. |
PEERDB_SWEEP_MAX_MIRRORS | unset (no limit) | Optional cap on mirrors read per alert-sweep tick. Leave unset so every listed mirror is read; see Partial coverage. |
See the full list in Environment Variables.
Connection status
The header shows a status pill that distinguishes:
- Connected — API reachable and authenticated.
- Auth failed — credentials rejected (check
PEERDB_PASSWORD). For the PeerDB UI this is the UI login password. - Unreachable — wrong
PEERDB_API_URLor a network/DNS issue. - Not configured —
PEERDB_API_URLis unset.
Security
Read-only proxy
CHM proxies only a read-only allowlist of PeerDB endpoints
(app/api/v1/peerdb/[...slug]). Mutating endpoints (create/drop/pause, alert
config, maintenance) are rejected with 403. The PeerDB credential is attached
server-side and never reaches the browser bundle, and secret-shaped peer config
fields are masked in the UI.
The section also respects
Feature Permissions — gate it with
CHM_FEATURE_PEERDB_ACCESS=authenticated or disable it with
CHM_FEATURE_PEERDB_ENABLED=false. The proxy enforces the same gate, so it
cannot be reached directly when the feature is disabled or restricted.
Local development (mock)
To preview the full UI without a real PeerDB instance, run the bundled mock server:
pnpm run peerdb:mock # serves :8113
PEERDB_API_URL=http://localhost:8113 pnpm run dev # → /peerdbLarge mirror fleets
PeerDB's only index on flow_errors is flow_name, so the mirror-logs
endpoint filtered by level=ERROR scans every error row for that mirror. On a
fleet with a multi-million-row flow_errors (QRep writes most of it at info)
that read is expensive and frequently returns nothing useful.
To make it cheap, add a composite index to the PeerDB catalog Postgres:
CREATE INDEX CONCURRENTLY flow_errors_flow_name_error_type_idx
ON flow_errors (flow_name, error_type);Without it, per-mirror ERROR-log reads degrade from a few hundred milliseconds to seconds, and enough of them in parallel will time out.
The alert sweep bounds its own load regardless of the index:
- it reads ERROR logs only for mirrors whose status is not
STATUS_RUNNING, or that report anerrorMessage— a healthy running mirror costs no log read; - it runs at most
PEERDB_SWEEP_CONCURRENCYreads at a time; - it stops starting new work after
PEERDB_SWEEP_BUDGET_MS, deferring the remaining mirrors to the next tick rather than overrunning the sweep.
On a 72-mirror fleet this turns ~100 log reads per tick into zero when the fleet is healthy.
Partial coverage
The alert sweep reads every mirror PeerDB lists. There is no hidden ceiling on fleet size: concurrency and wall-clock budget limit how fast it reads, not which mirrors it is allowed to read, and the mirrors most likely to alert are the ones it reads first.
A tick that could not read the whole fleet — because PEERDB_SWEEP_BUDGET_MS
elapsed, or because PEERDB_SWEEP_MAX_MIRRORS is set — says so rather than
reporting a clean result over a fleet it never looked at:
- the sweep summary reports the mirrors checked out of the mirrors listed, and flags the run as partial;
- an
alert_eventsaudit row with decisionpeerdb-coverage:partialrecords the counts, e.g.50 of 72 mirrors checked — partial result: 22 unchecked.
A partial run is a coverage fact, not an incident: nothing is dispatched for it
and no mirror's alert state is changed. The mirrors that were read are still
classified normally. Mirrors left unread are retried on the next tick, so a
fleet too large for one tick needs either a larger PEERDB_SWEEP_BUDGET_MS or
a smaller fleet — not a cap.
Troubleshooting
| Symptom | Likely cause |
|---|---|
| Every PeerDB check reports errored on a large fleet | The sweep is timing out against the PeerDB catalog. Add the flow_errors (flow_name, error_type) index above, and lower PEERDB_SWEEP_CONCURRENCY. |
| Pill shows Auth failed | Wrong PEERDB_PASSWORD. For the PeerDB UI, use the UI login password; for a raw flow-api, match its configured password (or leave empty). |
| Pill shows Unreachable | PEERDB_API_URL host/port wrong, or the API is not reachable from the CHM server. |
| Mirrors load but charts are empty | The mirror has no recent CDC graph/batch data yet, or the PeerDB version doesn't expose those endpoints. |
| Large fleets show partial KPI totals | Per-row metrics load lazily above 24 mirrors; the Throughput/Rows-synced cards label how many mirrors are loaded. Expand a row to load its metrics. |
Alert sweep reports peerdb-coverage:partial | The tick ran out of budget (or PEERDB_SWEEP_MAX_MIRRORS is set) before reading every mirror. Raise PEERDB_SWEEP_BUDGET_MS, lower PEERDB_SWEEP_CONCURRENCY if reads are timing out, or unset PEERDB_SWEEP_MAX_MIRRORS. See Partial coverage. |
Related
Feature Permissions
Gate or disable the PeerDB section with CHM_FEATURE_PEERDB_ACCESS and CHM_FEATURE_PEERDB_ENABLED.
Environment Variables
Full reference for all PeerDB and connection variables.
Blog: Monitoring PeerDB: snapshot progress, batch history, fleet lag, and slot health