Concepts
Four concepts, one autonomous loop.
Nightjar runs on the same handful of nouns regardless of the stack underneath it: a service you point a watcher at, an anomaly the watcher detects, an incident the anomalies roll up into, and a remediation that closes it. This page walks through each one in plain English and points at where it lives in the platform.
Concept · 01
Services
A service is the Node.js code Nightjar watches. It runs a worker that streams jobs from a BullMQ queue, or any HTTP/TCP surface you want to keep an eye on.
In Nightjar a service is whatever you point a watcher at — a BullMQ worker, a Redis Streams consumer, or the surrounding API that talks to it. Each service owns a set of streams and the configuration that tells the watcher what “healthy” looks like for them (consumer-group lag tolerance, expected PEL depth, the right retry/replay thresholds).
Services are also where ownership lives. Naming a service in the dashboard sets who gets paged when something it owns goes wrong, and the service’s name is what every anomaly recovers against — so the same service can show up in Slack, in your PagerDuty rotation, and in an audit row after the fact.
Where it lives
Sign in first — this surface is owner-scoped and reads from your service rows.
Concept · 02
Anomalies
An anomaly is a single detected deviation from the baseline a service runs against — a consumer-group stall, a PEL creep, a lag spike, an unusual error rate.
A watcher emits anomalies, not alerts. One consumer-group stall on a slow queue is an anomaly; three correlated stalls across queues in the same deploy window are also anomalies, joined into a single incident below. The platform keeps the granularity high — every anomaly stays addressable, with its service, stream, and timestamp attached — so the on-caller sees individual evidence rows, not a smoothed-over summary.
Anomalies are owner-scoped and short-lived in the UI but durable in storage: the same row drives the dashboard feed, the open-incident timeline, and the audit log the day after. Triaging an anomaly is mostly about asking “is this part of an open incident or a fresh signal?” — the tooling answers that for you before you read.
Where it lives
Sign in first — this surface is owner-scoped and reads from your service rows.
Concept · 03
Incidents
An incident is the human-facing bundle: one or more correlated anomalies, the joined timeline, and the plain-English summary you page on.
When anomalies correlate — usually across the same service or the same deploy window — Nightjar promotes them into an incident and stops firing independent pages. The incident record carries the headline summary a tired on-caller can scan at 3 a.m., plus the underlying anomaly rows so anyone joining the call can drill in.
Incidents are also the unit of work between humans and Nightjar. A proposed remediation attaches to an incident, not to the underlying anomalies; an approval (or decline) lives on the incident’s record; and the post-incident audit pulls from the same timeline. Once an incident closes, its anomalies stay queryable — they don’t vanish when the page does.
Where it lives
Sign in first — this surface is owner-scoped and reads from your service rows.
Concept · 04
Remediations
A remediation is what Nightjar proposes to do about an incident — and the rule set decides whether it runs itself, asks first, or refuses outright.
Every remediation is classified the same way: reversible actions (drain a stuck worker, fail over a healthy consumer, replay a bounded job batch) apply themselves and post a Slack summary; destructive actions (rollback a deploy, mass-replay, anything that mutates queue state at scale) sit behind a one-click PagerDuty approval link. The classification is the source of truth — the platform will hold on human confirmation whenever the action would be hard to undo, and auto-apply whenever it wouldn’t.
Nightjar proposes a remediation when an incident matches a rule; the dashboard shows the proposal alongside the incident timeline so the on-caller can sanity-check the call before approving. The audit row records the proposal, the approver, and the executed action, tied to the same anomaly rows that opened the incident.
Where it lives
No auth required — this surface is public.
Walk the loop in order
Services in, remediations out. The four concepts above are the whole loop.
Once a watcher is pointed at a service, the rest of the platform runs itself. The API reference below has the exact request/response shapes; this page has the nouns.