How an alert becomes a decision.

Jev answers four questions about each production alert, and plain code turns the answers into a decision. The server returns and shows decisions; it doesn't send pages, messages or tickets to anything yet.

Scroll sideways to see the whole diagram.

server pipeline
Sources Prometheus rules fire alerts Alertmanager groups by service webhook bearer token Other providers Grafana, Datadog PagerDuty, generic JSON Fail-open route by configured severity: critical PAGE, warning TICKET, info LOG on error decision jev-oncall server 01 Normalize the alert one id per firing: fingerprint + startsAt resend: skipped, resolved: cancels its REVIEW 02 Non-production goes to LOG a rule, before any model call env other than prod never pages 03 Pick candidate causes prod alerts from 30 min before to 2 min after narrowed by topology when listed, at most 50 04 Ask Jev: one call, four questions actionable, severity, team, duplicate_of 2 s timeout, 1 retry 05 Route on the probabilities P(page) = P(SEV1) + P(SEV2) ≥ 0.80 page, 0.20 to 0.80 REVIEW, else TICKET 06 Dedup as a graph link at P ≥ 0.70, root gets the most urgent action same team: DEDUP, other team: REVIEW 07 Respond and record decisions returned as JSON each REVIEW starts a 15 min ack clock, kept on disk Read decisions today Server endpoints /dashboard live page, every 15 s /recent decisions as JSON /pending REVIEWs awaiting an ack POST /ack/<id> on it, no page /health HTTP Not built yet PagerDuty, Slack, Jira nothing is sent anywhere today planned Jev System One API api.typesafe.ai model pinned to jev-1.13.0 questions probabilities Recent decisions and reviews decisions: last 500, in memory reviews: on disk, survive restart prior alerts record
Only step 4 leaves the server. If Jev errors or times out, the alert skips step 5 and is routed by its configured severity, so a Jev outage behaves like having no jev-oncall at all. Step 6 can also link an alert to one from an earlier webhook delivery, using the recent decisions, but only when that earlier alert already got an action at least as urgent.

The review clock

Being unsure costs a REVIEW, never silence.

review clock
REVIEW sent 0.20 < P(page) < 0.80 POST /ack/<id> Acked: someone is on it no page, alert not resolved resolved notification Cancelled no page, alert cleared itself 15 min, nobody acked Escalated to PAGE recorded once, not sent yet
An ack doesn't resolve anything: the alert stays open until the monitoring system clears it. If the same alert is sent again, its first deadline stands, so repeated webhooks can't keep pushing the page back, and a closed review stays closed. With [reviews] store set, every open, ack, cancellation and escalation is written to disk before it counts: a restarted server restores pending reviews with their original deadlines, pages the ones whose deadline passed while it was down, and never repeats an escalation or reopens an acked or cancelled review. That supports one server process per store file, on local disk. Without a store, reviews live in memory and a restart drops them, and the server warns at startup.

Default thresholds

Starting points, set in the [policy] section of the config file. Tune them with evaluate.py --sweep on replayed, labeled history.

keydefaultwhat it decides
page_bar0.80P(page) at or above this pages. PAGE_NOW when SEV1 is at least as likely as SEV2
no_page_bar0.20P(page) at or below this doesn't page. Between the two bars is REVIEW
drop_bar0.05Not paging and P(actionable) at or below this: DROP, the only outcome nobody sees
dedup_bar0.70Minimum P(duplicate_of = X) to link an alert to its cause
team_bar0.60Below this, a page or REVIEW also notifies the runner-up team
window_min, skew_min30, 2Minutes a cause may start before, or after, its symptom
max_candidates50Most candidate causes offered to Jev per alert
review_ack_min15Minutes before an unacked REVIEW pages