Jev judges.Code decides.
Open-source alert triage. Jev scores every alert, and plain, auditable code decides who gets paged.
No page0 to 0.20
2 alerts. Ticketed, or dropped when it's clearly noise.
Review0.20 to 0.80
tls-cert. Unsure costs a human's attention, never silence: if nobody acks it within 15 minutes, it pages.
Page0.80 to 1
orders-db, checkout and homepage page. Payment is a duplicate of checkout, which another team owns, so it goes to review.
One call, four questions, answered with probabilities.
Each question has a typed answer, so the code downstream can reason about how sure Jev is instead of parsing prose.
- actionable
- Yes or no
- Does a human need to intervene? An alert is dropped only when this is near zero and severity agrees.
- severity
- SEV4 to SEV1, a distribution
- P(page) is P(SEV1) + P(SEV2). Routing uses the whole distribution, not only the top label.
- team
- One of your teams
- Picks the owner from your team descriptions. Below 0.60, the runner-up team is notified too.
- duplicate_of
- A recent alert, or none
- Which alert causes this one. The answers become the edges of a dedup graph, so one incident pages once.
Non-production alerts are logged by a rule and never cost a model call. Every threshold lives in your config file, and evaluate.py --sweep replays stored answers under other thresholds so you can choose them from data.
A paging path that degrades to what you have today.
Every rule here exists so that a wrong or missing answer costs noise, not a missed page.
Fails open
If Jev errors, times out or returns a malformed answer, the alert is routed by its configured severity, exactly as it would be without jev-oncall.
The review clock
A review pages after 15 minutes unless someone acks it. An ack means someone is on it, not that it's resolved. Reviews are kept on disk, so a restart doesn't drop them.
Dedup never silences
A linked alert owned by another team gets a review. A page never folds into an earlier alert that got a weaker decision.
Invariant checks
Every run checks that no linked alert outranks its incident root, and that nothing was dropped without a model judgment.
A pinned model
Calls pin
jev-1.13.0, not an alias, so a new release can't shift probabilities under your thresholds.One idempotency key per call
A call and its retry share one key, so a client timeout never judges, or bills, the same alert twice.
One measured run: 300 alerts in under 8 seconds.
Synthetic alerts, 16 workers, the default 2-second timeout. It measures speed on one network path, not whether the routing was right. That takes replayed, labeled history.
- Model calls
- 233; the other 67 alerts were non-production
- Latency per call
- 418 ms median, 1,477 ms p95, 1,849 ms max
- Cost
- $0.0128 in total, about $0.04 per 1,000 alerts
- Fallbacks
- None
The tail nearly reaches the timeout, and 37% of judged alerts landed in review: a sign the default bars are wide for this mix. Read the full write-up.
Plugs into the alerts you already have.
jev-oncall judges and records decisions today. Sending them onward is next.
Takes webhooks from
- Prometheus Alertmanager bearer token
- Grafana Alerting HMAC
- Datadog HMAC header
- PagerDuty v2 and v3 HMAC
- Any JSON HMAC header
Gives you today
- Live dashboard, with Ack /dashboard
- Decisions as JSON /recent
- Reviews and their deadlines /pending
- Shadow mode beside your paging /shadow
- Offline replay and sweep evaluate.py
Not built yet
- Slack, with an Ack button
- PagerDuty and Opsgenie paging
- Jira and Linear tickets
git clone https://github.com/mingleiw/jev-oncall cd jev-oncall/demo echo "TYPESAFE_API_KEY=<your key>" > .env # optional docker compose up --build # then open http://localhost:8090/dashboard
Run the incident on your machine.
Prometheus fires a staged incident and Alertmanager delivers it to jev-oncall exactly as it would in production.
- A database runs out of connections, and checkout and payments fail because of it.
- Noise arrives with it: a staging disk, an expiring certificate, and an alert that clears on its own.
- Two severities are wrong: a slow report marked critical, and a homepage slowdown marked as a warning while conversion falls 22%.
- Without a key you see today's routing by configured severity. With one, see what Jev changes.
Your teams, your thresholds, one file.
Every section is optional. Unknown keys and out-of-range thresholds stop the server before it triages anything, because a typo that silently kept a default would change who gets paged.
[teams] database = "databases, connection pools, caches" network = "CDN, edge, DNS, TLS certificates" [policy] page_bar = 0.80 no_page_bar = 0.20 review_ack_min = 15 [reviews] store = "reviews.jsonl"
Judge every alert. Keep the page path safe.
MIT licensed. Read the code, run the demo, and tune the policy on your own history.