Jev judges.Code decides.

Open-source alert triage. Jev scores every alert, and plain, auditable code decides who gets paged.

No page0 to 0.20

2 alerts. Ticketed, or dropped when it's clearly noise.

Review0.20 to 0.80

tls-cert. Unsure costs a human's attention, never silence: if nobody acks it within 15 minutes, it pages.

Page0.80 to 1

orders-db, checkout and homepage page. Payment is a duplicate of checkout, which another team owns, so it goes to review.

Each dot is an alert from the demo's staged incident, placed at P(page), the probability Jev's answer gives SEV1 or SEV2, from the real jev-1.13.0 run the demo replays. The staging alert is left off: a rule logs it without asking Jev. The thresholds decide first; dedup and ownership rules can change the outcome after. Open the demo to ack, label and fail them yourself.

One call, four questions, answered with probabilities.

Each question has a typed answer, so the code downstream can reason about how sure Jev is instead of parsing prose.

actionable
Yes or no
Does a human need to intervene? An alert is dropped only when this is near zero and severity agrees.
severity
SEV4 to SEV1, a distribution
P(page) is P(SEV1) + P(SEV2). Routing uses the whole distribution, not only the top label.
team
One of your teams
Picks the owner from your team descriptions. Below 0.60, the runner-up team is notified too.
duplicate_of
A recent alert, or none
Which alert causes this one. The answers become the edges of a dedup graph, so one incident pages once.

Non-production alerts are logged by a rule and never cost a model call. Every threshold lives in your config file, and evaluate.py --sweep replays stored answers under other thresholds so you can choose them from data.

A paging path that degrades to what you have today.

Every rule here exists so that a wrong or missing answer costs noise, not a missed page.

  • Fails open

    If Jev errors, times out or returns a malformed answer, the alert is routed by its configured severity, exactly as it would be without jev-oncall.

  • The review clock

    A review pages after 15 minutes unless someone acks it. An ack means someone is on it, not that it's resolved. Reviews are kept on disk, so a restart doesn't drop them.

  • Dedup never silences

    A linked alert owned by another team gets a review. A page never folds into an earlier alert that got a weaker decision.

  • Invariant checks

    Every run checks that no linked alert outranks its incident root, and that nothing was dropped without a model judgment.

  • A pinned model

    Calls pin jev-1.13.0, not an alias, so a new release can't shift probabilities under your thresholds.

  • One idempotency key per call

    A call and its retry share one key, so a client timeout never judges, or bills, the same alert twice.

One measured run: 300 alerts in under 8 seconds.

Synthetic alerts, 16 workers, the default 2-second timeout. It measures speed on one network path, not whether the routing was right. That takes replayed, labeled history.

The jev-oncall dashboard for the 300-alert run: 63 flagged for paging, 86 for review, 33 linked to an incident that already paged, and 118 ticketed, logged or dropped, each a dot on the P(page) scale against the 0.20 and 0.80 policy bars.
Model calls
233; the other 67 alerts were non-production
Latency per call
418 ms median, 1,477 ms p95, 1,849 ms max
Cost
$0.0128 in total, about $0.04 per 1,000 alerts
Fallbacks
None

The tail nearly reaches the timeout, and 37% of judged alerts landed in review: a sign the default bars are wide for this mix. Read the full write-up.

Plugs into the alerts you already have.

jev-oncall judges and records decisions today. Sending them onward is next.

Takes webhooks from

  • Prometheus Alertmanager bearer token
  • Grafana Alerting HMAC
  • Datadog HMAC header
  • PagerDuty v2 and v3 HMAC
  • Any JSON HMAC header

Gives you today

  • Live dashboard, with Ack /dashboard
  • Decisions as JSON /recent
  • Reviews and their deadlines /pending
  • Shadow mode beside your paging /shadow
  • Offline replay and sweep evaluate.py

Not built yet

  • Slack, with an Ack button
  • PagerDuty and Opsgenie paging
  • Jira and Linear tickets
terminal
git clone https://github.com/mingleiw/jev-oncall
cd jev-oncall/demo
echo "TYPESAFE_API_KEY=<your key>" > .env   # optional
docker compose up --build
# then open http://localhost:8090/dashboard

Run the incident on your machine.

Prometheus fires a staged incident and Alertmanager delivers it to jev-oncall exactly as it would in production.

  • A database runs out of connections, and checkout and payments fail because of it.
  • Noise arrives with it: a staging disk, an expiring certificate, and an alert that clears on its own.
  • Two severities are wrong: a slow report marked critical, and a homepage slowdown marked as a warning while conversion falls 22%.
  • Without a key you see today's routing by configured severity. With one, see what Jev changes.

Your teams, your thresholds, one file.

Every section is optional. Unknown keys and out-of-range thresholds stop the server before it triages anything, because a typo that silently kept a default would change who gets paged.

See the architecture

jev-oncall.toml
[teams]
database = "databases, connection pools, caches"
network  = "CDN, edge, DNS, TLS certificates"

[policy]
page_bar       = 0.80
no_page_bar    = 0.20
review_ack_min = 15

[reviews]
store = "reviews.jsonl"

Judge every alert. Keep the page path safe.

MIT licensed. Read the code, run the demo, and tune the policy on your own history.