jev-oncall
Sep 24, 2026, 14:04 UTCAck means someone is on it: the review closes and won't escalate to a page. It doesn't resolve the incident; the alert stays open until your monitoring clears it. If nobody acks before the deadline, the review escalates to a page.
databasetriage-rotation env=staging: non-production never pages (rule, no model call)compute probably noise, but P(actionable) > 0.05: ticket, not dropdatabasecompute duplicate_of checkout (0.87); incident root checkoutnetwork unsure whether to page (0.2 < P(page) < 0.8): a human decides within 15m or it pagescompute probably noise, but P(actionable) > 0.05: ticket, not dropcompute2 labeled alerts is a smoke test. Trust these numbers, and the calibration especially, only after replaying a few hundred labeled alerts.
| What happened | jev-oncall | Your current routing |
|---|---|---|
| Needed a human, reached no one | 0 | 0 |
| Should have paged, got a ticket | 0 | 0 |
| Should have paged, got a review | 0 | 0 |
| Paged when it shouldn't have | 0 | 1 |
| Paged twice for one incident | 0 | 0 |
| Reached the wrong team | 0 | n/a |
| Linked to the wrong incident | 0 | n/a |
| Review that wasn't needed | 0 | 0 |
| Ticket for noise | 1 | 0 |
| Pages sent, all 8 alerts | 3 | 5 |
| Reviews sent, all 8 alerts | 2 | 0 |
Rows above the line count the 2 labeled alerts; pages and reviews count every alert, as decided on arrival. Your current routing means configured severity only: critical pages, warning tickets, info logs, in every environment. jev-oncall also never pages for non-production alerts, which is part of the difference.
| Question | Matched | 95% interval |
|---|---|---|
| Actionable | 2 of 2 | 34% to 100% |
| Page or not | 2 of 2 | 34% to 100% |
| Severity | 1 of 2 | 9% to 91% |
| Owner | 2 of 2 | 34% to 100% |
| Cause | 2 of 2 | 34% to 100% |
| P(page) | Alerts | Predicted | Paged by label |
|---|---|---|---|
| 0.0 to 0.2 | 1 | 0.00 | 0.00 |
| 0.8 to 1.0 | 1 | 1.00 | 1.00 |
Brier score 0.000, calibration error 0.000. Lower is better for both.
Since Sep 24, 2026, 14:03 UTC: 8 alerts. Your current routing paged 5; jev-oncall paged 3 on arrival. 2 alerts labeled so far. Select any alert to label it in the inspector.
Your current routing means configured severity only: critical pages, warning tickets, info logs, in every environment. jev-oncall also never pages for non-production alerts, which is part of the difference. The evaluation below compares the same two.
| Latest differences | Your routing → jev-oncall |
|---|---|
| StagingDiskFull: disk 97% full on staging CI runner | PAGE → LOG |
| PaymentServiceErrors: payment authorization failures 18% | PAGE → REVIEW |
| TlsCertExpiringSoon: TLS certificate for api.example.com expires in 25 days | TICKET → REVIEW |
| ReportJobSlow: nightly revenue report finished 12 minutes later than usual | PAGE → TICKET |
| HomepageLatencyHigh: p99 latency 2.4s against a 400ms SLO | TICKET → PAGE |
| # | Decision | P(page) | ms | |
|---|---|---|---|---|
| #1 | page | 1.00 | 753 | |
| #2 | quiet | - | - | |
| #3 | ticket | 0.07 | 234 | |
| #4 | page | 1.00 | 200 | |
| #5 | review | 1.00 | 221 | |
| #6 | review | 0.47 | 257 | |
| #7 | ticket | 0.00 | 228 | |
| #8 | page | 1.00 | 178 |