Scroll sideways to see the whole diagram.
The review clock
Being unsure costs a REVIEW, never silence.
[reviews] store set, every open, ack, cancellation and escalation is written to disk before it counts: a restarted server restores pending reviews with their original deadlines, pages the ones whose deadline passed while it was down, and never repeats an escalation or reopens an acked or cancelled review. That supports one server process per store file, on local disk. Without a store, reviews live in memory and a restart drops them, and the server warns at startup.Default thresholds
Starting points, set in the [policy] section of the config file. Tune them with evaluate.py --sweep on replayed, labeled history.
| key | default | what it decides |
|---|---|---|
| page_bar | 0.80 | P(page) at or above this pages. PAGE_NOW when SEV1 is at least as likely as SEV2 |
| no_page_bar | 0.20 | P(page) at or below this doesn't page. Between the two bars is REVIEW |
| drop_bar | 0.05 | Not paging and P(actionable) at or below this: DROP, the only outcome nobody sees |
| dedup_bar | 0.70 | Minimum P(duplicate_of = X) to link an alert to its cause |
| team_bar | 0.60 | Below this, a page or REVIEW also notifies the runner-up team |
| window_min, skew_min | 30, 2 | Minutes a cause may start before, or after, its symptom |
| max_candidates | 50 | Most candidate causes offered to Jev per alert |
| review_ack_min | 15 | Minutes before an unacked REVIEW pages |