Alert Enhancement: x failures in y minutes
Just wondering if you would consider implementing an alert rule for "x failures in y minutes"? My reason here is that our applications are load balanced between web servers, so it might fail once on a bad server, but the next check is OK because it hits a different server.
We'd like to see if theres the option to have a "X fails in Y minutes" so that we can filter out the random failures (Alert fatigue), and only get alerted when there's frequent failures.
Thanks!
original GH issue link https://github.com/checkly/public-roadmap/issues/122
Log in to comment and vote
Comments1
Andrea Nguyen
Jan 8
We’d like to +1 this request based on a real incident we recently ran into.
We've had a network-related issue that was not detected right away. The problem was intermittent, and while we do have multiple Checkly alerts in place, retries often succeeded after an initial failure. Because of that, alerts never fired even though customers experienced ongoing, intermittent impact.
We could have caught this much earlier if there is below alert strategy provided
X failures within Y minutes, or
% of failures over a rolling time window
These types of alerts are really important for detecting flaky network behavior or partial outages where things “usually” work but reliability is clearly degraded