Skip to main content

Alert Enhancement: x failures in y minutes

Just wondering if you would consider implementing an alert rule for "x failures in y minutes"? My reason here is that our applications are load balanced between web servers, so it might fail once on a bad server, but the next check is OK because it hits a different server.

We'd like to see if theres the option to have a "X fails in Y minutes" so that we can filter out the random failures (Alert fatigue), and only get alerted when there's frequent failures.

Thanks!

original GH issue link https://github.com/checkly/public-roadmap/issues/122

1 comment

Log in to comment and vote

Comments1

  • Andrea Nguyen

    •

    Jan 8

    We’d like to +1 this request based on a real incident we recently ran into.
    We've had a network-related issue that was not detected right away. The problem was intermittent, and while we do have multiple Checkly alerts in place, retries often succeeded after an initial failure. Because of that, alerts never fired even though customers experienced ongoing, intermittent impact.
    We could have caught this much earlier if there is below alert strategy provided

    • X failures within Y minutes, or

    • % of failures over a rolling time window

    These types of alerts are really important for detecting flaky network behavior or partial outages where things “usually” work but reliability is clearly degraded