Configurable Alert Debouncing for Degraded Response Time Thresholds
Problem Statement
Currently, when a user sets a "degraded" response time threshold (e.g., 10s) in Checkly and connects an alert channel, an alert is triggered every time a check exceeds this threshold. This can lead to alert fatigue, as a single transient issue or a brief period of instability may result in multiple, potentially redundant alerts being sent in rapid succession.
Proposed Solution
Introduce a configurable alert debouncing mechanism for degraded response time thresholds. This feature would allow users to specify that an alert for degraded performance should only be sent after a defined number (x) of consecutive or cumulative degraded check results, rather than after every individual occurrence.
Feature Details
Allow users to set a parameter (e.g., "Send alert after X degraded checks") when configuring a degraded response time threshold.
Support both consecutive and rolling window logic for degraded check counting:
Consecutive: Alert is sent only if X degraded results occur in a row.
Rolling window: Alert is sent if X degraded results occur within the last Y check runs or time window.
Reset the counter after an alert is sent, or provide an option to suppress further alerts until a recovery condition is met (e.g., a check returns to "OK" status).
Optionally, provide visibility in the UI or via API to see current degraded check counts and pending alert status.
User Value
Reduces alert noise and fatigue, especially during transient or intermittent issues.
Allows teams to focus on actionable incidents rather than being overwhelmed by redundant alerts.
Increases flexibility and control over alerting behavior, aligning with best practices in incident management and observability.
Example Scenario
A user sets a degraded threshold of 10s and configures "Send alert after 3 consecutive degraded checks." If three checks in a row exceed 10s, a single alert is triggered. If only one or two checks are degraded, no alert is sent, reducing unnecessary notifications for minor blips.
Log in to comment and vote
Comments3
Rickard Borgmäster
Aug 12
This matches our use case very closely. One important aspect for us is that the degraded debounce should be independent of the failed-run alert threshold.
For example, we would like to alert after 3 consecutive Degraded runs, while still alerting immediately after 1 Failed run.
The current failedRunThreshold cannot achieve this, since it also delays alerts for actual failures. We want to suppress transient performance degradation without delaying availability/failure alerts.
Alberto Gomez
Aug 4, 2025
Hey Sven, we are considering improvements in our alerting mechanisms for Q4 and this request is very helpful. We currently have partial support to what you are asking but we are not distinguishing between failures and degraded states. Our code currently allows users to configure the existing
failedRunThresholdsetting to achieve debouncing for BOTH failures and degraded states. Example:Set runBasedEscalation.failedRunThreshold to 3 → alerts sent after 3 consecutive failures OR degraded states. And this applies to both check-level and account-level alert settings
The other interesting ask is the rolling window, because currently we can account for consecutive failures, but not failures happening in a defined rolling window. That’s something we might also look into.
Thanks so much for the request and we will get back to you once we start evaluating the improvements.
Regards,
Alberto
Rickard Borgmäster
Aug 12
It would be useful to configure separate alert escalation thresholds for Failed and Degraded check results.
Our use case is that we run API checks every minute and use a response time threshold to mark slow responses as Degraded. We occasionally see isolated slow runs, while the next run is back to normal. We don't want these transient performance spikes to generate alerts.
At the same time, an actual Failed check is much more critical and should generate an alert immediately.
Ideally, we would like to configure a check like this:
Failed: Alert after 1 failed run
Degraded: Alert after 3 consecutive degraded runs
Currently, failedRunThreshold applies to both Failed and Degraded results. Setting it to 3 therefore reduces noise from transient Degraded results, but also delays alerts for actual failures.
Separate thresholds, for example failedRunThreshold and degradedRunThreshold, would allow different alerting policies for availability failures and performance degradation without requiring duplicate checks or custom scripting.
This would make it possible to keep fast failure detection while avoiding alert noise from short-lived performance degradation.