More control over service degradations
We receive frequent alerts due to alleged service degradation which can't be reproduced on our end. Our servers run in Germany (Hetzner, Falkenstein), and we perform checks using Frankfurt and other European regions. Sometimes we do receive degradation alerts from Ireland/Milan since the HTTP request took more than 3 seconds. These are always one-time incidents, the next check works fine, and we don't see slow response times in our access logs and almost never receive degradation alerts from Frankfurt.
It would be great if there could be additional thresholds for degradation alerts, e.g. only alert after three successive degradations were observed, or only if the degradation is observed from all regions simultaneously.
For now we're limiting the checks to only Frankfurt and Paris and enable double checks in order to hopefully reduce the amount of incorrect alerts we receive. Please let me know if you need more information.
original GH issue link https://github.com/checkly/public-roadmap/issues/293
Log in to comment and vote
Comments3
Steve Feldman
Mar 7, 2025
We really would benefit from this. Using degraded alerts is almost unusable because it’s just not reliable.
Gasper Jeklin
Nov 21, 2024
We would really love to see this.
1. Failures have retry/alert logic
2. Degraded state doesn’t have such benefits
Network or application can have a small glitch in the matrix and one response takes a bit longer and it triggers degraded state. it would make sense to implement similar logic for retries/alerts as for failures.
Amir Jaron
Feb 26, 2025
Joining on this one. We had to turn of degraded notification due to false positives.
We should be able to configure retry for degraded alerts as all on API checks.