Skip to main content

Customization of email alert content

Hello Team,

I am requesting a feature of customizing email alert content. Where we can attach image showing failure and url and other required information where test is being run.

Regards

Anil

6 comments

Log in to comment and vote

Comments6

  • IR

    •

    Apr 29

    Hi Checkly team,

    We spent all day yesterday standing up a fresh setup on Checkly for our marketing properties - about 28 checks (URL and Playwright/Browser), alert routing across email / SMS / call channels, and a public status dashboard - driven primarily through your REST API and CLI from a Cowork session that pairs an LLM with the account.

    Below is grounded feedback from that work. Every item traces back to specific friction we hit, with the moment it hit included so your team can reproduce. The biggest ask is in Part 1: an in-Checkly AI agent that understands and manages an account holistically. Part 2 is a list of smaller API and platform improvements.

    Part 1 - Rocky AI

    The biggest gap we felt today: there is no in-Checkly assistant that understands your account holistically and can act on it. We ended up doing "AI-assisted Checkly management" by glueing together REST API calls and human judgment from outside the product. Pulling that work into Checkly itself would close a meaningful loop.

    What we would want such an agent to do, from most to least obvious:

    1. Understand the entire account state. Read every check, alert channel, dashboard, status page, group, snippet, env var, maintenance window, private location. Hold a mental model. Answer questions like "which checks page on-call via SMS?", "which targets are geo-blocked outside US/Canada?", "what would break if I downgrade to Hobby tomorrow?", "are there duplicate alert subscriptions across channels?".

    2. Bulk operations and structural changes via natural language. Setting up our account today took ~50 REST calls and a hand-rolled mapping of alert channel IDs across many checks. We would have liked to say: "create checks for this URL list at the right tier, wire them to these three channels, leave SMS off on the lower-priority ones, and clean up the onboarding-default checks Checkly auto-created" — and have the agent just do it, with a diff and approval gate.

    3. Generate Playwright / Browser scripts the user actually wants. Today we wrote a generic "find first visible button and click" Playwright script for 5 landing pages. A good agent would visit each page, identify the actual conversion CTA (not the video play button or the cookie-consent banner), and write a check that asserts a meaningful post-click state — modal opened, form submitted, URL changed to a thank-you page. The current "browser check" creation flow is "write code or paste a recording", with no help bridging the gap.

    4. Detect drift and anomalies proactively. Flag SSL certs clustering near expiry (we saw 7 in the 44–45-day range, all probably from the same Let's Encrypt batch), p99 outliers (we have a 17.5s peak on a high-frequency check that's silently passing), checks running from the wrong region, alert channels with no active subscriptions, dashboards filtering on tags that don't exist on any check.

    5. Suggest and apply tagging / naming conventions. Our 28 checks currently have zero tags — and Checkly's dashboard, group, and alert routing systems all key on tags. The agent should look at our URLs and propose a coherent tagging scheme (tier, purpose, location-policy, and so on), then apply with one approval.

    6. Plan changes before applying. Show a diff: "this change would deactivate 10 checks, modify 3 alert channels, create 1 dashboard, total cost delta -$X/month." Approval gate, then apply atomically with rollback. The Checkly CLI's import apply / import commit / deploy pipeline gestures at this but isn't conversational.

    7. Live-debug a failing run. "Why did this check fail at 12:14 UTC?" — pull the result, inspect the trace, surface the specific assertion that broke or the underlying network error, suggest a fix. Today this requires several manual UI clicks and remembering where to look.

    8. Connect to chat (Slack, Teams, Discord). When a check fires, the agent should be the responder, not just another notifier — answer "what does this alert mean? was it flaky last week? should we mute or escalate? show me the trace" right inside the alert thread, where the on-call already is.

    9. Plan-aware pre-flight. Today we learned through experimentation that Hobby has hard caps but grandfathered-resources are partly tolerated, while Trial behaves like Team for entitlements. The agent should know exactly what each plan unlocks and pre-flight every change against the active plan ("you're about to add a 30-second check from 5 regions; that requires Team or higher post-trial — here's what would happen at the trial-end downgrade").

    10. Compose with the user's connected ecosystem. If the user has Slack + Notion + GitHub connected (we use all three at G6), the agent should be able to open an incident in Notion with the failing check details, post to Slack with a runbook link, and check whether a deployment was the root cause — without the user wiring a webhook for each.

    Part 2 - API and Platform

    Smaller asks, every one of them grounded in friction we hit this week.

    Bugs:

    1. PUT /v1/dashboards/<id> returns HTTP 500 when changing customUrl. Tried partial body, full body, and multiple slug variants. Always HTTP 500 with no detail. Workaround: delete and recreate the dashboard. Repro: any account, any existing dashboard, any new slug.

    2. GET / PUT / DELETE /v1/dashboards/<numeric_id> does not work — you must use the alphanumeric dashboardId. The list endpoint returns both id (numeric, e.g. 1079188) and dashboardId (alphanumeric, e.g. 17dae609). DELETE on the numeric returns HTTP 404. PUT on the numeric sometimes silently no-ops. Inconsistent with /v1/checks/<id> and /v1/alert-channels/<id> where the same id works for everything.

    3. PUT semantics differ unpredictably across resources. /v1/checks/<id> accepts a partial body and merges fields. /v1/alert-channels/<id> rejects partial body with HTTP 400 ("type" is required). /v1/dashboards/<id> accepts partial body but appears to reset omitted fields to defaults under some shapes (we observed isPrivate: true reverting to false after a PUT that only modified the keys array). Three different behaviors on three resources is a footgun. Standardize on PATCH for merges and PUT for full replacement.

    API coverage gaps:

    1. No REST API for password-protected dashboards. The Checkly UI offers a "shared password" mode where viewers enter a key. The REST API and the CLI's Dashboard construct only expose isPrivate: boolean (Checkly account login). For our public status dashboard, we had to choose between fully public and Checkly-account-login-only — there was no way to set up "public with a shared password" without using the UI.

    2. No bulk operations. Creating 28 checks during initial setup required 28 sequential POSTs; later disabling 18 of them required 18 sequential PUTs. A PATCH /v1/checks accepting a list of {id, body} pairs with atomic semantics would save real round-trips and reduce partial-failure risk on bulk operations.

    3. No "trigger run now" REST endpoint. After creating or modifying a Browser check we wanted to verify the script ran end-to-end without waiting for the next scheduled run. The CLI's npx checkly trigger is great; surface it via REST too.

    4. No diff / audit endpoint. "What changed in this account since 2026-04-28T12:00Z?" requires inspecting updated_at on every resource individually. An audit-log API would cover migration verification, change review, and post-incident forensics.

    5. No dry-run for checkly deploy. We disciplined ourselves into a "do not deploy" invariant in our repo all day for fear of wiping the production account from an unsynced import preview. A --dry-run flag that shows the diff before applying would let us actually use deploy with confidence.

    6. Entitlements API is descriptive, not actionable. /v1/accounts/me/entitlements returns human-readable descriptions ("Maximum number of browser checks") but doesn't reliably surface the actual cap values, and doesn't pre-flight an attempted resource creation against the plan. Today we hit Hobby's caps by trial-and-error.

    Inconsistencies

    1. Inconsistent list-endpoint shapes. Most list endpoints return [...] arrays. /v1/status-pages returns {length, entries, nextId}. /v1/private-locations rejects ?limit= with HTTP 400 while every other list endpoint accepts it. /v1/environment-variables returns 404 while the same data is at /v1/variables. A consumer can't write one generic "list every resource type" function.

    2. CLI vs REST divergence on Browser check code. The CLI's BrowserCheck construct uses code.entrypoint (a file path that gets bundled at deploy time). The REST API takes a script field (an inline string). Same concept, two shapes, two code paths.

    3. checkly import and checkly init rely on TUI prompts. Won't work in CI, scripts, or sandboxed shells. We worked around by pre-creating the Checkly project record via direct REST POST /next/projects and passing explicit resource specifiers (check:<uuid>) to skip the interactive prompts. A --non-interactive flag with sensible flag-driven equivalents for every prompt would help.

    Wishlist

    1. Auto-issued dashboard keys are surprising. POST a dashboard, get back keys: [{... rawKey ...}]. Run a subsequent unrelated PUT and the keys silently disappear. Either don't auto-issue, or persist them as proper first-class resources.

    2. Document that DELETE returns HTTP 204 with empty body. Fine, but consumers (including the Python standard library) blow up trying to JSON-parse an empty response without explicit handling. A short note in the API docs would prevent the trap.

    3. Better error messages on the API. "An internal server error occurred" with no detail blocks debugging. Even a request ID we could quote when reporting back would help.

    4. In-product playbook for plan downgrades. Today we learned about Hobby's grandfathering-vs-rejection asymmetry empirically (existing over-quota resources are tolerated; new ones via API are rejected). A "what happens when I downgrade?" preview, with concrete predictions for our exact account, would prevent that learning from happening during an outage.

    Most of the above is small and concrete. The AI agent in Part 1 is the headliner, it would consolidate a lot of the manual operational thinking we've been doing outside the product and ground it in Checkly's own state. Happy to go deeper on any of these, share the request / response logs we built up across the day, or jump on a call if useful.

    Thanks for reading!

  • Azhan Khalid

    •

    Feb 20

    Hi! We need the ability to set custom subject names for our alert emails ourselves. Do you have plans to make this editable, and if so, when will it be available? Thanks

  • stephan.natis@elia.be

    •

    May 14, 2025

    Same request for us. Checkly is used to test our API data. Depending on the data, emails are sent to different people across our company. It's common these person come back to me to find out the error that caused the alert.

    We would like the reason for the error to be included in the body of the email.

  • Susanne Tünker

    •

    Feb 27, 2025

    Another relevant use case: Companies managing multiple clients and projects struggle to differentiate alert emails. A helpful improvement would be the ability to modify subject lines, such as adding the project name at the beginning. For example: [Checkly/Project Name] Failure Alert.

  • Sven Müller

    •

    Sep 4, 2024

    Same request here. We would like to be able to fully customize the email body (similar to davanced webook request body)

    Use case:

    • we are monitoring a 3rd-party service (used internally)

    • in case of a failed API check, we want to sent an email to the support address (3rd-party service) to open a ticket/incident on their side

    • therefor we would like to customize the email content

  • spr@grabango.com

    •

    May 6, 2024

    I'd like to be able to add the failing check's response body to the email. And potentially remove links to checkly, so I can send emails to external partners automatically (whereas right now I have to manually format an email that's appropriate for them).