Operator alerts
This page lives at Staff → Operator alerts and requires a staff claim. Anyone on staff can read it; changing a setting or sending a test takes the super staff role, and every change is written to the admin audit log.
An operator alert is an event whoever runs the install has to hear about: a card dispute lost on a subscription, a data erasure that failed, a health check that went red, a webhook failing its signature. Each one is a row in one registry, and every row is delivered the same way, whether the install is Aglyn's own cloud, a self-hosted deployment or an application built on the open-source packages.
Where alerts go
Every alert does three things, in order:
- Writes a console notification for every staff member, in the bell. This always happens, whatever the settings below say.
- Emails the operator, as the Operator alert system email, which you can
redesign on Staff → System emails. The address is
STAFF_ALERT_EMAIL; when that is unset, the operator support address (NEXT_PUBLIC_OPERATOR_SUPPORT_EMAIL); when that is unset too, every staff account's own address. A preview deployment never falls back past the first. - Posts to the out-of-band webhook, when
OPERATOR_ALERT_WEBHOOK_URLis set. The body is JSON with a Slack-compatibletextfield, so a Slack incoming webhook works as is. WithOPERATOR_ALERT_WEBHOOK_SECRETset, each post carriesx-operator-alert-timestampandx-operator-alert-signature:sha256=followed by the hex HMAC-SHA256 of{timestamp}.{body}under the secret.
The webhook exists for the failures email cannot report: the mail provider down, or rejecting its key. Without it, those alerts reach only the console bell.
The page shows which of the three email sources is in use and whether the webhook is
set. Send test on any row sends one [Test] alert of that type to the inbox and
the webhook, past every switch, so you can prove both reach you.
Switching alerts
Every row has two controls:
- On: off keeps the console notification and sends no email and no webhook post.
- Delivery: Immediate sends it as it happens. Daily digest batches it into one email and one webhook post a day, sent after the hour you choose at the top of the page (UTC).
Until you change them, the coded defaults apply:
| Tier | Default |
|---|---|
| Must know | On, immediate. Money moving with nobody looking, a legal duty, data not erased. |
| Should know | On, and immediate unless the row says otherwise. New support tickets and replies go to the digest. |
| Routine | On, in the digest. Listings waiting for review, resources missing a sharing scope. |
A row you set back to its default is removed from the stored settings, so the page only ever holds the answers staff actually gave.
In the bell, the tier sets the color. A must-know alert is red, a should-know alert amber, and a routine one blue (see how urgent each notification is). A few alerts name their own color because their tone and importance differ. A degraded health check is a must-know, but it is amber because degraded is not down. A recovery is green. Support tickets are blue. The color changes nothing about delivery: the switches above decide that.
Repeats are told once. Each alert names what makes two raises the same event (a dispute, a workspace and month, a sending domain), and a repeat inside the alert's window is counted rather than sent again. The next alert that does go out says how many it held back. The endpoints anyone on the internet can reach (a webhook failing its signature) alert only once the failures recur.
The alerts
Plugins add their own rows: an install without the commerce plugin shows no commerce alerts.
| Area | Alert | Tier |
|---|---|---|
| Security | Urgent abuse, fraud or risk alert: phishing, card testing, a flagged payment, a held phishing email | Must |
| Security | An automatic security hold on a new workspace that published a phishing page | Must |
| Billing | Card dispute with no owner; billing webhook half applied | Must |
| Billing | Subscription dispute lost; Stripe billing a workspace that does not exist; a closed month's metered usage not reported to Stripe | Must |
| Billing | Billing webhook failing its signature check; a delivery that moved nothing | Must |
| Data protection | A person's erasure failed; a workspace's erasure failed; the database export failed | Must |
| Legal | DMCA counter-notice; a DMCA takedown, impersonation or illegal-content report | Must |
| Operations | A health check went degraded; a published site not rendering pages; the render monitor cannot see a site; the production canary found a deploy not rendering pages, rolled back or not | Must |
| Operations | A health check recovered; a published site rendering again; a reaper stuck; a plugin job, consent group change, publish outbox or sending-domain provisioning failing; the bandwidth ceiling reached | Should |
| Billing | Automatic billing lock; an invoice voided or uncollectible; a connected account's payout or transfer failed; a workspace's payment failed | Should |
| Security | Staff access granted or a role raised; a plugin verifier regression | Should |
| Support | A ticket past its response time; new tickets and replies (digest) | Should |
| Email deliverability | Send-rate governor saturated or unavailable; a recipient gateway blocking a shared domain; the email provider rejecting the key or the shared pool unhealthy | Should |
| Payments (commerce) | A chargeback that could not be routed; sales tax not reversed; a seller's share not reversed; a platform fee correction refused | Must |
| Email deliverability (marketing) | The events webhook failing, unconfigured or rejecting signatures, so bounces go unsuppressed | Must |
| Email deliverability (marketing) | Campaigns blocked by the reputation breaker | Should |
| Operations (AI) | The platform's AI provider refusing its key or out of credit | Should |
| Routine | A plugin listing waiting for review; resources missing a sharing scope | Routine |
These stay in the console bell and never email: new subscriptions, cancellations, plan changes, sign-ups, new workspaces, impersonation, and successful backups.
Health checks alert on their own
Every health endpoint remembers its last verdict. When one goes from healthy to degraded, the Health check degraded alert goes out once. When it comes back, Health check recovered goes out once. A check that flaps across the line inside an hour is told once, not on every swing. The page lists each check's current state and since when.
Nothing has to be watching for this to work. The operator alerts tick
(/api/admin/operator-alerts/tick, every 15 minutes) asks every health endpoint on
the console's own origin. It also asks the published-site runtime's endpoints when
OPERATOR_HEALTH_TENANT_ORIGIN names one of your sites. The same tick flags support
tickets past their response time and sends the daily digest.
If the scheduler stops, the tick stops with it, and so does every alert that depends
on the sweep. Point an external uptime monitor at /api/health/crons: it records its
own state on every read, and the tick is one of the jobs it watches. The webhook is
the other safety net: a channel that normally hears something every day and goes
quiet is a signal too.
The render monitor
A published site can look up while it renders nothing new. Pages a visitor has already loaded come out of a cache, so the home page keeps answering while every page that has to be drawn fresh fails or hangs. Health checks don't catch that either: they answer from inside the runtime and never draw a page.
The render monitor (/api/admin/render-monitor, every 5 minutes) asks each
site it watches for two pages no cache can hold, over the internet, from the
console:
/search, which every site has and which is drawn on every request, inside the site's full layout.- A path nobody has asked for before, which the site answers with its own "page not found" page, drawn in the same layout as every published page.
Those two draw the layout but not a page's own content, and a fault can hide
there: on 2026-10-05 every page with an image hung while the search page and
the "page not found" page drew fine. So a site can also be given real pages to
draw fresh. Each run draws one of them (taking turns), asked for under a spelling
of the site's name no cache has seen, through a deployment address of your
published-site runtime that only your scheduled jobs can reach. Name the pages
with RENDER_MONITOR_PAGES and that address with RENDER_MONITOR_RENDER_ORIGIN;
without both, no real page is asked for.
A run passes when every page comes back complete within 25 seconds. When a site fails two runs in a row, Published site not rendering pages goes out once, naming what failed (a timeout, an error status, an unfinished page). The first passing run after that sends Published site rendering again. Each watched site is listed among the health checks on this page with its state and since when.
A run where bot protection answered instead of the site has seen nothing. It
does not count as a failure and never sends Published site not rendering
pages. Two such runs in a row send Render monitor cannot see a site once,
which names the setup to fix: AGLYN_PROBE_TOKEN on the console, and the
firewall rule on your published-site runtime that admits it. The site keeps the
last state the monitor actually saw.
The same verdicts are public at /api/health/pages on the console, per site with
?site=<host>, so an external uptime monitor can watch real pages without
carrying a firewall bypass. It answers 503 only when a site is not rendering,
when a site it was asked about is not watched, or when the verdicts cannot be
read.
It watches your demonstration site by default. Add your own sites, or your
marketing site, with RENDER_MONITOR_ORIGINS; see
Self-hosting environment.
A GET to the route fetches and grades the pages and records nothing, so you can
ask what the monitor would see right now without moving its state. The scheduled
POST records and alerts.
Self-hosting
Nothing here is specific to Aglyn's cloud. The environment variables are listed in
Self-hosting environment. Schedule
/api/admin/operator-alerts/tick every 15 minutes and
/api/admin/render-monitor every 5 minutes with the rest of your
scheduled jobs.