Lockdown (the panic button)
Lockdown controls live in the staff console at Staff → Lockdown and require a staff claim; locking or lifting anything requires the super staff role. Every action — lock and unlock — writes an audit row.
Lockdown is the control you reach for when something has gone wrong: a compromised site, an account being abused, a billing suspension that has run its course, or a maintenance window that needs the doors closed. One mechanism, five scopes:
| Scope | What it covers | Where the state lives |
|---|---|---|
| Platform | Everyone except staff | lockdowns/platform |
| Workspace (org) | The org's console access (writes), all of its sites | orgs/{id}.suspendedAt family |
| Site (host) | One published site | hosts/{id}.suspendedAt family |
| Custom domain | One attached domain name; the site keeps serving elsewhere | lockdowns/domain--{hostname} |
| Account (user) | One person | lockdowns/user--{uid} + Firebase Auth disabled |
Precedence is platform → org → host → domain → user: the widest active scope decides the notice a person sees.
What a lockdown does
A lockdown is enforced server-side at the chokepoints, not hidden in the UI:
- Console sessions — the session mint and the cross-subdomain exchange refuse locked users with HTTP 423 and clear their session cookies. User-scope locks also disable the Firebase account and revoke its refresh tokens, so "logged out" means logged out.
- Published sites — the tenant middleware checks every request before the
page cache, so a taken-down site serves a real 503 notice (with
Retry-After) immediately — cached pages are also evicted at lock time. - APIs — org-scoped API routes refuse with
423 { "error": "locked", "reason": … }, so an API consumer sees suspended, not a mystery 403. - Outbound email — a full org, site or platform lock stops the mail a
workspace sends. The site's sending identity refuses, so
sendEmailanswerssuspendedfor workflow and automation steps, CRM one-off mail, inbox replies, member posts, receipts and reminders. Event runs (workflows, actions, inbound webhooks, which answer423) do not start at all. Campaigns, including scheduled sends and batches already in progress, are deferred, not failed, and outreach sequences for the site are held. So lifting the lock resumes them. A read-only lock pauses campaigns and sequences. The engine and the sending identity let it through, because an order paid before the window began still owes its receipt. The doors that raise events still apply their own read-only refusals (see below).
Reasons and the notice
Every lock carries a reason — security, abuse, billing, maintenance, or
manual — which picks the notice the locked-out person sees. An optional custom
message replaces the notice body (it is shown to customers — keep internal
rationale in the audit note, not here). billing notices point at billing
settings; security/abuse/manual notices point at support@aglyn.com;
maintenance shows the window when one is set.
abuse is for phishing, fraud or malicious content, and it is at least as strict
as security everywhere. Sessions are revoked, media stops serving, download
tokens rotate, and billing cancellation and paused site money are ticked by
default. On an account it is a permanent ban. The account is kept only so
its address can never sign up again, and after the lock notice it is sent nothing
at all — see A ban stops all mail. Its notice says the account is
closed for a Terms of Service violation, and never says which one.
Nothing places abuse automatically. The one lock the platform places by itself
for content is the page screen's automatic security hold: a held phishing page
from a workspace less than 14 days old locks that workspace, its site (as a
takedown) and the publishing account, all as security. Its audit rows name the
actor system:page-screen. Confirm it by re-placing the workspace and account locks
as abuse here, or lift all three if it is a false positive. See
the automatic security hold.
Modes: full, or read-only
Every lock is armed in one of two modes. The dropdown sits beside the reason on the platform card and the workspace/site card.
| Full (the default) | Read-only | |
|---|---|---|
| Customer sites | 503 notice, cached pages evicted | keep serving normally |
| Visitor forms, cart, checkout | refused | refused, with an inline "temporarily paused" |
| Visitor analytics beacon | frozen entirely | page views still count, host automations do not fire (why) |
| Console reads (sign-in, viewing) | refused | work |
| Console/API writes | refused | refused with 423 |
| Member sessions | revoked (security/manual) | never revoked |
| Staff | bypass everything | bypass everything |
Reach for read-only whenever the reason is our own maintenance — a schema migration, a data repair, a suspected-corruption investigation. Those need the writes frozen so nothing races the repair; they do not need the customer's shop taken off the air, and taking it off the air costs them real money for our convenience. Full lockdown is for takedowns: abuse, compromise, a workspace that must stop existing publicly right now.
Read-only is available on the platform, workspace (org) and site
(host) scopes. It is refused on the user, feature and domain
scopes, because none has a milder setting to offer — a user lock's teeth are
the Firebase account disable and token revoke, every feature key already names
a single write, and a domain lock stops serving one name while read-only is
defined as continuing to serve it. The route answers 400 rather than silently
arming a full lock, and the card hides the mode dropdown on those scopes rather
than offering a choice that will be rejected.
domain scope used to accept it, and liedUntil AGL-1621 a domain lock armed with mode: "read-only" returned 200
and was stored as a full takedown. The scope arrived after read-only mode
did, was never added to the refusal list, and its carrier write never gained the
line that persists the field — so the mode was dropped on the floor and the
document read back as full, because absent means full. An operator asked for
the lighter action, was told it worked, and a whole custom domain went dark.
Persisting the field would have been the wrong repair. Read-only has no enforcement surface at that scope at all: every path that resolves a domain lock is a path that serves (the edge verdict, the page loader, the locked notice), and neither write gate resolves the domain scope — both are keyed by host id, a domain lock by hostname, and a site can carry several attached names. A stored read-only domain lock would have refused nothing anywhere while this console reported LOCKED.
Staff writes bypass read-only exactly as they bypass everything else, which is the whole point: you perform the migration while the world keeps reading.
What "reads keep working" does and does not cover
A request is classified by what it DOES, not by how it looks. Most reads are
GET and pass automatically. A handful of console operations are queries that
send their arguments in a POST body — where an asset or component is used, a
plugin's impact, signing a media URL, minting a presence token, signing in —
and each of those is declared a read individually, in the route, with the
reason written next to it.
Anything not declared refuses, deliberately: a chokepoint nobody has audited is treated as a write, because an over-refused read costs a customer some friction and an under-refused write costs the data the freeze exists to protect. One consequence worth knowing before a customer reports it:
- the tenant edit bar stops appearing on published sites for the locked workspace. That is intended — the bar leads to an editor whose saves the freeze denies anyway.
That is a recorded decision rather than an oversight. If you find another operation that only reads and still 423s during a window, it is worth filing — that is how the declared list grows.
A declaration is usually per route, but it does not have to be. The media folder-sharing preview — the "also apply to the 47 files in this folder and its subfolders?" count the library quotes before a sharing change — is declared per request, because it shares a route with five actions that write and with the cascade it previews. During a read-only window that count still answers; the cascade it precedes, and every other folder operation, still refuses. An author therefore sees the size of the change they cannot yet make, which is the question they were actually asking.
How fast read-only takes hold
The two halves converge at very different speeds, and the difference is the opposite of what most people assume. Measured against the emulator and a real production-mode tenant (AGL-1626 — the numbers below are observed responses, not derived):
| Observed | |
|---|---|
| First visitor write after arming | 423 on the very first request — 32–86 ms across four runs |
/api/lockdown-verdict and the staff probe | up to ~60 s (32 s against a cold cache, 60 s against a warm one) |
| Customer pages | unchanged — 10 samples over 100 s, all 200 with content |
The freeze is immediate; the view of the freeze lags. Visitor write chokepoints read the workspace and site records live on every request, so a migration may start as soon as the lock is armed — there is no window in which the console says "locked" and writes are still landing. What lags is the verdict route the staff panel and the tenant middleware read, which is cached for about a minute. So during the first minute the panel may still say a workspace is unlocked while its customers' forms are already being refused. That is the safe direction, but it will confuse you if you are watching the panel to decide when to begin.
The same minute applies to a full lock's 503, plus one more effect worth
knowing: a page rendered while a full lock was in force is the 503 notice, and
it is cached like any other page. Lifting through the staff surface or
/api/admin/lockdown clears those pages as part of the lift. Editing the
workspace record directly in Firestore does not — the site keeps answering 503
from cache until the pages regenerate on their own.
What read-only has been proved against
Read-only shipped with unit coverage at every layer and nothing observed on the
wire, which is a weak proof for this particular mode: a cached page still
serving is indistinguishable from a lock that never engaged. AGL-1626 closed
that with npm run e2e:lockdown:readonly — the emulator, a next build /
next start tenant, and a real refusal captured for each branch:
- the site keeps serving —
/homereturned200with its content on every sample across 100 seconds of an armed read-only workspace lock, with no rewrite to the maintenance notice; - a visitor form answered
423with"Temporarily paused"and "Nothing you typed has been lost", carrying no support address — support belongs to the site owner, not to us — and spent nothing: no submission stored, no change to the month's counter; - the note you type stays behind the door. The same lock's staff message ("Scheduled data migration") appeared in the console refusal and in none of the visitor ones, which carried only the pause copy. Whatever you write in that box is for the account holder; a stranger on their site never sees it;
- the basket answered
423with "Basket changes are paused… Browsing works as normal"; - checkout answered
423with its own title and the promise no generic copy can make — "this is not a payment problem and you have not been charged" — refused before the handler, so no payment session is created; - strictness outranks width (below) was forced rather than reasoned about;
- expiry restored writes with no staff action: refused while the window was open, accepted once it passed.
In the console, with the same lock armed:
- a customer's edit answered
423, while the same customer's usage scan ("what is this asset used by") went through — the read that would otherwise push someone to delete on a guess; - a staff account performed the identical edit successfully, which is the whole point of the mode;
- under a platform read-only lock a customer could still sign in (the
mint is a read) and their first edit afterwards answered
423naming the platform scope.
Re-run it after any change to the chokepoints. It needs the emulators, the e2e
seed, and port 4500 free — the tenant only recognizes localhost:4500
locally.
The one row that is not a wire observation: "never revoked"
Every other claim in the mode table above was forced on the wire. "Member sessions — never revoked" was not, and this is what stands behind it instead (AGL-1724).
It resisted the harness for a structural reason worth writing down, because the
same reason will defeat the next attempt: lockdown-readonly-wire.mjs arms its
locks by writing the suspended* carrier directly, not by POSTing to
/api/admin/lockdown. That is deliberate — it keeps the harness independent of
a running console — but revocation lives in applyOrgLockdown, which a direct
carrier write never reaches. A "the session still works" probe added to that
harness as it stands would pass against a revocation path that had been deleted
entirely. Forcing it for real needs the emulated console up, a super-staff
token, a minted member session cookie held across the arming, and the paired
mode: 'full' run to prove the probe can fail at all.
What it rests on instead is apps/console/specs/lockdown-revocation-wiring.spec .ts, which drives the real applyOrgLockdown — the module the three other
lockdown specs all mock away, and which until then had no test of its own. It
pins both directions against the same reason (security, the reason that does
revoke under a full lock, so the discrimination is on the mode): read-only
revokes nothing, full revokes both members pool-aware, and neither ever revokes
a staff account on the roster. Each assertion was confirmed to fail against a
deliberately broken copy of the module.
And the reason a surviving session is safe rather than merely intended: the
write freeze is enforced in two places that have nothing to do with the session.
A read-only lock writes orgSuspended: true onto every member doc exactly as a
full lock does, and orgNotSuspended() in cloud/firebase-firestore.rules
gates every client-direct write on it — besigner saves included. Admin-SDK
writes are refused separately, at the wired chokepoints
(lockdown-423-coverage.spec.ts). So a member who keeps browsing under
read-only holds a read capability, not a write one, and revoking the session
would buy no enforcement — it would only sign the workspace out during our
maintenance window, which is the outcome the mode exists to avoid.
The corollary, and it has bitten once. Because that projection is written for both modes, it answers "is this workspace locked at all" and cannot answer "does this lock refuse this request". A chokepoint that consults the projection instead of the verdict is therefore mode-blind, and it will refuse a read the verdict just passed. That is what 404'd every private media preview in the console under a read-only lock (AGL-1790): the route declared its read, the verdict honored it, and a legacy line beside the verdict refused on the projection alone. The projection is still a fine signal — it is how these routes avoid an org-doc read on the happy path — and it still refuses when it disagrees with the workspace document, which is a stale projection rather than a mode. It is just never the answer. If a read 423s or 404s during a window and the mode table says it should work, this is the first shape to look for.
A gentler lock never softens a stricter one
This is the highest-consequence rule in the feature, so it is worth stating on its own: if a platform-wide read-only maintenance window is running and one workspace is under a full security takedown, the takedown wins. The wider, gentler lock does not readmit that workspace's visitors.
Forced on the wire (AGL-1626) with both armed at once: the verdict reported
"mode":"full","reason":"security" — the workspace's lock, not the platform's
— and the site answered 503 with Retry-After. Arming a maintenance window
across the platform can never quietly reopen a site staff has taken down.
To arm one from a terminal, add mode to the usual body:
curl -X POST https://app.aglyn.com/api/admin/lockdown \
-H "Authorization: Bearer $ID_TOKEN" -H 'Content-Type: application/json' \
-d '{"action":"lock","scope":"org","targetId":"ORG_ID",
"mode":"read-only","reason":"maintenance",
"untilMs":1786695133044}'
mode is optional and defaults to full, so every existing runbook command and
saved script keeps doing exactly what it did before.
Standard or takedown: what happens if we cannot reach the database
Every lock carries a second, independent choice, on the console as "If Aglyn can't reach the database":
| Choice | Stored as | If the lockdown record cannot be read |
|---|---|---|
| Release — standard lock (default) | nothing is stored | The lock stops being enforced |
| Keep holding — takedown | enforcement: 'takedown' | The lock keeps being enforced |
Lockdown normally fails open: if Firestore is unreachable, the verdict is "not locked". That is deliberate and stays the default — a database blip must not take every customer site down at once.
A legal or abuse takedown has the opposite cost. "We kept serving it because our database was down" answers no court order and no CSAM report. So a takedown holds through the outage, and only a takedown does.
Choose takedown only for a legal or abuse order — a court order, a DMCA
takedown you have accepted, a CSAM or malware removal, a domain hijack or
dispute. Everything else is a standard lock: maintenance windows, billing
suspensions, precautionary holds, and incident response, including the
security ones. Getting this wrong takes sites down for a reason unrelated to
the incident that took the database out.
Three properties worth knowing before you rely on it:
- It is never inferred. Neither the scope nor the reason implies the
class — a
securitylock is a standard lock unless you say otherwise, and any of the six scopes can be either. A lock is fail-closed only because somebody chose it, and the audit row records who and when. - It is not retroactive, and it is not magic. Enforcement holds on a server process that has already seen the takedown. A process that starts up during the outage has never read the record, has nothing to hold, and fails open like everything else — so does a takedown you try to place while the database is already down. In practice the case this covers is the real one: an order placed hours or days ago, and a blip today.
- The expiry still wins. A takedown with an until time still releases on schedule, even mid-outage. Classifying a lock does not make it un-liftable.
A takedown cannot be read-only, and the route refuses the combination: a
read-only lock keeps serving the content and only refuses writes, which is the
opposite of what a takedown is for.
From a terminal, add enforcement to the usual body:
curl -X POST https://app.aglyn.com/api/admin/lockdown \
-H "Authorization: Bearer $ID_TOKEN" -H 'Content-Type: application/json' \
-d '{"action":"lock","scope":"domain","targetId":"seized.example",
"enforcement":"takedown","reason":"security",
"message":"This domain is subject to a dispute."}'
enforcement is optional and defaults to standard, so every existing runbook
command and saved script keeps failing open exactly as before. A value the
server does not recognize is rejected rather than defaulted, in either
direction — the response echoes the class back as verified.enforcement so you
can confirm the one you meant is the one that landed.
Maintenance windows and expiry
A lock may carry an until time. When it passes, the lockdown simply stops — access restores with no staff action and no write. Use it for maintenance windows; leave it empty for anything that should stay locked until a person lifts it.
Who keeps access: the un-panic invariant
Staff are never locked out, by any scope, ever. A platform-wide lockdown leaves every verified staff session able to reach the staff console and lift it — this is spec-enforced (a panic button that panics its own operator is worse than none). For the same reason, a staff account cannot be user-locked: revoke its staff claim first if it truly must go.
Feature scope
The beta-week abuse-response kit: kill one capability platform-wide while
everything else keeps serving. Feature locks live in the same lockdowns
collection (feature--{key} docs), use the same reasons/messages/expiry, the
same audited writer, and appear as a checklist on the same Staff → Lockdown
page.
| Feature key | What it stops | Reach for it when |
|---|---|---|
signups | New account creation (all four doors — the signup form, both Google flows, and the sign-in page's new-account bounce). Accounts created after the lock began are refused a session; every existing account signs in untouched. With the blocking function registered (see below) the Auth records are refused too, so nothing is created at all. | Bot registration wave, free-tier abuse storm |
uploads | New media bytes (upload, signed-URL upload, replace). Browsing, organizing, restoring, and serving existing media all keep working. | Malware/abuse report in the DAM |
checkout | New Stripe checkout sessions only — console plan upgrades and marketplace purchases. Existing subscriptions, invoices, and the pay-your-way-out path for billing-locked orgs are untouched, and the notice says explicitly that it is not a payment failure. | Stripe integration bug mid-charge |
marketplace-installs | Installing anything from the marketplace (all artifact kinds, including re-copying an updated artifact). Everything already installed keeps working; publishing, reviews, and abuse reports stay open. | A malicious listing slips review (the per-plugin kill switch takes out one listing; this is the wider valve) |
ai-assist | The AI assist endpoint. The switch works even while the feature is unconfigured — it predates the API key on purpose. | Provider incident, cost runaway |
ai-generate | The generative doors (sections, pages and automations written by a model), behind the release_ai_generative flag. Separate from ai-assist because generation spends at a different rate and an incident on one need not stop the other. | Provider incident, cost runaway on generation, an abuse wave on the Free AI taste — this is the first response, ahead of any per-account measure |
The Free taste and this key. Free workspaces carry 300 AI credits a month
with no invoice behind them, and a platform-wide daily ceiling on free-tier
spend pauses free generation on its own when the sum of a day's free spend
reaches AI_FREE_DAILY_PLATFORM_CEILING_USD (staff are mailed at 80%, the
assist signals page shows the day so far). That pause is automatic and
Free-only. ai-generate is the manual lever for anything the ceiling has not
caught yet — a wave that is spending fast but is still under the day's ceiling,
or generated content that is abusive rather than expensive — and it stops
generation for every plan, paid included, so it is the wider valve: pull it
first, then read the signals page to see which workspaces drove the spend.
Composition, not ranking: a platform lock implies every feature; a feature lock implies nothing about the platform, workspace, site, or account scopes.
Pausing AI for one workspace
A feature lock can also be scoped to one workspace: the same lockdowns
carrier at feature--{key}--org--{orgId}, written by the same route with an
orgId in the body, audited the same way, and read only by the AI doors that
name that org. The staff org page offers it as Pause AI in its Staff actions,
which writes both ai-assist and ai-generate for the org in one request, and
Resume AI lifts both. A request may name several levers in targetIds
instead of one targetId: each is written and audited as if asked for alone,
and the workspace's owners get one email naming every lever by its customer
name ("AI assist and AI generation"), never the checklist label.
Reach for it to stop a customer's AI spend without touching what they bought:
the plan, the add-on and every entitlement override stay exactly as they are, so
resuming restores them untouched — unlike forcing aiAssist off in the override
dialog, which changes the entitlement record and has to be remembered and undone.
Members of the paused workspace see the ordinary feature-pause notice on every AI
request; every other workspace is unaffected; staff calls still pass, to verify
the pause. The chip AI paused on the org page reflects the state the route
reads back.
Confirm weight: feature locks do not require the type-to-confirm phrase. The platform phrase exists because one request can take everything down; a feature lock is one named capability with the platform still serving — the same blast-radius class as an org or site lock, and incident response wants the narrow lever fast.
Staff bypass, per feature: staff keep uploads, marketplace-installs,
ai-assist and ai-generate through a lock — responding staff need to upload
a test file, reproduce an install, or make one AI call or generation to verify
the fix before lifting it.
checkout grants no staff bypass: a staff-created checkout session is
still a real charge, and verification belongs in Stripe test mode. signups
is decided by account age, not claims — there is no bypass to grant.
signups also refuses account CREATION — if the valve is armed
Everything else on this page is enforced by code that ships with every deploy.
signups is the exception, and the difference is worth understanding before
you need it.
The lock has always refused the session mint, the legal-acceptance recorder, and the signup-page doors. That makes a wave's accounts unusable — but account creation itself is client → Firebase Auth, with no Aglyn server in front of it, so the Auth records were still being created: unusable, but accumulating in the pool and against the Auth quotas.
Refusing creation needs a Firebase Auth beforeUserCreated blocking
function, which lives in cloud/functions and is registered in Identity
Platform, not in this repo. Two consequences:
- Merging does not deploy it. It ships with
firebase deploy --only functions, and thebeforeCreatetrigger must then show up in the Identity Platform blocking-functions config. - The switch looks identical either way. So the Lockdown page reads the
Identity Platform config on load and states, under the
signupsrow, which world you are in: "Account creation is REFUSED too", "Account creation is NOT refused", or "UNKNOWN". Unknown is never rendered as armed — if the page cannot confirm the valve, treat the lock as sessions-only.
It fails OPEN when it cannot read the lever, and keeps holding one it has
already seen. If the function cannot read lockdowns/feature--signups —
Firestore unreachable, or the read exceeding its 2.5 s budget — the account is
admitted, and the function logs signups lock unreadable at creation, account admitted. The reasoning: an unreadable lock means the platform does
not know whether the lever is pulled, and this is the only gate in front of
account creation, so refusing on "don't know" turns away every signup for as
long as reads are slow — which, where signup traffic is light enough that the
instance is cold on most attempts, is every signup. The opposite error is
bounded: an account created while Firestore is unreadable cannot finish
signing up anyway (the profile, the acceptance record and the workspace all
live in Firestore), so it buys a removable orphan Auth record rather than a
working account.
That covers not knowing. A lever the running instance has actually seen
pulled is a different fact and survives a later failed read: every read that
completes is recorded in the instance, and a read that then fails is answered
from that record (logged with cause: "held"). So the obvious objection — a
bot wave makes reads fail and thereby releases the brake aimed at the wave —
needs the wave to land on an instance that has never once read the lock.
Lifting the lever is itself a Firestore write, so nothing can lift it during
that outage either, and a lock's own expiry is still honored, so a dead-man
untilMs cannot become un-liftable by being remembered.
This is stricter than the tenant takedown ledger below, which remembers
only locks classified takedown: here any active lock is remembered,
because holding one refuses new signups rather than keeping every visitor off
a customer's whole site.
What it does not do, stated rather than implied: an instance that has never completed a read has nothing to remember, so a lever pulled during a total Firestore outage does not reach a cold one. Closing that needs a carrier more available than Firestore.
The cold read is warmed. The first Firestore read in a container is
initializeApp, a token fetch and a fresh gRPC channel, and on a cold
instance that is most of what the read costs. The function starts that read at
module scope — gated on its own FUNCTION_TARGET, so the every-minute job
beat and the deploy-time trigger scan do not pay for it — and the first
account creation consumes the result instead of paying for it inside the
2.5 s budget.
If it ever needs to be taken out of the path in a hurry: unregister the
beforeCreate trigger in the Identity Platform console. No deploy, no code
change, and it is the same console you are already in.
It never touches sign-in. There is deliberately no beforeUserSignedIn
sibling — that one fires for existing accounts and would put every sign-in,
including the permanent break-glass account, behind this read at all. The lock
stops accounts being born; it never stops one coming home.
All three doors, both pools. Email/password, Google and SSO all end at
Firebase Auth account creation, and the handler reads no provider, no email
and no tenantId — so SSO's per-org GCIP tenant pool is refused on the same
terms as the project pool.
Expiry works the same as every scope: when the optional end time passes, the feature restores itself with no staff action and no write.
What a customer sees. Every console surface a feature lock can refuse — the billing upgrade buttons, every marketplace install and purchase button, the theme and template installers, and both AI-assist doors — renders the lock's own notice rather than a generic failure toast. So a checkout lock reads "Checkout is temporarily unavailable — this is not a payment failure, and your account, subscription, and sites are unaffected", and an installs lock reads "installs are paused; everything already installed keeps working". This matters for the message you type: it replaces the body of that notice, so write it for the customer, not for the incident channel. An end time is restated in the reader's own local time; without one, no return time is promised. Genuine failures are untouched — a real error still shows a real error.
Custom-domain scope — one name, not the site
When the problem is the domain rather than the content — an ownership
dispute, a hijacked or lapsed registration now pointing at us, or a trademark
complaint about the name itself — a host takedown is disproportionate. The
customer's site is fine; the name is the problem. This scope locks one
attached domain while the same site keeps serving on its *.aglyn.app
subdomain.
curl -X POST https://console.aglyn.com/api/admin/lockdown \
-H "Authorization: Bearer $ID_TOKEN" \
-d '{"action":"lock","scope":"domain","targetId":"acme.com",
"reason":"security","message":"Pending registrar review."}'
Three things about it are worth knowing before you use it:
It is keyed on the NAME, not on the site. Every other narrow scope keys on the thing it locks. This one cannot, because the incidents it exists for are exactly the ones where the domain moves — a disputed name gets detached and re-attached, sometimes to a different workspace. Keying on the hostname means the lock follows the name, survives a detach and re-attach, and can be placed on a domain that is currently attached to nothing at all, which is the state a dispute is usually resolved in.
The notice never names the address that still works. The site is still up on its platform subdomain, and every other scope's copy would happily be helpful about that. Here it must not be: the person reading the notice may be precisely the party the site is being withheld from, and "try acme.aglyn.app instead" would lift the lock in one sentence. The copy also does not say whose the name is — a dispute is exactly the case where we do not know, and a notice is a publication.
A platform subdomain cannot be domain-locked. {sub}.aglyn.app is our own
name, and the tenant resolves that space by subdomain without ever consulting
this scope — so a lock placed there would write a document no reader looks at.
The route refuses it rather than accepting a control that silently does
nothing. To take a platform subdomain down, use the Site (host) scope.
Expiry and modes work as everywhere else. Read-only is accepted but rarely what you want here: a name dispute is a full refusal or nothing.
One device, not the account
The narrowest thing on this page, and the only one that is not a lockdown scope. It lives on the user's detail page in Users admin rather than on the lockdown route, and it is here because it is what you reach for instead of the user scope when the incident is "someone stole my laptop".
The user scope disables the account. That is right for a compromised or abusive account, and wrong for a compromised device: the person calling has done nothing wrong and still needs to work. Sign-in history → Sign out ends the sessions without disabling anything — see Sign one device out for how to operate it and what to say.
What it actually guarantees, because it is easy to overstate in both directions. The refresh-token revocation underneath is account-wide — Firebase offers nothing narrower — so every device signs out once. The per-device part is the refusal afterwards: the signed-out device carries a revocation stamp and is refused at the session boundary every time it comes back, because it cannot produce a fresh authentication. The account holder signs in again and keeps working. So: everyone signs out once, they sign back in, that device does not.
Why it is not true single-session revocation, decided rather than skipped (AGL-1513). Dropping the account-wide revocation would leave the other sessions untouched, which sounds like the better product. It was measured and rejected: the per-device stamp is enforced at the session mint and the cross-subdomain exchange, and a browser that is already signed in passes through neither. Without the account-wide revoke it would keep minting fresh ID tokens from its own refresh token and keep opening every Bearer-token console API route — not for an hour, indefinitely. That is not a narrower control, it is no control. Making it real would mean carrying the device identity into the server-wide revocation check that every API door already runs, which is a change worth its own issue and is not this one.
The residual, stated plainly. Anything that goes through our servers stops within about fifteen seconds (the cached revocation epoch). A tab already open on the signed-out device can keep reaching the database directly until its ID token expires, up to an hour, because security rules key on that token and not on our cookie. That residual is not read-only: the rules carry no assertion about revocation, so the tab keeps every client write the account already had — publishing, content edits, media metadata, presence and co-editing. Object storage is the one exception, because it denies the client outright. The tab cannot obtain another token.
Asset quarantine — one file, not the site that serves it
When the problem is one uploaded file — malware in a PDF, an abusive image, a DMCA-noticed asset — locking the host punishes a customer for one object. Quarantine is the proportionate lever: the CDN refuses that file worldwide while everything else in the workspace keeps serving.
It is reversible, and that is the whole reason it exists instead of deletion: a false-positive scan or a successful counter-notice is undone by lifting the quarantine, with no re-upload and no lost URL.
Keyed by the file's content digest, not by the document. One quarantine covers every media document that shares the bytes — a template duplicated into forty workspaces is forty documents and one digest — and it keeps biting if the same file is uploaded again.
| Where the state lives | mediaQuarantines/index — one document holding the whole deny list |
| Who may set or lift | super staff role, same bar as a lockdown |
| Reasons | malware, abuse, dmca, legal, manual |
| Audited | Every set and lift, with reason, actor, expiry, and the message the customer sees |
| Expiry | Optional, same semantics as a lockdown — when it passes, delivery restores with no action and no write |
Which digest to send
A media document may carry two digest fields, and they are not interchangeable. Send the strong one.
| Field on the media document | What it is | When to send it |
|---|---|---|
contentSha256 | The full 64-hex sha256 of the bytes, written by the routes that actually held them | Whenever the document has one |
contentHash | A 16-hex (64-bit) truncation of one of two algorithms — sha256 on the direct upload/replace routes, GCS's md5 on the signed-upload route | Only as a fallback, when there is no contentSha256 |
Documents with no contentSha256 are every asset uploaded before the field
existed, plus video larger than 50 MB that came in through the
signed-upload route. The first of those two groups is closed: every route that
mints a media document now writes the digest, and the assets that predate it
have been filled in.
That second class used to be everything except SVG on the signed route, because the browser PUTs straight to storage and the server never held the bytes. Since AGL-1629 that route streams the object back through sha256 at finalize, up to a 50 MiB ceiling — which is exactly the largest non-video cap it accepts, so an image, a PDF, an archive or a deck can never be too big for a strong digest. Only video can exceed it, and only past 50 MB. The bound is deliberate: a 200 MB video is the one shape where a full read costs more than the strengthening is worth, and it is also the shape least likely to be a chosen-prefix collision target.
The request field is called contentHash for both — the route accepts any
8–64 character hex digest and does not care which document field you copied it
out of. So paste the value of contentSha256 into "contentHash" and read the
request field name as "the digest".
Why the distinction is worth a paragraph. Two files can be made to share a
contentHash: 64 truncated bits is collision-resistant by accident rather than
by design, and the md5-derived half is the sharp end — a chosen-prefix md5
collision is an afternoon of ordinary compute, and two files sharing a full md5
share its truncation. So an entry keyed on the weak field can be aimed at a
stranger's file. contentSha256 is one algorithm at full width and has no such
property.
Choosing the weaker key is a missed strengthening, never a missed takedown. The CDN checks every key an asset can present — strong digest, legacy hash, per-asset — and any single match refuses. That is deliberate: entries in force were written under whichever field existed when staff pressed the button, and dropping the legacy key would mean a live takedown quietly lifting itself the first time a replace stamped a strong digest onto the document it covered. An entry written under either field keeps biting.
What each key covers
| Key | Set it with | Reach |
|---|---|---|
hash--{sha256} | contentHash: "<the contentSha256 value>" | Every document sharing those bytes, in every workspace — at delivery and at ingestion wherever the server hashed the bytes itself |
hash--{legacy} | contentHash: "<the contentHash value>" | The same reach with the weaker key. For signed-upload video over 50 MB, and for anything uploaded before the strong digest existed, it is the only one there is |
asset--{scopeSegment}--{mediaId} | by: "asset" plus scopeSegment and mediaId | Exactly one document in one workspace, matched on identity, so it needs no digest at all |
Three limits, each a real hole rather than a caveat:
- The mint leg of a signed upload is not gated.
POST /api/media/upload-urlhands out a signed URL before a single byte exists, so there is nothing to look up. The refusal lands at the finalize step instead, which deletes the orphaned object — nothing is registered, counted or billed — but the bytes do briefly reach the bucket. - For video over 50 MB, a takedown bites within an ingestion path, not across them. That video is the one class the signed route still keys on a truncation of GCS's md5, while the direct upload and replace routes key on sha256 — so the same clip pushed through the other route presents a different digest and therefore a different key. Everything under the 50 MiB digest ceiling now shares one sha256 across every route, so the cross-path promise holds for it. Where a large video must be un-re-uploadable through every path, lock the scope.
- A composite object carries no digest at all. GCS reports no
md5Hashfor one, so a very large signed upload can reach the DAM with neither field set. Quarantine itby: "asset", and know that it is then covered at delivery only — with no digest to compare, the ingestion gate has nothing to match on a fresh upload. (A replace aimed at that document is still refused: the replace route checks the target document's per-asset key too.)
Use by: "asset" deliberately, as well, when the same bytes are legitimate
elsewhere and only this workspace's copy is the subject of a report.
What each audience is told
What a fetcher sees: a neutral 410 Gone, byte-identical to the lockdown
refusal. It deliberately says nothing about why — that a takedown notice or a
malware finding exists on a specific file is not something an anonymous fetcher
has standing to learn. The owning workspace is told the reason in the console;
the internet is told the file is gone.
What the owner sees if they upload it again: a 403 that does explain
itself. Quarantine is enforced at ingestion as well as at delivery — the upload,
replace and large-file-finalize routes all consult the deny list before they
write anything — so a re-upload of quarantined bytes is refused outright rather
than accepted and then served as a 410. Nothing is stored, nothing is billed,
and the customer gets the same "this file was disabled … it has not been
deleted" notice with the support address. Replacing the bytes of a quarantined
asset is refused for the same reason: the takedown would keep biting on the new
bytes, so a "successful" replace would have produced a file that still refuses
to load.
The ingestion gate inherits every limit in What each key covers — the unmintable signed-URL leg, the within-a-route matching, and the digest-less composite object. Delivery is the only place that covers all three.
Billing is untouched, on purpose. Quarantine does not delete the object, does not modify the media document, and does not change the storage counter. The file still exists and still belongs to the workspace — it is suppressed, not erased — so the customer's storage usage and invoice are unchanged. The customer notice says so explicitly, because someone whose file stops loading will otherwise assume their data was deleted.
How fast it bites, and what it cannot reach (AGL-1615). The console now
states this to the operator on the Disabled files page itself, from one shared
model (mediaTakedownReachLines()), so this section and that page cannot drift
apart. In full:
| Surface | Stopped? | Worst case |
|---|---|---|
| Our origin | yes | ~15 s — and a lift is just as fast, because the refusal is never cached |
| The raw Storage download link | yes, immediately | the object's token is rotated; permanent, see below |
| A browser holding the ordinary URL | yes | 60 s (max-age=60) |
| The CDN edge, for an image | yes | ~1 h (s-maxage=3600) plus one stale serve. Video, PDFs and other types are private since AGL-1515 and are never edge-held |
A browser holding an image as a published page names it (the versioned ?v= URL, AGL-3485), or the content-hashed permanent URL | no | both forms promise never to change, so nothing can expire them early, and there is no per-file purge. The edge still drops the versioned form within the hour above |
| Anything already downloaded | no | a browser cache, a corporate proxy, a downstream CDN, a scraper, an archive snapshot |
A takedown stops new delivery. It is not a recall. Say that to a complainant in those words. "Stopped within 15 seconds at origin, up to an hour at the edge for an already-cached image" is what safe-harbor "expeditious" contemplates, and it is not the same promise as "the file is gone". Treat any public asset with real traffic as already distributed.
The raw download link, and why it is the one irreversible part. A media
document also carries a url pointing at
firebasestorage.googleapis.com?alt=media&token=…, served by Google, where
none of our code runs — so the deny list is not slow there, it is never
consulted. That is the delivery path for every free-tier workspace, every
private asset and every embed predating AGL-829. Quarantine therefore rotates
that object's token, which kills the published link at once. It cannot be
undone: releasing the quarantine restores CDN delivery but does not
resurrect that particular URL, so any page still embedding it stays broken.
The switch is on by default (a legal or malware takedown with a live public
URL is a takedown that failed) and can be turned off for a precautionary one
you expect to lift. Rotation reaches the one object whose document the
operator looked up — a digest key covers copies in other workspaces, and those
keep their own links until they are taken down individually.
What was considered and rejected. A Vercel edge purge would close the
image window, but it puts a credentialled third-party call on the takedown
path: fail hard and a Vercel outage becomes a takedown outage; fail soft and
you have a purge you cannot rely on, which is worse than none because you
believe in it. Shortening the image s-maxage trades a measured hit rate on
the DAM grid's hot path (AGL-1515) for a faster worst case in a rare event.
Neither reaches bytes already delivered, which is the part that actually
matters.
Where it shows up. A quarantined asset carries a red Disabled badge in
the DAM grid, for staff and for the workspace that owns it, and the badge
carries the customer notice — the reason, the reassurance that the file was not
deleted, and the support address. The internal note is never part of that
payload. Before this, a disabled file looked exactly like a broken one, which is
the state most support conversations about it started from.
Operating it from the console
Staff → Disabled files is the form. Reach for it first — it removes every transcription step the curl below still has.
- Pick Workspace (org) or Site (host), paste the id, paste the media
id. Both halves are in the file's CDN URL: the scope segment, then the id
after
/media/. Look it up. - The panel shows every key that could refuse this file, which of them are set, the reason and internal note behind each, and the deny list's size against its 2000-entry cap.
- Pick the reason, an optional customer-facing message, an optional internal note and an optional end time, then Disable this file.
Two things it does that the curl cannot, and they are the reason to prefer it:
- It never asks you for a digest. You name the file; the server reads the
document and picks the strongest key it has —
contentSha256, then the legacycontentHash, then the per-asset key. The Which digest to send decision is made for you and cannot be made wrong. The scope segment is derived the same way, so a per-asset key always matches the one the CDN actually looks up. - Release clears everything that is biting, not just the preferred key. An
asset can be covered by two entries at once — a legacy-keyed one set before a
replace stamped a strong digest onto it, plus a per-asset one — and a lift
that dropped only one would leave the red badge up and look exactly like a
lift that failed. The page reports
NOT CONFIRMEDunless no key can still refuse the file, and logs every action that reached the server.
Disable only this copy is the same deliberate narrowing as by: "asset":
use it when the same bytes are legitimate elsewhere and only this workspace's
copy is the subject of the report. The key that is about to be written, and
what it reaches, is on screen before the button.
Setting and lifting needs the super staff role, on the page and in the route. Looking a file up does not — during an incident "is this already disabled?" is usually a support question.
The whole deny list
Everything above is per-file: you name a file and learn what refuses it. That is the right shape for an incident that starts with a report, and the wrong shape for the one situation the 2000-entry cap exists for — the list is full, the next takedown is refused with a 409, and the remedy is "release stale entries". You cannot release what you cannot enumerate, and the only other route to a stale entry is knowing a media id it covers, which for a hash-keyed entry set months ago is exactly what nobody remembers.
The second half of Staff → Disabled files renders the deny list as a
table. Reading it is open to every staff role; releasing from it needs
super, like every other write here. Three things to know before using it:
- A row is a key, not a file. Release from the table clears exactly that one entry. The file-mode release above — "clear every key that could refuse this file" — is correct there and wrong here: the entry may cover a document that has since been deleted, and a hash key covers files in workspaces the row knows nothing about.
- Oldest first, because the whole point is finding what has been sitting there. An entry with no set-time at all predates the field and sorts first.
- Expired-but-unreleased rows are called out. Enforcement stops the moment an entry's end time passes, with no write — so those rows refuse nothing and still consume the cap. They are the safest thing to clear first, and nothing else on the platform would ever have told you they were there.
The table pages, and its toolbar filters and searches the whole list, not only the page shown: Reason, State (enforcing, EXPIRED, UNREADABLE) and Key kind are pickers; the key and the note filter as typed text (contains, does not contain, equals, starts with, ends with, empty, not empty); and the search matches anywhere in the key, reason, note or the copy it was set from. Every filter combines with every other and with the search. The deny list is one document rather than one record per entry, so the console reads all of it — at most 2,000 entries, the cap — and answers the filters over every entry. A filter the table cannot apply shows a notice above it, "Filter is not applied: why", and is left out rather than applied to some rows. The count, the full-list warning and the "enforce nothing" count above the table always describe the whole list. See Filter and search a list. The Library picker and the id fields above belong to the per-file lookup, not to this table.
From a terminal
The page cannot do anything this cannot; it just makes the two mistakes above
unavailable. POST /api/admin/media-quarantine with a staff bearer token:
# Disable one file. The value is the media document's `contentSha256` —
# see "Which digest to send" above; fall back to `contentHash` only when
# the document has no `contentSha256`.
curl -X POST -H "Authorization: Bearer $STAFF_TOKEN" \
-H 'content-type: application/json' \
-d '{"action":"quarantine",
"contentHash":"9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08",
"scopeSegment":"org:acme","mediaId":"m1","reason":"dmca",
"note":"Notice #4417 — staff eyes only",
"message":"Disabled pending review of a copyright claim."}' \
https://app.aglyn.com/api/admin/media-quarantine
# Lift it — the SAME digest that was used to set it. A release removes the
# one key it names, so lifting a legacy-keyed entry means sending the legacy
# `contentHash`, even if the document has since gained a `contentSha256`.
curl -X POST -H "Authorization: Bearer $STAFF_TOKEN" \
-H 'content-type: application/json' \
-d '{"action":"release",
"contentHash":"9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08"}' \
https://app.aglyn.com/api/admin/media-quarantine
# An asset with no digest at all (composite object, or a pre-digest upload).
curl -X POST -H "Authorization: Bearer $STAFF_TOKEN" \
-H 'content-type: application/json' \
-d '{"action":"quarantine","by":"asset",
"scopeSegment":"org:acme","mediaId":"m1","reason":"malware"}' \
https://app.aglyn.com/api/admin/media-quarantine
# What is quarantined right now (open to every staff role). `count` against
# `maxEntries` is worth reading — a full list refuses the next takedown.
curl -H "Authorization: Bearer $STAFF_TOKEN" \
https://app.aglyn.com/api/admin/media-quarantine
message is shown to the customer; note is the internal rationale and
never leaves the audit trail. Like every lockdown write, the response carries
the server's re-read of what it wrote — if confirmed is false, the write
returned and the state still disagrees. Treat that as an unresolved incident.
How this surface came to be
The arc, for whoever inherits an incident and wonders why the page is shaped the way it is:
- AGL-1512 shipped the enforcement — take one infected file down, not the host that serves it — keyed on the content digest so one takedown covers every copy of the bytes in every workspace.
- AGL-1613 closed the re-upload chokepoints: quarantined bytes had been refused at delivery but accepted back through upload, replace, and large-file finalize, so a takedown could be undone by uploading the file again.
- AGL-1612 gave quarantine a staff surface at all — before it, the DAM did not show a disabled file (it looked exactly like a broken one) and no staff page existed. It added the red Disabled badge, for staff and for the owning workspace.
- Setting and lifting stayed a curl with a bearer token, which is a fine runbook and a bad incident tool: the operator transcribes a digest and a scope segment, chooses between two digest fields, and then believes the result. AGL-1631 exists because this runbook named the wrong digest field — the reason Which digest to send is a section and not a footnote.
- AGL-1687 built the form, which asks for the file rather than any key so
the digest decision cannot be made wrong, and imported the read-back
discipline (
NOT CONFIRMEDover claimed success) from the Lockdown page (AGL-1571). - AGL-1700 added the deny-list table: the
GEThad returned the full listing since AGL-1512, and nothing had ever rendered it, so the cap-full remedy — release stale entries — required remembering a media id nobody remembers.
The tenant API surface — where a lock does and does not reach
Worth knowing before you press the button, because it is the one place a lock is narrower than it looks.
The tenant middleware is what takes a locked site off the air, and its
matcher deliberately excludes /api. It exists to stop cached pages serving.
So under a full org or host takedown the site's pages 503 while its API routes
keep answering unless each one gates itself. Lockdown also fails open on
every read path — an unreachable Firestore answers "not locked", because a
database blip must not weld every customer site shut — which means the per-route
gate is the whole of the enforcement in both the normal and the degraded case.
Since AGL-2495 every route under apps/tenant/app/api has a written
disposition, held by apps/tenant/specs/lockdown-tenant-api-coverage.spec.ts.
A new route fails that spec by existing until someone decides which it is:
| Disposition | What it means |
|---|---|
| Wired | the route asks the verdict and returns the refusal itself — visitor writes (forms/submit, the plugin dispatcher), the analytics beacon, protection/unlock |
| Delegated | a shared module runs the verdict for it, named in the file — the media CDN handler, the page loader, the edit-access mint |
| Exempt | it must stay reachable while locked, and says why in the file and in the spec's audit table — the 503 notice page, the verdict probe, the abuse intake, the DMCA counter-notice, the health probes |
| Known-open | it still serves site content or metadata through a takedown, recorded by name and held write-free |
What a full lock stops on the tenant API, that it did not before:
POST /api/protection/unlock— a password-protected page's node tree. It verified the password and returned the composed page while the site around it was 503ing. It now answers the same 423 a paused write gets, after the brute-force counter and before the page read. Read-only locks are unaffected: a site that is still serving still unlocks.GET /api/screen/not-found— the host's designed 404 body. The loader behind it was a second entry point that never resolved a verdict, so a locked site kept handing out its own header, footer and nav to anyone hitting a missing URL. It now falls back to the platform status page.
What a full lock still does NOT stop, and it is deliberate that you can read
the list rather than discover it: collections-rss, sitemap, robots,
manifest, host/{hostId}, the published screen list, and the plugin fetch
proxy. All are reads or metadata, all are frozen by name in that spec, and each
is asserted to stay write-free — an ungated read is a disclosure decision
somebody made; an ungated write is the defect. Closing them needs a
read-refusing gate across the tenant API, which is a decision with its own
blast radius and has not been taken.
The same spec holds the background jobs, for the reason the AGL-1621 drill
found: apps/tenant/utils/publish-schedule-job.ts ran a scheduled publish on
platform credentials from a secret-gated route, outside every gate, so a
publish could fire on a locked host. It is gated now and the spec names it if
that line ever leaves. Six plugin jobs (four in commerce, two in bookings) are
recorded as still ungated — the runner, not the call site, is where that should
be fixed.
The analytics beacon, which a read-only lock splits in half
POST /api/analytics/collect is the one route where "read-only refuses every
write" is deliberately not the whole answer, so it is worth reading before you
arm a maintenance window and wonder why the dashboard kept moving.
Under a full lock it is frozen completely (AGL-1627). The middleware does
not match /api, so a page already sitting in a visitor's browser keeps
beaconing long after the site went dark — and these are not private counters.
/api/billing/report-usage meters the same hosts/{id}/analytics/{day}
documents the beacon increments, deliberately, so that the bandwidth ceiling and
the invoice can never disagree. A page view recorded against a site we switched
off therefore reaches an invoice. It now refuses, silently, keeping the 204
the beacon always returns: nothing renders that response, so a 423 would buy the
browser a retry and tell nobody anything.
Under a read-only lock the counters keep counting. The site is still serving
— that is the entire point of the mode — so we are still paying the egress, and
those same counters are the meter that the free plan's bandwidth band and the
abuse ceiling are computed from. Freezing them would turn a read-only lock into
a window where a site serves traffic that is unmetered, unbilled and outside the
abuse ceiling, and billing and security are two of the four reasons a lock
gets armed. So the counter stays honest about what we served.
Under a read-only lock the host automations do not fire. The last thing the
beacon does is emit a pageView host event, and that is not telemetry: the
workflow and action runners behind it create records, merge values onto
contacts, add people to campaigns, send email and call outbound webhooks. Those
are visitor-triggered content writes — exactly what every sibling route already
refuses under read-only — and they only escaped the gate because they ride
inside a route named "analytics". A migration or repair now has nothing racing
it here except a commutative counter increment.
If you need the counters frozen as well — a reconciliation of the analytics documents themselves — use a full lock on that site. "Read-only, but also stop the meter" is not a mode we offer.
Owner notices: every lock emails the people it locked
Every lock and every lift at the org, host, domain and user scopes — and a feature lock placed for one workspace — emails the people it affects, from the platform's own sender. That sender is the point: a full lock stops the workspace sending anything, and it signs its members out, so our mail is the only thing that can still reach them.
| Lock | Who is emailed |
|---|---|
| Org | The workspace's owners and admins. |
| Host, custom domain | The owning workspace's owners and admins, and the site's managers. |
| User | The person, at their sign-in address. With Also lock and cancel workspaces this user solely owns, each of those workspaces gets its own workspace notice too. |
| Feature (one workspace) | The workspace's owners and admins. |
| Platform, platform-wide feature | Nobody: it names no workspace. |
The email carries, and only carries:
- the lock's customer-facing message — the same words the lock serves on its notice page (your message, or the per-reason default beneath it);
- what it affects: the sites, whether everyone was signed out, a canceled subscription, paused renewals and payouts;
- how to appeal: reply, or write to the support address with the reference.
It never says why. Anything you want recorded about the reason belongs in the audit row, not the message.
A lift sends the matching "restored" notice to the same people, saying what happens now (renewals resumed, payouts restored, a canceled subscription stays canceled).
Nothing else is sent while the lock stands
After the notice, a locked person gets no automated platform mail: no usage or budget alerts, notification emails, CRM digests or task reminders, AI insight digests or allotment alerts, monthly usage summaries, and no workspace notices such as a paused outreach mailbox. The same holds for an account whose sign-in is disabled, and a suspended workspace's usage alerts and summaries stop too. Console notifications are still recorded, so the history is there after the lock lifts.
Aglyn's own marketing stops too. Every address the account holds is put on the suppression list of each of Aglyn's own sites, with the reason Account locked. That stops campaigns, product updates and sequences. The rows are written only where the address wasn't already listed, so a real unsubscribe or bounce is never replaced. Unlocking the account removes exactly the rows the lock wrote.
Still sent: the lock and lift notices themselves, security alerts, password reset and verification mail, erasure confirmations, and Stripe's receipts. Mail a customer's own site sends is unaffected. A locked person who is a contact on someone else's site still gets that site's mail. Unless the lock is a ban.
A ban stops all mail
An account locked with the reason abuse is banned. Its lock notice is the
last mail it receives. Every address the account holds goes on the platform
suppression list as Banned account, and the check that runs before every
message the platform sends refuses those addresses:
- transactional mail — password resets, verification, security alerts, receipts;
- Aglyn's own mail — workspace notices, usage alerts, digests, marketing;
- every other site's mail, campaigns and one-to-one alike.
Only a lock or lift notice gets through, so Resend owner notice still reaches a banned account.
A Banned account row can't be released from the Emails page. An opt-in doesn't lift it, and a later bounce or complaint doesn't overwrite it. It is lifted only by lifting the ban: unlock the account, or re-place the lock with another reason. If the address was already suppressed for another reason when the ban was filed, lifting the ban restores that reason rather than releasing it.
An account banned under security before abuse existed is not a ban. Re-place
its lock with the reason abuse.
Appeals
Every lock a person can contest offers an appeal — security and abuse locks
on an account, a workspace, a site or a domain. The account holder sees it in
three places:
- The notice email. It asks them to reply, or to write to the support address, with the lock's reference. Replies go to support, because that address is the email's Reply-To.
- The sign-in form. A locked account's sign-in is refused. The refusal names the support address both "to restore access, or to appeal the decision".
- The console notice for a locked workspace or site. It reads "To appeal this decision, write to …" in place of "Questions?", and a custom message cannot remove the line.
The reference is LK- and ten characters, the same for every notice about the
same account or workspace. Search the Audit log for it: the lock's row
carries it as its note. The appeal comes from the person's own address, so you
can also find the account by searching Users for it.
Answer an appeal from the support mailbox. A ban refuses what the platform sends, not what a person sends, so your reply reaches a banned address. To overturn the decision, unlock the account, or re-place the lock with another reason. That lifts the ban and sends the "restored" notice. To uphold it, say so in your reply, and leave the lock as it is.
Email the owners
The checkbox is on for every reason and resets to on when you change the
reason. Untick it only for a legal hold. Either way, the outcome is its own
line in Actions taken in this session — Emailed … the workspace-locked notice — 2 of 2 (verified), NOT sent — Staff chose not to email the owners
(verified, because that is what was asked), or FAILED / NOT sent — nobody to email (NOT CONFIRMED). A lock never skips its email silently.
Resend owner notice
For locks that already stand — placed before owner notices existed, or whose
email failed — the Resend owner notice card sends the lock email the lock
would have sent, with its stored message. Super role only, audited as
lockdown.resend-notice.
- Enter one
scope:idper line (org,host,domainoruser), or press Add the target above. - Each distinct person gets ONE email listing everything locked for them: a user lock and the workspace that user owns are one email.
- A (lock, person) pair already sent is reported as "already sent at …" and not emailed again, unless you tick Send again. A lock that was lifted and placed again is a new lock, and is sent afresh.
# The same action outside the console (super role).
curl -X POST "$CONSOLE/api/admin/lockdown" \
-H "Authorization: Bearer $ID_TOKEN" -H 'Content-Type: application/json' \
-d '{"action":"resend-notice","sendAgain":false,
"targets":[{"scope":"user","targetId":"UID"},{"scope":"org","targetId":"ORG_ID"}]}'
# → { ok, confirmed,
# targets: [{ scope, targetId, locked, recipients, error }],
# recipients: [{ uid, email, lockKeys, outcome: 'sent'|'already-sent'|'failed',
# sentLockKeys, alreadySentAtMs, error }] }
confirmed is true only when every target was locked and every person was
either emailed now or had been already.
Stopping billing: cancel the subscription
A lock does not touch Stripe. A locked workspace keeps its subscription, and the subscription keeps renewing. For a fraudster that renewal usually lands on a stolen card and comes back later as a chargeback. So when the reason is fraud or abuse, stopping billing is part of the lock.
You can cancel billing from two places. Both use the same server helper and write the same audit row.
-
With the lock, on Staff → Lockdown.
- With the Workspace (org) scope, tick Also cancel its subscription now (no refund).
- With the Account (user) scope, tick Also lock and cancel workspaces
this user solely owns. This locks every workspace the account owns
(
orgs.ownerUid) with the same reason, through the same org-lock path, and then cancels each one's subscriptions. Workspaces the account only belongs to are not touched. A workspace that is already locked keeps its existing lock and still has its billing canceled. One account lock covers at most 25 owned workspaces. Lock any beyond that by org id.
-
On the staff org page (Staff → Organizations → the org). The Subscription card lists every subscription with its plan, status and next renewal. Cancel subscription… is in the card header. It asks for:
- When: now or at period end.
- A reason. Other also needs a note.
- The workspace's slug, typed in to confirm.
Use at period end when the customer asked to leave and has paid for the rest of the period.
Which lock cancels billing by default
The checkboxes start on only when the reason is security. Changing the
reason resets them to that reason's default. You can still untick them for a
security lock, or tick them for any other reason.
| Scope | security | billing | maintenance | manual |
|---|---|---|---|---|
| Workspace (org) | cancels (box on) | does not (box off) | does not (box off) | does not (box off) |
| Account (user) | locks and cancels owned workspaces (box on) | does not (box off) | does not (box off) | does not (box off) |
| Platform, site, domain, feature | never | never | never | never |
The renewal and payout pauses follow the same rule for workspace and site locks; see Stopping a tenant's money.
Billing, maintenance and manual locks must never end a subscription unless
someone chooses to. A billing lock exists so the customer can fix their card
and come back. The API infers nothing from the reason: without
cancelSubscription: true (org) or lockOwnedWorkspaces: true (user), no
lock cancels anything.
What a cancellation does, and what it never does
- It finds every subscription. It lists the org's stored Stripe customer
(
status=all) and also searchesmetadata['orgId']. The search finds a subscription created against a different customer. If either lookup fails, the result is not confirmed: "we canceled what we found" is not the same as "nothing is billing any more". - Now: Stripe deletes the subscription with
invoice_now=falseandprorate=false. There is no final invoice, no proration credit and no refund. - At period end: it sets
cancel_at_period_end. If a pending downgrade schedule is holding the subscription, it releases that first, the same way a customer's own cancel does. - It never refunds. Money already collected stays collected. If a refund is owed, issue it separately from the Refund a charge card, which has its own audit trail.
- It is idempotent. A subscription that has already ended, or is already set to end at the period end when that is what you asked for, is reported and not written again.
- It records the reason at Stripe. The reason goes into Stripe's
cancellation_details.comment, along with the channel (staff-consoleorlockdown) and your uid. - It answers with a read-back. Each subscription is re-read after the
write and reported with
confirmed, the same way the lockdown route reports a lock. - It leaves the org doc to the webhook. The existing webhook projects
plan: freeandbillingStatus: canceledonto the org, as it does for every other cancellation.
Lifting a lock never recreates a subscription. Unlocking the workspace or the account gives the customer access back, but their plan does not come back. To have a plan again, they have to subscribe again. An account unlock also does not unlock the workspaces its lock locked. Lift each workspace separately.
A cancel that fails does not undo the lock
The lock is written and audited first. The cancel runs only after that.
The lock response keeps its own confirmed. The cancel is reported beside it
as subscriptionCancel (org scope) or as ownedWorkspaces[].subscriptionCancel
(user scope), each with its own confirmed. Actions taken in this session
shows the cancel on a separate line. A failed cancel appears there as
NOT CONFIRMED, and an error snackbar says the lock is in place. In that case,
finish the cancel from the org page's Subscription card. The cancel is
idempotent, so running it again is safe.
# The same action outside the console (super role). No refund, ever.
curl -X POST https://app.aglyn.com/api/admin/billing/cancel-subscription \
-H "Authorization: Bearer $STAFF_ID_TOKEN" -H 'Content-Type: application/json' \
-d '{"orgId":"<org id>","when":"now","reason":"fraud","note":"phishing kit"}'
# → { ok, confirmed, changed, lookupErrors, subscriptions: [{ id, outcome,
# confirmed, error, verified: { status, cancelAtPeriodEnd, … } }] }
GET …/cancel-subscription?orgId=<id> returns the subscriptions without
changing anything. Any staff role can call it.
Each cancellation writes an adminAudit row: action: org.subscription-cancel,
target: orgs/{id}, the reason and note, via, when, refunded: false, and
the outcome for each subscription. The row is written for a failed attempt too.
Stopping a tenant's money: renewals and payouts
A locked site takes money from its own customers as well as paying us. Its membership renewals keep charging, and the seller's connected account keeps paying out. For a workspace or site lock, two more checkboxes stop both:
- Also pause the membership renewals it sells.
- Also pause the seller's payouts.
Both start on for security and off for every other reason, and reset
when you change the reason. The API infers nothing: without
pauseRenewals: true or pausePayouts: true in the request, nothing is
paused. Both run after the lock is written and audited, and neither can undo
it.
What pausing renewals does
- It finds the subscriptions from the plugin's own records. Each plugin
that sells subscriptions tells the lockdown which live ones a site sells.
Commerce reads its
hosts/{hostId}/subscriptionsrecords. A workspace lock covers every site of the workspace (up to 200). Nothing is searched in Stripe. - It pauses collection, with
void. Each live subscription getspause_collection[behavior]=void. While paused, every invoice is voided, so no charge goes through. Nothing is canceled or refunded, and the customer is not emailed. We chosevoidoverkeep_as_draftbecause these subscriptions live on Aglyn's platform account, where the merchant cannot see or act on a draft invoice. A security lock usually means the cards may be stolen, and a draft is a charge waiting for someone to finalize it later. A voided invoice is a final "not charged" record. - It leaves a merchant's own pause alone. A subscription that is already
paused when the lock lands is reported as
already-paused. It is not recorded, so the lift never resumes it. - It records what it paused. Each pause is written to
lockdownBillingPauses(Admin SDK only; no client can read it), naming the lock that holds it.
What pausing payouts does
- It finds the seller's account. This is the connected account on the workspace owner's profile, the one every sale of that owner's sites pays into. If the owner runs several workspaces, their payouts pause too.
- It saves the schedule, then switches to manual. The account's current
payout schedule (interval, anchor, delay) is saved in
lockdownBillingPauses, and the account is set tosettings[payouts][schedule][interval]=manual. Money stays in the account balance and is not paid out. - A Standard account is not controllable. Aglyn creates Express accounts, and the platform can set payouts only for Express and Custom accounts. For a Standard account the result says not controllable — pause them in the Stripe Dashboard. That is not a failure of the lock. Pause the payouts by hand in the Dashboard: Connect → Accounts → the account → Payouts.
- A workspace lock pauses every account the workspace is paid through.
Besides the storefront account, a plugin can declare other connected
accounts it pays the workspace through. The marketplace declares the
publisher's payout account. Each account gets its own record, its own saved
schedule and its own line in the result (
storefront,marketplace publisher). A site lock pauses only the storefront account, because a workspace's marketplace payouts do not belong to one site. - A new publisher's payout delay waits for the lock. A publisher in its first 30 days has its payouts held 14 days. That hold is never changed while a lock holds the account. The lift restores the delay the lock saved, and the publisher's next sale moves it on from there.
What a workspace lock does to its marketplace listings
A workspace lock of any reason takes the workspace's marketplace listings
out of browse and search, and their pages read as unavailable. Anything
already installed from them keeps working. A new sale is refused under any
lock, and a new install or update is refused under a security lock. The
result shows one line, for example marketplace: Hid 4 marketplace listing(s) of 4.
The lock marks each listing without changing its own state, so the lift restores exactly what was visible before. A listing the publisher had unpublished or made private, or one staff had taken down, stays that way.
What the lift restores
Unlocking the workspace or site resumes exactly the renewals that lock paused, and restores exactly the payout schedule it saved. It does not resume anything it did not pause.
If two locks cover the same subscription or account (for example the workspace and one of its sites), each lock joins the hold instead of pausing again. The money is restored only when the last of those locks is lifted. Lifting the site while the workspace is still locked leaves everything paused.
A lift that fails to resume keeps its record, so lifting again retries exactly what is left.
Reading the result
Actions taken in this session shows each step on its own line, apart from the lock:
Paused N membership renewal(s), with counts of any already held by another lock or already paused by the merchant;Payouts for acct_… set to manual (was weekly; saved for the lift), ornot controllable (Standard account) — pause them in the Stripe Dashboard;- on a lift,
Resumed N membership renewal(s)andPayouts for acct_… restored to weekly.
A step that failed shows as NOT CONFIRMED and the lock still stands. Each
step writes its own adminAudit row (lockdown.renewals-pause,
lockdown.payouts-pause, lockdown.renewals-resume,
lockdown.payouts-restore) with refunded: false and canceled: false, and
failed attempts are recorded too.
# The same lock outside the console (super role).
curl -X POST https://app.aglyn.com/api/admin/lockdown \
-H "Authorization: Bearer $STAFF_ID_TOKEN" -H 'Content-Type: application/json' \
-d '{"action":"lock","scope":"org","targetId":"<org id>","reason":"security",
"pauseRenewals":true,"pausePayouts":true}'
# → { confirmed, verified, renewalsPause: { confirmed, subscriptions: [...] },
# payoutsPause: { outcome, accountId, schedule, confirmed, error } }
Operating it
- Open Staff → Lockdown (or suspend a workspace from its org detail page — same mechanism underneath).
- Pick the scope and target, the reason, an optional customer-facing message, and an optional end time.
- Platform locks require typing the confirmation phrase — in the UI and in the API, so no script can take the platform down with a one-field request.
- Lift it from the same page. Lifting also evicts stale notice pages, restores member write access, and is audited like the lock was.
Never take a lock or a lift on trust
A click is a request. A click that misses — the page settles, a banner collapses, the button moves — looks exactly like one that worked, and the dangerous half is a lift you believe happened: a controlled 60-second action becomes an outage nobody is watching.
So the page never claims a state it has not read back:
- Every lock and lift answers with the server's re-read of the target, and
the workspace/site/account card shows that verdict —
LOCKEDorNOT LOCKED— stamped with the time it was read. It is a snapshot, not a live view, which is why the time is on it. - Check state re-reads one target without touching it. Use it freely; it
is available to every staff role, not just
super. - The verdict is discarded the moment you change the scope or the target id — a panel about the previous target is worse than no panel.
- The target id now stays after a submit, so
Unlockis live immediately after a lock instead of being a disabled control. - Actions taken in this session lists everything that reached the server, with the time. If you clicked and no new line appeared, the click did not register — check the state and click again.
- A write that returns but whose re-read disagrees is reported as
NOT CONFIRMED, loudly. Treat it as an unresolved incident, not a success.
What a caller is told
You cannot check a lockdown by trying it yourself. Staff bypass every scope — that is the un-panic invariant, and it is deliberate — so your own request succeeds no matter what is locked. Signing out does not help either: without a credential the request is refused as unauthenticated long before the lockdown verdict runs. The customer-visible refusal lives in a band between those two that a staff operator has no way to stand in.
What would this caller be told? answers it from the other side. Describe the caller — a user uid, a workspace id, a site id, or any combination — and the server runs the same verdict every API route runs and shows you:
- whether that caller is refused for reads, for writes, or for both, in one line — under a read-only lock it says "reads pass, writes refuse", which is the answer to the question read-only mode creates: their site is up but they cannot save — is that us?;
- under which scope and reason the refusal falls;
- the exact response body they receive, built by the same code that builds the real 423 — not a summary of it. Under a read-only lock this is what their write receives; their reads get the real data;
- which capabilities (signups, uploads, checkout, marketplace installs, AI assist) are refused for them, since a feature lock bites without touching any scope;
- whether the account you named is itself staff, in which case it bypasses everything and the answer says nothing about whether a lock is engaged.
Two honesty rules the panel follows, and you should read it by:
-
It is computed, not observed. It is what this server derives from state it reads at that moment. It does not prove that any route returned it, and other server processes converge within about 15 seconds, so a lock armed seconds ago may not yet be enforced everywhere.
A customer's public site takes longer than that, and the 15 seconds is not the number to quote them. A site page is gated in the tenant middleware, which memoizes its verdict for a further 30 seconds. Measured 2026-08-23 (AGL-1621) against the emulator, not production: an
orgorhostlock reaches an already-warm isolate in ~30s, and aplatformordomainlock in ~45s — the two caches are in series. A lift takes the same time, in the same direction. A cold isolate refuses on its first request, so a refresh that lands on one shows the lock instantly; that is luck, not the bound.These are FLOOR figures, and production is slower. See what has and has not been measured before quoting any of them to a customer or a complainant.
-
A scope you leave blank is not evaluated. "Not refused" for a bare uid says nothing about that person's workspace. The panel lists exactly which scopes the answer covers, and a workspace or site id that matches nothing is reported as such rather than counted as clear.
Reading is open to every staff role — during an incident the person who needs
to answer "what is this customer actually seeing right now" is usually support,
not the super-role operator who armed the lock.
To confirm the refusal on the wire rather than in the abstract, you need a
caller who is genuinely refused. An org API key is the cheapest one: it carries
no staff claim and no uid, so /api/v1 refuses its own holder. See
Verifying a lockdown on the wire.
What has and has not been measured
The timing figures in this document have one provenance, and it is not production. Quoting them as if it were is the mistake that produced the figures they replaced.
| Figure | Where it was measured | Status |
|---|---|---|
| Verdict-reader TTL, 15.1s both directions | Emulator, driving the real route against a real Firestore | Measured 2026-08-23 (AGL-1621) |
| Tenant middleware memo, 30.2s lock / 30.2s lift | Emulator, against the real middleware | Measured 2026-08-23 (AGL-1621) |
| Composed ~15s console/API, ~30–45s site pages | Derived from the two above | Measured 2026-08-23 (AGL-1621) |
| Anything on production | — | NOT MEASURED. No production drill has been run against the current build. |
Treat the emulator numbers as a floor. Three things production adds, all of which make it slower:
- Edge-isolate fleet spread. The emulator has one warm process. Vercel has many, each with its own 30s memo, converging independently.
- Real Firestore round-trip time. The emulator's is a loopback.
- The ISR fan-out is serial, one host at a time with a 5s timeout each
(
revalidateHostAfterLockdown). Anorglock over twenty sites can take ~100s to return while already being in effect. Thehostscope pays this once, not per host.
The old "lock visible in ≤10s" figure was a cold isolate — which refuses on its very first request — recorded as if it were a bound. Do not repeat that by blending a surface: say which surface a number came from, or do not quote it.
Why the production drill has not been run
A production drill was scoped on 2026-08-23 and stopped before anything was locked, because it has no safe subject. Recorded here so the next person does not rediscover it at the worst moment:
- There is no throwaway production host. Every seeded fixture
(
seed-e2e.mjsand friends) is emulator-gated and refuses to run against a real project. - The
demohost is not a throwaway. The tenant middleware falls back to it forapp.aglyn.comand for every Vercel preview deployment, so locking it takes down far more than one site. - It would deliberately turn a monitored canary red. The render canary
grades
demoby loading its home page; a full lock makes that page compose an empty node tree, which is exactly the failure the canary exists to catch. Its 5-minute memo can hold the red after the lock lifts. - A 2-minute dead-man expiry is shorter than one propagation cycle (~45s each way). The lock could expire before every isolate has observed it engage.
- A dead-man expiry is not the same code path as a lift. Expiry is evaluated at read time and performs no write, so it fires no revalidation fan-out. Measuring release via expiry would measure a different mechanism from the one an operator actually uses under pressure.
A safe production drill needs a genuinely disposable host — its own org, not
referenced by any fallback, canary or monitor — provisioned first. Until then
the emulator harnesses
(route.drill.emulator.spec.ts, and LOCKDOWN_DRILL=1 for the middleware
timings) are the only repeatable source of these numbers.
Verifying a lockdown on the wire
A read-only API key on a disposable workspace turns the whole 423 sweep into one curl, because the customer REST API deliberately evaluates the verdict with neither a staff claim nor a uid:
# Unlocked: 200 with the service document.
curl -i -H "Authorization: Bearer $AGLYN_DRILL_KEY" https://app.aglyn.com/api/v1
# With that workspace locked: 423 Locked, Retry-After, and the notice body.
# {"error":"locked","scope":"org","reason":"billing","title":"Account on hold",
# "message":"…","contact":"support@aglyn.com","untilMs":1786695133044}
The same key proves the platform scope (lock the platform, the same call answers
"scope":"platform"). Feature scope needs a caller on a feature chokepoint —
signups grants no staff bypass, so a staff token on
POST /api/auth/legal-acceptance is refused under a signups lock and can be
checked without any extra credential.
What the audit row records
Every lock and lift writes an adminAudit row carrying the actor, the scope,
the target path, and — in before and after — the reason, the
customer-facing message, the mode, the enforcement class, and the end
time as untilMs. mode and enforcement are always stated rather than
omitted, so a row about a lock written before either field existed reads
full / standard instead of a gap you have to interpret — and a takedown is
exactly the row somebody will later have to produce as evidence. Recording the end
time is the point: it is the only thing that distinguishes a deliberate
time-boxed lock from an indefinite one nobody came back to, and on a lift it
says whether a time-boxed lock was released early or a forgotten one was
cleaned up. A null in any of those three means the lock genuinely carried no
reason, no message, or no expiry.
Two locks are placed by the platform rather than a person, and their rows say so
with a system: actor and after.automated: true: the billing sweep below
(system:billing-auto-lock), and the page screen's automatic security hold
(system:page-screen), whose note names the page, the version, the signal and the
abuse row.
Billing locks for lapsed subscriptions are manual by default. The automated
30-days-past-due sweep exists but ships disabled; it is enabled by setting the
AUTO_LOCK_BILLING_FROM environment variable to a start month (YYYY-MM) — a
deliberate operator decision, never a default.
What the sweep counts as delinquent, and why it is not just past_due
(AGL-1877). A Stripe test-mode test-clock drill of a failed renewal
measured the timeline: the subscription retries five times, stays past_due
throughout, and Stripe cancels it at 21.08 days with
cancellation_details.reason: 'payment_failed'. It never becomes unpaid. So the
30-day grace clock outlives the past_due/unpaid statuses by nine days, and the
predicate had no reachable true branch at all until it also accepted a
canceled-for-non-payment subscription. That reason is now mirrored onto
orgs/{id}/billing/stripe as subscription.canceledReason, and the sweep locks
only on the literal 'payment_failed' — a workspace that canceled on purpose, and
every cancellation recorded before this shipped, fail closed and are never locked.
The LIVE dunning schedule — read from the Dashboard, unreadable by API (AGL-2430)
Every number in the paragraph above is a test-mode measurement. Stripe's retry schedule, the Smart Retries flag, the after-the-final-retry behavior and the subscription-email toggles are Dashboard settings held independently per mode (Settings → Subscriptions and emails). Test and live do not share them, and this account has already been shown to diverge between modes on a neighboring setting — product tax codes were live-only until AGL-1877 reconciled them.
What was measured in LIVE, read-only, on 2026-08-20 (live account,
GET requests only — no Dashboard setting was
changed):
| Question | Live answer |
|---|---|
| Retry count / interval / terminal behavior | Not readable. No API surface exposes it |
GET /v1/account | No field matching dunning|retry|smart_retr anywhere in the payload |
/v1/billing/settings, /v1/subscription_settings, /v1/billing/dunning, /v1/billing/retry_settings, /v1/account/settings | All 404 Unrecognized request URL |
/v1/billing_portal/configurations | 200 — but it is the customer portal, not dunning |
| Live invoices | 3, all paid |
Live subscription_cycle invoices | 1 — amount 0, attempt_count: 0 |
| Live charges / failed | 1 / 0 |
| Live subscriptions | 2, both canceled, both cancellation_details.reason: 'cancellation_requested' |
So it is unreadable twice over: no endpoint returns the setting, and no live
renewal has ever attempted a real charge from which it could be inferred. The
one subscription_cycle invoice on the account is a zero-amount renewal that
never touched a card, and no live subscription has ever ended for non-payment.
AGL-1877's audit note said the account had produced zero subscription_cycle
invoices; that is now stale in letter — there is one — but correct in
substance, because a zero-amount renewal produces no dunning evidence.
Reading it required a human, and on 2026-08-24 one did. Live Dashboard → Settings → Billing → Subscriptions and emails → Manage failed payments → Card payments → Cards → Manage:
| Live setting | Value |
|---|---|
| Retry strategy | Smart Retries — "Retry up to 4 times within 3 weeks" |
| Attempts | 1 initial + 4 retries = 5 |
| Window | 3 weeks / 21 days |
| If all retries fail | Cancel the subscription (NOT mark unpaid) |
| Invoice after the final retry | Leave the invoice past-due — the debt survives the cancellation |
| If a dispute is opened | Leave the subscription past-due — a chargeback does not cancel |
| Upcoming renewal events | 15 days before renewal |
Live and test agree. The divergence this issue was opened to catch did not
happen: 5 attempts, ~21 days, terminal canceled with reason payment_failed
in both modes. Everything AGL-1877 pinned from a test-clock drill does describe
live behavior.
What that read is, and is not:
- It is a transcription of a human reading a screen on one day. There is no API behind it (see the table above — the endpoints 404, and they 404 in test mode too, because the endpoint list is a property of the API and not of the mode). If somebody edits that Dashboard screen tomorrow, no constant, spec or build in this repository changes.
- Nothing customer-facing may quote the number anyway. That is
MAY_QUOTE_RETRY_WINDOW_IN_COPY, which staysfalseand is a separate constant fromLIVE_RETRY_WINDOW_IS_KNOWNon purpose: knowing a number today does not make it safe to print one that can go stale silently. The console banner and the customer billing docs describe only the shape — access continues while Stripe retries, and the plan stops if the retries run out — andapps/console/specs/billing-dunning-banner.spec.tsxfails if a count, a duration or the phrase "retry window" returns to that copy. BILLING_LOCK_GRACE_DAYS = 30is reconciled. Stripe gives up at 21 days and cancels, so 30 sits nine days past the terminal state with the slack in the customer's favor. It also stood before the read, and that reasoning is worth keeping: it is reachable under all three terminal settings the Dashboard can hold — cancel (thecanceled+payment_failedbranch at ~21 days), mark unpaid and leave past_due (both at day 30 on their own branches). No live value can make it unsafe, only more or less generous.- Stripe does email the customer on a failed payment. That toggle is ON,
along with trial-ending, upcoming-renewal, expiring-card and bank-debit
failure notices (the last was turned ON on 2026-08-24). This matters because
system-email-catalog.tscatalogsstripe-payment-failedasdeliveredBy: 'stripe'precisely because the code cannot see the toggle; Aglyn composes no failed-payment email of its own, and the in-app notification it does send is suppressed entirely by a mutedbillingcategory. Had that toggle been off, a failed renewal would have had no customer-reachable signal beyond the console banner.
The only thing watching the setting: npm run check:stripe-dunning-drift.
Because the setting cannot be read back, tools/scripts/check-stripe-dunning-drift.mjs
watches for drift the long way round. It is GET-only and safe against a live
key — which is the point, since live is the mode that matters.
- Re-probes whether Stripe has shipped an endpoint that exposes the schedule. If one appears, the "unreadable" premise recorded above is stale and the checker exits 1 so it gets replaced by a real one.
- Watches the account's behavior for anything the record forbids — an
unpaidsubscription when the terminal state is recorded ascanceled, an invoice past the recorded attempt count, a retry scheduled beyond the recorded window. Behavior is downstream of the setting, so a changed setting eventually surfaces here. - Checks the record against the code that depends on it, chiefly
BILLING_LOCK_GRACE_DAYS.
Exit codes are 0 in sync, 1 differs, 2 could not check — and it reaches 2
readily and on purpose: no key, an unreadable key, a key whose prefix does not
name a mode, an API failure, a schedule recorded as null, or constants it
cannot parse out of the source. Step 2 is vacuous until a real renewal
fails, so the report separates CHECKED claims from UNVERIFIED ones and prints
the observation count it worked from; --require-observed turns "nothing to
look at" into exit 2, which is the flag the launch-day runbook should use once
live subscriptions exist. Every run names the mode it read, because a green run
against a test key says nothing whatsoever about live.
The mode-tagged constants, and the probe evidence above, live in one place:
apps/console/utils/stripe-dunning-schedule.ts.
What the live Dashboard did say, once someone opened it (AGL-2430)
Settings was read at → Billing → Subscriptions and emails on the live account on 2026-08-23, a day before the retry schedule above was read. What that first pass read was the email half, and it found a defect worse than an unknown number — one that is still open, because its fix is gated on a deploy.
| Setting | Live value | Verdict |
|---|---|---|
| Send emails when card payments fail | ON | Good — the customer is told |
| Payment method updates | Use a mix of both (Legacy), all four links → https://aglyn.com/ | The defect |
| Include a link for customers to manage their subscriptions | OFF | Deliberate — see below |
The defect. All four "Payment method updates" destinations — free-trial reminders, expiring cards, card-payment failures, upcoming renewals — point at the marketing homepage. A customer whose card fails is emailed a link to a page with no way to update a card. Stripe then retries, and at the end of the retries the live terminal setting cancels the subscription. The email works, the notification works, and the customer still cannot pay.
The fix, and the one that was refused. Pointing those fields at Stripe's own hosted card-update page sends a paying customer outside the product to manage their own subscription, and it routes around a bug rather than fixing it: a past-due or locked workspace that cannot reach its own billing page is the actual defect. See The billing recovery path must survive a billing lock below.
What to paste, and the one-way door. The URL is
https://app.aglyn.com/billing (Route.BILLING_ENTRY), org-agnostic because
Stripe stores ONE link for the whole account. ⛔ Switching off Use a mix of both (Legacy) is irreversible — Stripe shows a confirmation dialog saying so,
and the legacy option cannot be restored afterwards. That dialog has been
declined deliberately; accepting it is the account owner's call and nobody
else's.
⛔ "Include a link for customers to manage their subscriptions" stays OFF
Not an oversight — a decision, recorded here and in
LIVE_STRIPE_SUBSCRIPTION_EMAIL_SETTINGS so nobody switches it on later as a
tidy-up.
The toggle appends a Stripe-hosted "manage your subscription" link to subscription emails, which lands the customer in Stripe's billing portal — where cancel is a button. Cancellation at Aglyn goes through the retention funnel (AGL-1859/AGL-1863): survey, downsell, winback, and only then the cancel. A portal link in an email routes around all of it.
The asymmetry is the design, and it is the same one the console's billing page already implements: recovery is self-serve and as frictionless as we can make it; leaving goes through the funnel. Friction belongs on the way out, never on the way back in.
This is a different question from the card-update hand-off, which legitimately opens Stripe's portal from Manage payment methods inside the console. That starts from a page we control, after the customer has already arrived somewhere that knows which workspace they are fixing.
The billing recovery path must survive a billing lock
A lock imposed for non-payment that also prevents payment is a deadlock: it can never be lifted by the one action that would lift it. Nothing on this path is exotic, which is exactly the hazard — each link was written correctly and independently, and any one of them can be closed by a well-meaning change.
The set, as it stands:
- Sign-in works.
/api/auth/sessioncallsgetLockdownVerdictwithstaffanduidand noorg— and that helper evaluates the org scope only when handed an org doc. A suspended workspace's members can still get a session cookie. - The console shell renders.
PlatformLockdownGateis platform-scope only (/api/lockdown-statusreports nothing about an org), so an org lock does not replace the app with the notice page. - The reads succeed.
orgNotSuspended()in the Firestore rules gates writes. The org doc isisOrgMember()-readable andbilling/stripeiscanManageOrg()-readable, neither conditioned on suspension. - The card-update route answers.
api/billing/subscriptioncarries// lockdown-423: exempt— itsportalaction is the card-update path. applyOrgLockdownkeeps sessions for abillinglock. The env-gated auto-lock sweep passesrevokeMemberTokens: falseoutright, and the staff route gates it onreason === 'security' || reason === 'manual'— precisely so members of a billing-locked workspace can reach billing and fix the thing.- The entry point does not filter.
/billingroutes a suspended, past-due workspace exactly like a healthy one.
apps/console/specs/billing-recovery-reachable.spec.ts holds all six in one
place, with the reason attached to each. It reads source rather than driving a
locked session — a real drive needs a live delinquent org — so treat a green as
"nobody has removed the exemption", never as "a locked customer was observed
paying".