The outage rule: what should happen when we are down
A licensing check that takes your site down during our outage is worse than no licensing check.
There is a question every licensing system has to answer and most answer badly: what happens to your site when ours is unreachable?
The two obvious answers are both wrong.
Fail closed — refuse everything when the check cannot be made — means our bad afternoon becomes your bad afternoon. It converts our availability problem into yours, at a multiplier: every one of our customers goes down together, and each of them takes their own customers with them. No serious integration can accept that, and any vendor who ships it has quietly made themselves a single point of failure for your business.
Fail open — allow everything when the check cannot be made — means the licensing system can be defeated by a firewall rule. Block our domain and the check evaporates. You have built an honour system with extra steps.
The rule we settled on
The only thing that may bypass the check is us being unreachable. Not a setting, not a flag, not a configuration.
And unreachable is bounded. A refusal from us closes the gate immediately and permanently. Silence from us falls back to the last answer we gave, and that answer expires.
That is what entitlement tokens are for. When we say yes, we say it in a signed token good for several days. If we vanish, your site keeps running on that token until it lapses. If we say no, the token is not issued and the gate closes now.
Why the distinction survives an attacker
Blocking our domain does not grant indefinite access — it grants exactly the remaining life of the token you already had. Then it stops. An attacker who firewalls us buys days, once, and cannot renew.
Meanwhile a genuine outage on our side is invisible to your users, which is the entire point. You will find out from our status page rather than from your support queue.
The branch everyone gets wrong
In code, this comes down to one thing: catch the network error, do not catch the 403.
try {
const res = await fetch(VERIFY_URL, { ... });
if (res.status === 403) {
// We answered. The answer was no. Close the gate now.
return deny(await res.json());
}
const body = await res.json();
cache.set(body.entitlement, body.entitlement_expires);
return allow();
} catch (networkError) {
// We did not answer at all. Fall back to what we last said.
return cache.valid() ? allow() : deny('unreachable_and_expired');
}Conflating those two is the most common integration mistake we see, and it fails in the worst direction: a revoked key keeps working for days, which defeats the one thing revocation exists to do.
What this costs us
Being honest about the trade: this design means a customer who stops paying keeps working until their token expires. We accept that. The alternative — cutting somebody off mid-request the instant a card fails — creates far more damage than a few days of grace, and most failed cards are an expiry date, not a decision.
So the token life is a deliberate number, not an accident. Long enough that our outage is invisible. Short enough that revocation means something.