Security
The strike counter you could reset by yourself
I built a cool-off to slow down people abusing the chat box. Then I noticed the counter was keyed to a number their own browser hands me — one they could change at will. Here is the fix, and the deferred list it came from.
- Published
- Author
- Shubham N Datarkar
- Read
- 6 min
A few days ago I wrote that I learned to say no — a flat, boring redirect for the visitors who show up to argue with a chat bubble instead of asking it anything. What I did not say in that post is that the mechanism behind it had a hole in it big enough to walk through. This is the post where I go back and close it.
There is a special kind of work that only exists because you were honest the first time. When I shipped the abuse gate I wrote down the parts I had not finished — the follow-ups, the corners, the things that worked in the happy case and nowhere else. Deferred work is where good intentions quietly go to die. The only reason I got to come back to these is that they were written down instead of forgotten.
The counter you could reset yourself
The gate keeps a strike count. Push at it repeatedly and the strikes add up until you hit a short cool-off — the machine equivalent of leaving the room for a minute. Reasonable, except for one detail: I was counting the strikes against the conversation id.
The conversation id is a number your browser hands me. It is convenient, it groups your messages together, and it has one property that ruins it for this job: you can change it whenever you like. Start a fresh conversation and the count goes back to zero. So the cool-off worked perfectly against someone who did not know it was there, and not at all against the exact person it was built for. A lock you can open by asking for a new key is a decoration.
Anything the client sends you is a suggestion, not a fact — the same rule that means a price never comes from the browser. A conversation id is fine for grouping messages. It is not fine for anything the sender has an incentive to lie about. I used a display value as a security value, which is the tidiest way to build something that only looks like it works.
The strikes now key on the IP address instead — the one identifier the abuser is not casually handing me and reshuffling between messages. And because we run on more than one instance, the count lives in a shared store, so a strike earned on one server is a strike everywhere, not a fresh start each time the load balancer picks a different door. Where there is no shared store configured, it falls back to keeping the count in memory; where there is no IP at all, it falls back to the old per-conversation count, which is worse but is only ever the fallback.
The block that forgot everything when I restarted
Above the cool-off sits a heavier response, for the small number of messages that are not rudeness but a genuine threat. Those get the sender's IP blocked. The problem was where I was keeping that block: in a variable in memory. Which means it survived exactly until the next time the server restarted — a deploy, a crash, a quiet Tuesday — and then it was gone. It was also invisible to the admin screen that lists blocked IPs, so a human reviewing the blocklist would not even see it there to know it had happened.
A block that evaporates on restart and that nobody can see is not really a block. It is a note I left for myself and then lost. So threats now go through the same durable, database-backed blocklist everything else uses — the one an admin can actually open, review, and lift. It is time-bounded, it respects the allowlist so a shared office IP does not take out a whole building, and it is reserved for the genuinely nasty tier, not for someone being merely tiresome.
The tiresome tier gets handled a rung lower now too: a stream of threats also counts toward the cool-off, weighted heavier than ordinary noise, so someone who only ever sends threats trips the lock rather than sliding under it because none of their messages were an ordinary strike.
Failing in the safe direction
One more corner, and it is the one I am most glad I went back for. To decide how to respond I sometimes ask a small, cheap model for a judgement, and every such call is metered against a budget so a flood cannot run up a bill. The question nobody asks until the bad afternoon is: what happens when the meter itself errors — when I cannot tell whether there is budget left?
The old answer was to assume yes and carry on, which is the wrong answer for the abuse path specifically. A flood is the exact situation where the meter is most likely to be under strain, and 'assume there is budget' during a flood is how a flood gets to keep spending. So on the abuse path, a budget error now means no — I skip the extra work rather than risk paying for it. The ordinary, non-abuse paths keep the friendlier default and carry on, because there the failure you are guarding against is a real person getting an unhelpful shrug, and a shrug is worse than a small bill.
A control that fails open is only as strong as the day nothing goes wrong. The interesting question is always what it does on the day something does.
The bit worth keeping
The honest part
None of this is switched on for most of you, and that is deliberate. The whole abuse gate sits behind a flag that is off by default; these changes alter what happens only once someone deliberately turns it on. I am describing plumbing, not a wall you will ever walk into, because if you are here to ask a real question you were never going to meet any of it.
The reason to publish a fix to something nobody could see failing is the same reason to publish the security inventory nobody claps for: a thing written down in public is a thing you can be held to. I told you I learned to say no. The follow-up is that the first version of 'no' had a back door, I knew it did, I wrote it down, and I came back and shut it. That last part is the only bit that actually counts.
Common questions
Can someone reset their abuse strikes on the Bolt chat by starting a new conversation?
No longer. Strikes used to key on the browser-supplied conversation id, which was resettable; they now key on the client IP address and persist across servers via a shared store, so opening a fresh conversation no longer clears the count.
Does Bolt block IP addresses permanently?
No. Blocks are time-bounded, stored in the same database-backed blocklist an admin can review and lift, and respect an allowlist so a shared IP is not caught in the net. Only the genuinely threatening tier triggers a block; ordinary rudeness gets a short cool-off instead.
What happens to the abuse gate if a system error occurs during a flood?
On the abuse path it fails closed — if the budget meter errors, Bolt skips the extra model call rather than risk spending during an attack. User-facing, non-abuse paths keep the friendlier default so a genuine question is never blocked by a transient error.
