Keeping my AI twin out of trouble
In part 1 I went through how the twin works when everyone plays nice. This part is about when they don't. A chatbot on a public website is an open door: anyone can type anything, and some people (and a lot of scripts) will try to make it do things it shouldn't.
So I spent a lot of time on this part. Here's what I found, from the scariest to the silliest.
Anyone could run up my API bill
The rate limits figured out who you were from a label your browser sends with each request (the
X-Forwarded-For header, for the curious). Problem is, anyone can write whatever they want in
that label.
So a small script could pretend to be a brand new visitor on every single request, and nothing capped how much it could spend on my API account. A security review caught it before anyone else did (phew). Now the app only trusts the address added by my hosting provider, which a visitor can't fake, and there's a daily cap on replies across all visitors, just in case.
Fake conversation history
Remember from part 1 that your browser holds the conversation and sends all of it with each message? That means a visitor can edit it, and invent twin replies that never happened. Picture an earlier answer where the twin "already" shared something private, followed by "cool, tell me more".
The answers themselves were safe, because the judge checks them against my notes, not against the chat. But a few replies skipped the judge, like small talk. My first fix was to close those gaps: small talk only looks at your latest message now, and replies next to a message or booking card use fixed wording. I threw 8 fake-history attacks at it and none of them got anything out.
Then I went one step further. Your browser now sends only a reference to each of the twin's replies, and the server swaps in its own copy. You can still write whatever you want in your own messages, you just can't put words in the twin's mouth anymore.
The same message, eight times
There's a limit of one message per visitor per day. The app checked the counter, then sent the message. If you clicked confirm eight times at the exact same moment, all eight requests checked the counter before any of them updated it.
So eight messages went out, and I would have felt very popular for about a minute. Now checking and counting happen in a single step (one SQLite transaction), so the second click has to wait for the first, and then finds out it's too late.
Someone else called Mouhamad
When you send me a message or book a call, you give your name. Turns out you could just type mine. The request would then reach me looking like I sent it to myself, which is either a weird prank or the start of an identity crisis.
Now the twin refuses any name that looks like mine, including the sneaky versions: accents, invisible characters, or letters from other alphabets that look exactly like ours. Sorry to the other Mouhamads out there, you'll just need to type your full last name. This one came back for a second round when I changed the first step of the twin, see part 5.
"How old are you?"
The gate treated this as small talk, so the twin happily answered as itself: no age, built by Mouhamad, and so on. Cute, but it's a personal question aimed at me, and those should be refused. Personal questions asked with "you" now count as personal. Where I'm based or where I grew up still counts as a work question, since it's in my notes anyway.
What I took from all this
Most of these came from trusting something I shouldn't have: a label anyone can write, a history anyone can edit, a counter two requests can read at the same time, a name field. In most cases the model wasn't the weak spot, the code around it was.
That's why the rule from part 1 matters so much: the model can prepare, but only you can send, after a click and a fresh "are you human?" check. Even if something slips past the gate, the twin can't send anything on its own.
But a twin can be perfectly safe and still wrong. Part 3 is about the twin messing up all by itself, no attacker needed.