How my AI twin works (so far)
There's a small chat on this site that answers questions about my work. I call it my twin. It only knows what I wrote down for it. It's an experiment and it will keep changing, so bear with me while I try to keep it up and running, and make it even safer along the way.
I started with one post about it. It kept growing every time something broke, so I cut it into five. This first one covers why I built it, what it's made of, and what happens to your message when you ask it something. The other four are about what broke, how I test it, and my latest experiment.
Why I built it
Two reasons. First, I wanted a way to interact with my experience that's more engaging than LinkedIn. You can just ask what I do at Chainguard instead of scrolling through a list of job titles. Second, I wanted a real project to learn how these systems behave once real visitors show up (safety, latency, etc.).
The stack
- Site: a static Next.js export. The chat keeps the conversation in your browser and sends it along with each message.
- API: Python 3.14, FastAPI and LangChain. A small SQLite database stores the logs, the feedback, and the messages and call requests waiting for your confirmation.
- Hosting: both run in Chainguard containers on Railway. I work at Chainguard, so I'd have a hard time explaining any other choice.
- Models: Jev, from TypeSafe, for the very first step: sorting your question. Claude Haiku 4.5 for the rest of the fast stuff: writing the answer, small talk, helping you send me a message or book a call, and taking over when Jev isn't sure. Claude Sonnet 5 for the judge, the step that double-checks every answer.
- Search: Qdrant, running in memory inside the API, so there's no separate database to run
or pay for. The embeddings (the numbers that let the app find notes close in meaning to your
question) come from BGE small (
BAAI/bge-small-en-v1.5) through fastembed. It runs on the CPU and costs nothing. It only speaks English though, which is why the twin translates your question into English before anything else. Multilingual models exist, and trying them is an ongoing experiment. More on that in the trade-offs below. - Notes: plain Markdown files. My career, a FAQ, my certifications, a persona file that sets the twin's voice, and the posts on this site (yes, including this one).
- Monitoring: LangSmith, on its EU region. I can open any conversation and see every step with its timing. API keys, IP addresses and location are stripped before anything leaves the app.
- Abuse: Cloudflare Turnstile (the "are you human?" check), rate limits, daily caps, and a spend limit on my API account.
What happens when you ask a question
- Gate. Jev reads your latest message, plus the few before it for context, and puts it in a bucket: work question, small talk, personal, off-topic, a request to contact me, or an attempt to trick it. If Jev isn't sure, Haiku makes the call instead. Jev can't write, so when your message only makes sense with what came before, Haiku rewrites it as a full English question. "Which ones?" right after a question about my certifications becomes "Which certifications does Mouhamad hold?" That's what gets searched. French questions get the same treatment and come out in English. For a long time this whole step was Haiku alone. Part 5 is the story of the switch.
- Search. The app looks for the parts of my notes closest in meaning to that question and keeps the best 6, as long as they're close enough (a score above 0.45). My notes are cut into pieces by section heading, so each piece sticks to one topic.
- Draft. Haiku writes a short answer using only those pieces. If nothing matched well enough, the twin just tells you it's not in my notes.
- Judge. Sonnet 5 reads the question, the notes and the draft, and checks four things: is every fact in my notes, does it leak anything private, does it stay on topic, and does it sound like the twin. If the draft fails, the writer gets one more try, with a fixed hint for each problem. It never sees the judge's own comments, because those could carry something a visitor slipped into their question. If the second try fails too, you get a short "I garbled that one" instead of a wrong answer. I'd rather look a bit dumb than make stuff up.
Personal questions, off-topic stuff and attempts to trick the twin stop at the gate and get a canned reply. I'd rather the twin say no than invent something about my life.
Sending a message or booking a call
You can also ask the twin to send me a message or book a call with me. For this part, Haiku gets two tools: one that reads my free slots, and one that drafts a message.
A quick note on those slots. They're not my work calendar. They're times I've set aside where I consider myself free for a chat. And a call isn't booked until I've looked at it and approved it myself, so maybe don't block your afternoon until you hear back from me.
The rule I set early: the model can prepare, but it can never send. The tools only read slots or save a draft. You get a card, you can fix the details, and nothing happens until you click confirm and pass a fresh "are you human?" check. Part 2 shows why that rule mattered more than I expected.
Trade-offs I made
An English-only encoder. The encoder is the model that turns text into numbers so the search can compare meanings. BGE small is tiny (about 67 MB), fast on a CPU, and free. It's also trained on English only, which is exactly why French questions broke at first (that story is in part 3). Since the gate already calls a model, I ask it to translate every question into English on the way. Cheap, and it works.
I could do this differently:
- Use a multilingual model that also runs locally.
paraphrase-multilingual-MiniLM-L12-v2is almost a drop-in swap, but it's trained to compare sentences more than to search.multilingual-e5-largeis better, but it weighs over 2 GB, which is a lot for a small container. - Pay for a multilingual API, like OpenAI's
text-embedding-3-smallor Cohere's multilingual model. Better at French, but every question now goes over the network and shows up on a bill.
That's the ongoing experiment I mentioned. For now the translation holds up. If I start writing posts in French here, the encoder is the first thing I'll change.
Sonnet 5 as the judge. For a while the judge ran on Haiku too, to save time. But the judge is my safety net, and I want it smarter than the model it's checking. So it's back on Sonnet 5. I measured what that costs on my test questions (65 at the time). The check takes about 2.4 seconds instead of 1.8, and a full answer went from about 5 to 6 seconds. Sonnet is also pickier: it sent back 2 drafts for a rewrite where Haiku sent back 1. One extra second for a better safety net is a deal I'll take.
Search in memory. The search index lives in memory, inside the app. No extra service to run and no extra bill. Every time the app starts, it rebuilds the index from my notes, which takes about 4 seconds. My notes would need to grow a lot before that becomes a problem.
The browser holds the conversation. The server doesn't have to keep track of anyone's chat, which means that I needed to treat every single user input as untrusted. That includes the twin's own past replies, since they come back from your browser too. One of the bugs in part 2 came from forgetting that.
No streaming. The judge needs the whole draft before it can say yes, so the answer can't show up word by word. I really wanted that cool typing effect, but you won't get it here. Not yet, anyway.
The rest of the series
That's the happy path. The rest is where it gets fun:
- Part 2, keeping it out of trouble: a script that could have run up my API bill, fake conversations, eight messages from one click, and people signing their messages with my name.
- Part 3, when it gets it wrong: crashes, French questions that found nothing, and my own career turning into "I garbled that one".
- Part 4, how I test it: 77 hand-written test questions, scorecards, and what real visitors add.
- Part 5, swapping Haiku for Jev: a new kind of model for the first step, and answers more than a second faster.