{"slug": "character-ai-the-technical-story-of-what-keeps-users-hooked", "title": "Character.ai: The Technical Story of What Keeps Users Hooked", "summary": "Character.ai, co-founded by former Google engineers Noam Shazeer and Daniel De Freitas, keeps users hooked through a combination of technical optimizations including its Kaiju dense transformer models at 13B, 34B, and 110B parameters, multi-query attention, sliding-window attention, shared KV cache, INT8 quantization with quantization-aware training, and sticky sessions. The company's product features like Story Memory, Facts, Lorebook, and a 'good writing' scorecard, along with safety measures including supervised fine-tuning, online direct preference optimization, and classifier-guided beam search, are designed to sustain long, in-character role-play sessions.", "body_md": "[← AI Tools](/ai-tools)\n\n# Inside Character.ai: The Technical Story of What Keeps Users Hooked\n\n[Series · Inside AI Products](/series/inside-ai-products)\n\nBy[Manish Shahi](/about)Software Engineer • AI Developer\n\n## Table of contents(29)\n\nI run [manish.sh](https://manish.sh). I write about [AI](/glossary/artificial-intelligence) tools, how [LLMs](/glossary/large-language-model) behave, and the product work that makes people stay.\n\nThis post opens the ** Inside AI Products** series — I read what a company published, explain it in plain English, and tell the story in first person. If you prefer interviewing a model and then checking papers, that is\n\n**—**\n\n[Inside LLMs](/series/inside-llms)[Kimi](/writings/models/inside-kimi-k2-6-reverse-engineering-an-ai-assistant-by-interviewing-itself),\n\n[DeepSeek](/writings/models/inside-deepseek-reverse-engineering-an-ai-assistant-by-interviewing-itself),\n\n[Qwen](/writings/models/inside-qwen-3-8-max-preview-reverse-engineering-an-ai-assistant-by-interviewing-itself).\n\nI did not find [Character.ai](/glossary/character-ai) through a research paper.\n\nI typed **c.ai**.\n\nThe browser sent me to **character.ai**. Short domain, long rabbit hole. I tried a few characters, then one more turn, then another hour. The question for this series is simple: *what engineering makes people stay?*\n\n### 60-second TL;DR\n\n**How I found it:** Typed`c.ai`\n\n→ redirected to Character.ai → stayed for long role-play chats.**What it is:** A consumer chat app for talking to (and creating) characters — companions, fiction, practice — not a coding API.**Problem it solves:** Generic chatbots forget the plot and break character; Character.ai optimises for long, in-character sessions.**Why them:** Co-founded byand[Noam Shazeer](https://en.wikipedia.org/wiki/Noam_Shazeer?utm_source=manish.sh&utm_medium=blog&utm_campaign=inside-character-ai-the-technical-story-of-what-keeps-users-hooked&ref=manish.sh)**Daniel De Freitas**(ex-Google;[Transformer](/topics/transformer)+ Meena/LaMDA — see[Character.ai on Wikipedia](https://en.wikipedia.org/wiki/Character.ai?utm_source=manish.sh&utm_medium=blog&utm_campaign=inside-character-ai-the-technical-story-of-what-keeps-users-hooked&ref=manish.sh)). Chat and cheap serving were already their craft.**Model:**[Kaiju](/glossary/kaiju)— dense[transformers](/glossary/transformer)at 13B / 34B / 110B, tuned for speed and fun chats, not only exams.**Speed tricks:**[MQA](/glossary/multi-query-attention),[sliding-window attention](/glossary/sliding-window-attention), shared[KV cache](/glossary/kv-cache),[INT8](/glossary/int8)+[QAT](/glossary/quantization-aware-training), sticky sessions.**Stickiness:** Story Memory + Facts +[Lorebook](/glossary/lorebook)+ a “good writing” scorecard.**Safety:**[SFT](/glossary/sft)→ online[DPO](/glossary/dpo)→ classifier-guided[beam search](/glossary/beam-search).** Sources:**the ten Character.ai posts under[Research](#research)— company blogs, not an independent audit.\n\n### How to read this post\n\nRead it like a documentary. Each chapter usually goes:\n\n**Short scene**— what happened / what the company said** Hard words**— a** Like you are five**gloss when terms get dense** Picture or table**— mermaid diagram or quick comparison when it helps** Callout**— Key Idea / Why This Matters / Common Misconception (when useful)** Takeaway → Next**— one line to keep, then a hook into the next chapter\n\nBlue/[glossary](/glossary) links open fuller definitions. Jump around via [Chapters](#chapters). Sources sit at the bottom under [Research](#research).\n\n### Chapters\n\n**Part 1 — How I got here**\n\n[How I found the product](#how-i-found-the-product)[Introduction to the product](#introduction-to-the-product)[What problem it solves](#what-problem-it-solves)[What this series is](#what-this-series-is)[Who founded it — and why it works](#who-founded-it--and-why-it-works)\n\n**Part 2 — The model (Kaiju + training)**\n\n**Part 3 — Product loops**\n\n**Part 4 — Scale & open future**\n\n## How I found the product\n\nI was looking for a quick character chat. Someone had mentioned “c.ai” the way people mention a meme domain — short, easy to share, easy to type from memory.\n\nSo I opened `c.ai`\n\n.\n\nRedirect. Landing page. [Character.ai](/glossary/character-ai).\n\nThat redirect matters. Short domains get people through the door. What keeps them there is the chat: fast replies, a character that remembers the conversation, and a story that stays coherent for a long time.\n\nI live in India. I care about how these systems are built — not only demos. After enough long chats, I went looking for Character.ai’s own engineering posts. They publish a surprising amount: [Kaiju](https://blog.character.ai/inside-kaiju-building-conversational-models-at-scale/?utm_source=manish.sh&utm_medium=blog&utm_campaign=inside-character-ai-the-technical-story-of-what-keeps-users-hooked&ref=manish.sh), [inference](https://blog.character.ai/optimizing-ai-inference-at-character-ai-2/?utm_source=manish.sh&utm_medium=blog&utm_campaign=inside-character-ai-the-technical-story-of-what-keeps-users-hooked&ref=manish.sh), [Squinch](https://blog.character.ai/squinch/?utm_source=manish.sh&utm_medium=blog&utm_campaign=inside-character-ai-the-technical-story-of-what-keeps-users-hooked&ref=manish.sh), memory updates, Slonk, open-source fine-tuning. This article is my first-person walkthrough of that public trail.\n\n## Introduction to the product\n\n**Character.ai** is a consumer chat product: you talk to characters — celebrities, fictional people, tutors, custom personas you create — in a threaded chat UI. Sign-up is light (Google / Apple / email). The homepage pitch is millions of characters and “start in seconds.”\n\nIt is **not** primarily a developer API playground or a coding assistant. The unit of value is a **session**: many turns with one character, often role-play or companion chat, sometimes study or practice.\n\n**Like you are five:** Imagine a huge costume party of chatbots. You pick one costume (a character), then talk for as long as the story stays fun.\n\nUnder the hood (from their public posts) that product sits on:\n\n| Layer | What you feel | What they talk about publicly |\n|---|---|---|\n| Model | Replies feel “in character” |\n|\n\n[KV cache](/glossary/kv-cache),[MQA](/glossary/multi-query-attention),[INT8](/glossary/int8)serving[Lorebook](/glossary/lorebook)[SFT](/glossary/sft)→[DPO](/glossary/dpo)→ guided decode## What problem it solves\n\nMost general chatbots are built for short Q&A: ask, answer, done. Character chat fails that pattern in three boring ways:\n\n**Continuity**— After twenty turns the bot forgets your name, the plot, or the tone.** Character**— Replies drift into generic “helpful assistant” voice.** World**— Shared lore (places, rules, side characters) either vanishes or has to be pasted into every prompt.\n\nCharacter.ai’s public story is a bet that people want **companions and stories**, not only answers. So the product problem is:\n\nKeep a fictional (or companion) conversation\n\nalivefor a long time — fast enough to stay immersive, consistent enough to feel real, safe enough to ship at scale.\n\n**Like you are five:** A good bedtime storyteller remembers what happened yesterday. A bad one asks your name every page. Character.ai is trying to be the first kind — for millions of chats at once.\n\nThat is why later chapters obsess over sticky [KV caches](/glossary/kv-cache), Story Memory, [Lorebook](/glossary/lorebook), and a “compelling writing” scorecard instead of only MMLU points.\n\n## What this series is\n\n**Inside LLMs** asks a model about itself, then checks papers.\n\n**Inside AI Products** asks: *what did the company say it built, and how does that explain how the app feels?*\n\nCharacter.ai is a good first pick because:\n\n- People live in the product, not only in an API playground.\n- Their public posts cover\n**model → serving → memory → infra**. - The hook is clear: long role-play and companion chats.\n\nI will keep it technical but easy. New terms get a short definition and a [glossary](/glossary) link.\n\n## Who founded it — and why it works\n\n[Character.ai](https://en.wikipedia.org/wiki/Character.ai?utm_source=manish.sh&utm_medium=blog&utm_campaign=inside-character-ai-the-technical-story-of-what-keeps-users-hooked&ref=manish.sh) was co-founded by ** Noam Shazeer** and\n\n**Daniel De Freitas**, both ex-Google. (Shazeer has his own Wikipedia page; De Freitas is mainly on the\n\n[Character.ai](https://en.wikipedia.org/wiki/Character.ai?utm_source=manish.sh&utm_medium=blog&utm_campaign=inside-character-ai-the-technical-story-of-what-keeps-users-hooked&ref=manish.sh)page and\n\n[Wikidata](https://www.wikidata.org/wiki/Q130363750?utm_source=manish.sh&utm_medium=blog&utm_campaign=inside-character-ai-the-technical-story-of-what-keeps-users-hooked&ref=manish.sh).)\n\nThis is not a random founder story. It explains why they could build a chat product that holds under heavy load.\n\n** Noam Shazeer** helped write\n\n[— the paper behind the](https://arxiv.org/abs/1706.03762?utm_source=manish.sh&utm_medium=blog&utm_campaign=inside-character-ai-the-technical-story-of-what-keeps-users-hooked&ref=manish.sh)\n\n*Attention Is All You Need*[Transformer](/glossary/transformer). At Google he also worked on big serving systems. Later Character.ai ideas like\n\n[MQA](/glossary/multi-query-attention), smaller\n\n[KV caches](/glossary/kv-cache), and\n\n[Squinch](/glossary/squinch)fit that same goal: make big models cheaper to run.\n\n**Like you are five:** Shazeer helped invent the popular “brain recipe” (Transformer) and also cared about making the recipe run without melting the kitchen (serving cost).\n\n** Daniel De Freitas** helped design Google’s chat models that started as\n\n**Meena** and became\n\n**. In short: years of dialogue work, not only demos.**\n\n[LaMDA](https://blog.google/technology/ai/lamda/?utm_source=manish.sh&utm_medium=blog&utm_campaign=inside-character-ai-the-technical-story-of-what-keeps-users-hooked&ref=manish.sh)**Like you are five:** De Freitas spent years teaching computers to hold a conversation, not only finish a sentence on a quiz.\n\nThey left Google around **2021** and built Character.ai end to end — train or fine-tune, tune for character chat, then ship to users worldwide. In **2024** both went back to Google DeepMind in a licensing deal; Character.ai kept running with the remaining team. This post still follows the public engineering they left behind.\n\nSo when they talk about INT8 [inference](/glossary/inference), sticky KV caches, and ~20,000 queries per second, it does not sound like empty marketing. They already knew the hard truth: **chat models stay expensive unless you invent serving tricks**.\n\n## Kaiju: the in-house LLM family\n\nMain source: [Inside Kaiju](https://blog.character.ai/inside-kaiju-building-conversational-models-at-scale/?utm_source=manish.sh&utm_medium=blog&utm_campaign=inside-character-ai-the-technical-story-of-what-keeps-users-hooked&ref=manish.sh).\n\n** Kaiju** is Character.ai’s own family of large language models.\n\n**Like you are five:**\n\n- An\nis a very big word-guessing brain. You type words; it guesses what comes next, again and again, until a reply appears.[LLM](/glossary/large-language-model) - A\nis that brain’s wiring — lots of tiny number dials connected together.[neural network](/glossary/neural-network) - A\nis a tiny text brick (a short word or part of a word).[token](/glossary/token) are those number dials. Training turns the dials so guesses get better.[Parameters](/glossary/parameters)/[weights](/glossary/weights)\n\nThree production sizes:\n\n| Variant | Size |\n|---|---|\n| Small | 13 billion parameters |\n| Medium | 34B |\n| Large | 110B |\n\n**Like you are five:** “13 billion parameters” means the model has a huge pile of dials — enough to remember lots of language patterns. Bigger piles can be smarter, but they also cost more to run.\n\nMany labs chase exam leaderboards first. Character.ai says Kaiju was built for **fast, fun, safer chats**. Think “does this feel good at 2am in a long role-play?” more than “can we gain one more exam point?”\n\nKaiju is a **dense** [transformer](/glossary/transformer).\n\n**Like you are five:**\n\n- A\n**transformer** is the “brain recipe” most chat AIs use. It reads your words, decides what matters, and guesses the next little piece of text. - A\n**token** is one of those little pieces — often a short word or part of a word. **Dense** means: for every next guess, the whole kitchen stays open. Every cook (every layer) works on the order.**Feed-forward block** is one workbench inside each layer — the place that mixes and reshapes what the model just “heard.” In a dense model, that whole workbench stays on.**Sparse** is a different kitchen. You hire many specialist cooks (“experts”), but for each token only a few wake up. The others sleep. That can save cost, but it is a different design.[Mixture of Experts](/glossary/moe)(MoE)\n\nKaiju uses the dense style. That fits the founders’ background — they helped popularise dense transformers, and Kaiju follows that path.\n\n``` php\nflowchart TD\n  A[\"Your message\"] --> B[\"Kaiju dense transformer\"]\n  B --> C[\"Next-token stream\"]\n  C --> D[\"Safety-aware decode\"]\n  D --> E[\"Reply in chat\"]\n\n  classDef box fill:#f7f5f0,stroke:#2a2a2a,color:#2a2a2a,stroke-width:1px\n  class A,B,C,D,E box\n```\n\n## Architecture tricks that make chats feel instant\n\nLong chats cost a lot. A big reason is the ** KV cache**.\n\n**Like you are five:**\n\nis “what should I look at right now?” — like pointing at the important words in a story.[Attention](/glossary/attention-mechanism)**Key / Value (K and V)** are sticky notes the model saves about words it already saw, so it can look them up later.- A\nis the box of those sticky notes. Without it, the model would rewrite every sticky note from scratch every turn (slow and expensive).[KV cache](/glossary/kv-cache) means “using the trained model to chat” — not the training homework phase.[Inference](/glossary/inference)\n\nCharacter.ai worked hard to make that sticky-note box smaller and cheaper.\n\n### 1. Multi-Query Attention (MQA)\n\nNormal multi-head attention gives every head its own Keys and Values. ** MQA** shares K and V across heads. Less to store → smaller KV cache → more chats per GPU and faster inference.\n\n**Like you are five:** Imagine many flashlights (heads) reading a page. Old way: each flashlight keeps its own photocopy of the page. **MQA**: they share one photocopy. Same reading, less paper in the backpack.\n\n### 2. Sliding Window Attention\n\nMost layers only look at a recent window (**1024 tokens** in their write-up), with a few full “global” layers mixed in (about 5:1 or 6:1). That is ** sliding-window attention**: cheaper on long chats, still some long-range reach. Current Kaiju does not use\n\n[attention sinks](/glossary/attention-sink).\n\n**Like you are five:**\n\n**Sliding window**= mostly read only the last few pages of the story, not the whole library every time.** Global layers**= every so often, peek at the whole book so you do not forget the beginning.= a trick where the very first words keep grabbing attention forever. Kaiju says it does not lean on that trick right now.[Attention sink](/glossary/attention-sink)\n\n### 3. Cross-layer KV cache sharing\n\nAbout **2–3 neighbouring layers** share one KV cache. Same idea as MQA, applied up the stack: less memory, little quality loss in their reports.\n\n**Like you are five:** Floors 3, 4, and 5 of the building share one filing cabinet instead of each buying their own.\n\n### 4. INT8 + Quantization-Aware Training\n\nWeights, activations, and KV values use ** INT8**. On modern chips, INT8 matmuls are roughly\n\n**2×** faster than bf16-style compute. They train with\n\n**so the model learns to stay accurate at that precision — not only**\n\n[QAT](/glossary/quantization-aware-training)[quantize](/glossary/quantization)at the end and hope.\n\nThey also name **pre-layer RMSNorm** (normalise before the heavy work) and\n\n**dynamic clamping** so low precision stays stable.\n\n**Like you are five:**\n\n= draw with fewer crayons. A fine picture uses many colour shades; a smaller box of crayons is lighter to carry.[Quantization](/glossary/quantization)= an 8-crayon box for numbers (small and fast).[INT8](/glossary/int8)**bf16 / FP32**= bigger crayon boxes (more detail, heavier).** Matmul**= the model’s “times tables” on huge number grids — the heavy gym work GPUs do.= practise drawing with the small crayon box[QAT](/glossary/quantization-aware-training)*during training*, so the picture still looks good later.= gently rescale the numbers so they do not get too wild before the next step.[RMSNorm](/glossary/rmsnorm)**Dynamic clamping**= “no number may jump higher than this fence” — and the fence can move a bit so training stays safe.\n\n| Trick | What you feel in chat |\n|---|---|\n| MQA | Smaller memory per token |\n| Sliding window | Long chats cost less |\n| Cross-layer KV share | Even less KV memory |\n| INT8 + QAT | Faster replies without a quality cliff |\n\n## Training stack: Squinch and friends\n\n[Inside Kaiju](https://blog.character.ai/inside-kaiju-building-conversational-models-at-scale/?utm_source=manish.sh&utm_medium=blog&utm_campaign=inside-character-ai-the-technical-story-of-what-keeps-users-hooked&ref=manish.sh) covers training briefly. The longer story is [Optimizing Large-Scale Pretraining](https://blog.character.ai/squinch/?utm_source=manish.sh&utm_medium=blog&utm_campaign=inside-character-ai-the-technical-story-of-what-keeps-users-hooked&ref=manish.sh) — written after they stopped big pretraining and started using open-source bases more.\n\n**Like you are five:**\n\n**Pretraining**= the long “read lots of text” homework before the model becomes a chat buddy.** Open-source base**= someone else’s finished homework you may reuse, then add your own tutoring.** GPU / H100**= a very fast number engine.\n\nBefore that shift, the early training team (led in part by Shazeer) shared a few practical tricks.\n\n**Hardware.** NVIDIA H100 GPUs. [Model parallelism](/glossary/model-parallelism) inside a node (tensor + sequence). ** FSDP** across nodes.\n\n**Like you are five:** The giant Lego set does not fit on one table, so friends each hold a piece (**model parallelism**). ** FSDP** also shares the instruction booklet and bricks across tables.\n\n**Mixed precision.** Rough map:\n\n| Piece | Precision |\n|---|---|\n| Forward weights / KV | INT8 |\n| Activations / local gradients |\n|\n\n**Like you are five:** Cheap crayons for most drawing; keep one fancy master colouring book (FP32) so colours stay neat.\n\n** Squinch.** A 6-bit way to compress gradients by Noam Shazeer. Gradients move between machines in small blocks (8 numbers each). Less network traffic, still accurate enough.\n\n**Like you are five:** A **gradient** is a “turn this dial left/right” note after practice. **Squinch** folds many notes into a tiny postcard so machines can mail them fast.\n\nOther names from that era: **Attention Z-Reg** (keeps attention scores stable in bfloat16), **Visibility Mask** (compact rules for what can attend to what in packed chat data), **Virtual Scalars (Bungee)** (helps small INT8 models stay stable), and **ternary weight updates** (send only 0 / 1 / −1 in some cases).\n\n**Like you are five:**\n\n**Attention Z-Reg**= do not let attention numbers get so huge they go blurry.** Visibility Mask**= a seating chart for which words may look at which words.** Bungee**= elastic bands so tiny INT8 models do not snap.** Ternary updates**= sometimes only shout “stay / up / down.”\n\n**Data.** Two mixes: “MMLU Max” for exams, “Production Max” for engagement. Web text, code, synthetic data — then ** pretraining annealing** near the end (shift toward higher-quality / instruction data).\n\n**Like you are five:** Quiz books vs fun-chat practice. **Annealing** = near the end, switch to cleaner practice so good manners stick.\n\n``` php\nflowchart LR\n  A[\"Data mix\"] --> B[\"Dense transformer train\"]\n  B --> C[\"Squinch grads across nodes\"]\n  C --> D[\"QAT-aware INT8 model\"]\n  D --> E[\"Annealing → post-training\"]\n\n  classDef box fill:#f7f5f0,stroke:#2a2a2a,color:#2a2a2a,stroke-width:1px\n  class A,B,C,D,E box\n```\n\n## Safety pipeline\n\nAlso from [Inside Kaiju](https://blog.character.ai/inside-kaiju-building-conversational-models-at-scale/?utm_source=manish.sh&utm_medium=blog&utm_campaign=inside-character-ai-the-technical-story-of-what-keeps-users-hooked&ref=manish.sh). Before a model goes live:\n\non safety and instruction data.[SFT](/glossary/sft)- Modified online\nfrom swipe / feedback data — a lighter path than classic[DPO](/glossary/dpo)[RLHF](/glossary/rlhf). **Classifier training**, with an optional head that can score safety per** token**.\n\nAt [inference](/glossary/inference) they use **classifier-guided beam search**: try a few candidate replies, steer with safety scores, keep the chat creative without going fully bland.\n\n**Like you are five:**\n\n= show many good example answers and say “copy this style.”[SFT](/glossary/sft)**Swipe / feedback**= users tapping “this reply was better / worse.”= learn from those better/worse picks without building a giant separate “score teacher” first.[DPO](/glossary/dpo)= a heavier version of “humans score answers, then the model practises to please the score.”[RLHF](/glossary/rlhf)**Classifier**= a safety referee that gives thumbs up / down signals.= keep a few possible next sentences in play, not only one, then pick a good safe path.[Beam search](/glossary/beam-search)**Classifier-guided**= the referee helps choose among those paths.\n\n``` php\nflowchart TD\n  A[\"Base Kaiju\"] --> B[\"SFT safety + instructions\"]\n  B --> C[\"Online DPO from swipes\"]\n  C --> D[\"Safety classifier head\"]\n  D --> E[\"Classifier-guided beam search\"]\n  E --> F[\"User-facing reply\"]\n\n  classDef box fill:#f7f5f0,stroke:#2a2a2a,color:#2a2a2a,stroke-width:1px\n  class A,B,C,D,E,F box\n```\n\n## Memory and Lorebook\n\nProduct meets habit here. Two posts matter most: [Smarter Memory for Smarter Chats](https://blog.character.ai/memory/?utm_source=manish.sh&utm_medium=blog&utm_campaign=inside-character-ai-the-technical-story-of-what-keeps-users-hooked&ref=manish.sh) and [April Update: New Model, Memory, and Lorebook](https://blog.character.ai/pipsqueak2-and-more/?utm_source=manish.sh&utm_medium=blog&utm_campaign=inside-character-ai-the-technical-story-of-what-keeps-users-hooked&ref=manish.sh).\n\n**Like you are five:** Chat AIs forget easily. Memory features are sticky notes so the character does not ask “what is your name?” every five minutes.\n\n### Story Memory and Facts\n\n**Story Memory** tracks the plot and emotional thread. You can pin what matters so it is less likely to drop. **Facts** cover appearance, quirks, relationships, hobbies, and similar stable details. The Memory page got a cleaner layout, more categories (hairstyle, eye colour, quirks, and more), in-chat “saved” notices, and **Memory Usage** bars so you see when space is running low.\n\nThese upgrades were meant to run on newer styles like **PipSqueak 2 (PSQ2)** and improved **DeepSqueak** — more consistent, less drift, fewer cutoffs. DeepSqueak 2 was said to be training in the April update.\n\n**Like you are five:**\n\n**Story Memory**= “what happened in our adventure.”** Facts**= “your hair is blue; you love mystery books.”** Pin**= a gold star sticky note that should not get thrown away.** Memory Usage bar**= a juice glass — when it is almost full, old notes may get squeezed shorter.** PipSqueak / DeepSqueak**= names of chat “engines” (model styles) on the app — like different car engines under the same car body.\n\n### Lorebook\n\n** Lorebook** was the community’s #1 request for a long time. In the April update, Character.ai described it as a reusable\n\n**keyword-triggered** world database — lore that appears when the chat mentions that place, person, or rule.\n\nYou write entries for locations, side characters, rules, items. They only load when keywords match. So you can share one world across characters without stuffing everything into every prompt. c.ai+ creators got first access, then wider rollout.\n\n**Like you are five:** Lorebook is a backpack of world cards. You only pull out the “dragon castle” card when someone says “dragon” or “castle” — you do not dump the whole backpack on the table every turn.\n\n| Feature | Job |\n|---|---|\n| Story Memory | Keep the plot |\n| Facts | Keep stable details |\n| Memory Usage | See what is filling space |\n| Lorebook | Load world lore only when needed |\n\n## Compelling writing evaluation\n\nMost labs score exams and coding. Character.ai’s post [Evaluating Our Models Using Principles of Compelling Writing](https://blog.character.ai/evaluating-our-models-using-principles-of-compelling-writing/?utm_source=manish.sh&utm_medium=blog&utm_campaign=inside-character-ai-the-technical-story-of-what-keeps-users-hooked&ref=manish.sh) takes another path: a **Compelling Writing Evaluation Framework** built with professional writers.\n\nThey break “good chat writing” into measurable pieces: plot shapes (like the Hero’s Journey), archetypes, pacing, “show don’t tell”, dialogue quality, emotional intelligence, and novelty. Offline [LLM](/glossary/large-language-model) judges score writer-labelled data; online user ratings sit beside that.\n\n**Like you are five:**\n\n**Exam benches (like MMLU)**= quiz marks.** Compelling writing eval**= “was this story fun and clear?” report card.** LLM-as-judge**= another AI teacher marks homework with the writers’ checklist.** Hero’s Journey**= a famous adventure shape (leave home → trouble → grow → return).\n\nSimple reason: if people role-play for hours, **story quality is a real metric**, not a side hobby.\n\n## Inference at scale\n\nCharacter.ai has said they serve about **20,000 queries per second** (see their earlier [inference](https://blog.character.ai/optimizing-ai-inference-at-character-ai-2/?utm_source=manish.sh&utm_medium=blog&utm_campaign=inside-character-ai-the-technical-story-of-what-keeps-users-hooked&ref=manish.sh) posts). That only works if [prefill and decode](/glossary/prefill-decode) stay cheap.\n\n**Like you are five:**\n\n**Queries per second**= how many “please reply” knocks arrive every second.= read your whole message first (set the table).[Prefill](/glossary/prefill-decode)**Decode**= say the reply word by word (serve the food).** Kernel**= a tiny ultra-fast recipe written for the GPU.= a clever way to do attention with less memory traffic.[FlashAttention](/glossary/flashattention)= send you back to the same waiter so your sticky-note box (KV cache) is already there.[Sticky session](/glossary/sticky-session)= how long until the first word appears.[TTFT](/glossary/ttft)= how long between the next words.[TPOT](/glossary/tpot)= an even smaller crayon box than INT8 for some serving work.[FP8](/glossary/fp8)**DP / TP / EP**= ways to split one giant model across many GPUs (data / tensor / expert splits).= a popular engine for serving chat models efficiently.[vLLM](/glossary/vllm)**p90**= “for 90 out of 100 requests, we were at least this fast” — a fairness check on speed.\n\nFrom outside, traffic tools show the same habit: huge monthly visits and long sessions. Here is a Similarweb snapshot I saved (an estimate, not Character.ai’s official numbers). Notice visit length and pages per visit — people stay and chat.\n\nServing ingredients they name:\n\n- Custom\n**INT8 attention kernels** [FlashAttention](/glossary/flashattention)changes for MQA- Parallel work across query heads (~10–30% faster in prefill/decoding vs older kernels)\n— same chat stays on the same server so KV cache hits stay high (they have cited ~95%)[Sticky sessions](/glossary/sticky-session)\n\nThey also posted [Technical Deep Dive: DigitalOcean + AMD, 2× production inference](https://blog.character.ai/technical-deep-dive-how-digitalocean-and-amd-delivered-a-2x-production-inference-performance-increase-for-character-ai/?utm_source=manish.sh&utm_medium=blog&utm_campaign=inside-character-ai-the-technical-story-of-what-keeps-users-hooked&ref=manish.sh). On MI300X / MI325X, for a large [MoE](/glossary/moe) like [Qwen3](/topics/qwen3)-235B, they report about **2×** throughput with DP2/TP4/EP4, [FP8](/glossary/fp8), careful GPU placement, [vLLM](/glossary/vllm) tuning, and strict p90 ** TTFT** /\n\n**targets — with cost per token down by a similar factor in that write-up.**\n\n[TPOT](/glossary/tpot)| Metric | What you feel |\n|---|---|\n| TTFT | How fast the first word appears |\n| TPOT | How smooth the stream feels after that |\n| Sticky KV hits | Less cold start mid-chat |\n\n## Slonk: Slurm on Kubernetes\n\nFrom [Slonk](https://blog.character.ai/slonk/?utm_source=manish.sh&utm_medium=blog&utm_campaign=inside-character-ai-the-technical-story-of-what-keeps-users-hooked&ref=manish.sh):\n\nResearchers like ** Slurm** (\n\n`sbatch`\n\n, `squeue`\n\n, fair share). Ops teams like **for healing and containers.**\n\n[Kubernetes](/glossary/kubernetes)**Slonk** = Slurm on Kubernetes. Familiar research commands on resilient K8s StatefulSets (controller, workers, login pods), health checks, auto-fix, Git-sync, and a custom operator for machines. Training and inference can share the pool; production can take priority when needed.\n\n**Like you are five:**\n\n= the school timetable for GPU homework (“your turn on the big computer”).[Slurm](/glossary/slurm)= robot caretakers that restart broken classroom PCs and keep apps alive.[Kubernetes](/glossary/kubernetes)**Slonk**= keep the familiar timetable, but let the robot caretakers run the building.** StatefulSet / pod**= a named computer seat that Kubernetes can replace if it breaks.** Preempt**= “production chat needs this GPU now — training homework, please pause.”\n\nIf you have seen a cluster suffer because the scheduler and the container layer do not agree, this design will feel familiar.\n\n## Open-source shift\n\nIn [Breaking News: Our Open-Source Models Are A Lot of Fun!](https://blog.character.ai/breaking-news-our-open-source-models-are-a-lot-of-fun/?utm_source=manish.sh&utm_medium=blog&utm_campaign=inside-character-ai-the-technical-story-of-what-keeps-users-hooked&ref=manish.sh), Character.ai explained the move: take open-source bases, tune them for fun and emotional quality, and ship them as Character models — using their user feedback loop, smarter prompting, and post-training ([SFT](/glossary/sft), [DPO](/glossary/dpo), RL, [QAT](/glossary/quantization-aware-training)).\n\n**Like you are five:**\n\n**Open-source model**= a shared toy brain you can download and improve.** Post-training**= extra tutoring after the big homework (SFT / DPO / RL / QAT).** RL**= practise by trial and reward (“that reply earned a cookie”).** Ensemble inference**= ask a small team of models / settings and blend the vibe.** Retention**= people coming back tomorrow.\n\nTheir own test numbers:\n\n**22%** more time spent and**13%** more sessions**14%** better retention for lightly engaged users**60%** more messages rated positive (plus jumps in “funny”, “interesting”, “helpful”)\n\nThen they rolled out **PipSqueak** more widely. Treat the numbers as company launch metrics — but the direction is clear: entertainment over pure exam scores.\n\nWhether the base is Kaiju or a fine-tuned open model, the product still needs memory, safety, and fast serving.\n\n## pipeling-sft and Agent SDK\n\nTwo posts show what comes after pure in-house pretraining.\n\n### pipeling-sft\n\n[Character.AI Open Sources pipeling-sft](https://blog.character.ai/character-ai-open-sources-pipeling-sft-a-scalable-framework-for-fine-tuning-moe-llms-like-deepseek-v3/?utm_source=manish.sh&utm_medium=blog&utm_campaign=inside-character-ai-the-technical-story-of-what-keeps-users-hooked&ref=manish.sh) is a toolkit for full-parameter [SFT](/glossary/sft) of large [MoE](/glossary/moe) models ([DeepSeek](/topics/deepseek) V3–class). It supports multi-level parallelism, BF16 and experimental FP8, Hugging Face checkpoints, and helpers for stable long runs. Still experimental — meant for teams who do not want to rebuild MoE fine-tuning from scratch.\n\n**Like you are five:** A shared recipe book for teaching a giant “many-cooks” (MoE) brain new manners, without inventing the oven from scratch.\n\n### Agent SDK\n\n[Agent SDK: Teaching Machines to Use Tools with Claude Code](https://blog.character.ai/agent-sdk/?utm_source=manish.sh&utm_medium=blog&utm_campaign=inside-character-ai-the-technical-story-of-what-keeps-users-hooked&ref=manish.sh) looks past pure chat. Their backend is Go; Anthropic’s Claude Code Agent SDK was Python-first — so they built a Go port with tools, streaming, hooks, [MCP](/glossary/mcp), and SSE. Internal tools (they mention AgentX) use it for multi-step work, provider fallbacks, and quality checks. Direction: products that can **use tools**, not only talk.\n\n**Like you are five:**\n\n**SDK**= a toolbox of ready-made Lego bricks for builders.= the AI can press buttons (search, generate, check) instead of only talking.[Tool calling](/topics/tool-calling)= a shared plug standard so tools fit many AIs.[MCP](/glossary/mcp)**SSE / streaming**= words (and events) arrive like a live ticker, not one big sealed envelope.** Go / Python**= two different building languages; their servers speak Go.\n\n## Why users stay hooked\n\nBack to that first `c.ai`\n\nredirect.\n\nI stayed because the chat felt **alive enough** — fast enough to keep immersion, consistent enough to remember the plot, creative enough that the next turn felt worth typing.\n\nThe public engineering story maps to that feeling:\n\n``` php\nflowchart TD\n  A[\"Fast INT8 + MQA + sticky KV\"] --> E[\"Immersion holds\"]\n  B[\"Story Memory + Facts + Lorebook\"] --> E\n  C[\"Compelling writing eval\"] --> E\n  D[\"SFT + DPO + guided beam safety\"] --> E\n  E --> F[\"Long sessions / return visits\"]\n\n  classDef box fill:#f7f5f0,stroke:#2a2a2a,color:#2a2a2a,stroke-width:1px\n  class A,B,C,D,E,F box\n```\n\nCharacter.ai is not only “building models.” In their words, they are building **companions and stories** — and every trick from Squinch to Lorebook serves that loop.\n\nIf you have used it: what keeps you coming back — speed, memory, a favourite character voice, or Lorebook worlds? Do tell — it will help shape the next Inside AI Products entry.\n\n## Research\n\nPrimary Character.ai sources for this walkthrough (checked live):\n\n[Inside Kaiju — building conversational models at scale](https://blog.character.ai/inside-kaiju-building-conversational-models-at-scale/?utm_source=manish.sh&utm_medium=blog&utm_campaign=inside-character-ai-the-technical-story-of-what-keeps-users-hooked&ref=manish.sh)— in-house LLM architecture; conversation over benchmarks; MQA; sliding-window attention; KV cache sharing; INT8 + QAT; training overview; safety pipeline.[Optimizing Large-Scale Pretraining at Character.AI](https://blog.character.ai/squinch/?utm_source=manish.sh&utm_medium=blog&utm_campaign=inside-character-ai-the-technical-story-of-what-keeps-users-hooked&ref=manish.sh)— distributed training; Squinch; data/training optimisations; scaling transformers (pre–open-source-foundations era).[Evaluating Our Models Using Principles of Compelling Writing](https://blog.character.ai/evaluating-our-models-using-principles-of-compelling-writing/?utm_source=manish.sh&utm_medium=blog&utm_campaign=inside-character-ai-the-technical-story-of-what-keeps-users-hooked&ref=manish.sh)— fiction-writer-style eval instead of only MMLU/coding benches.[Breaking News: Our Open-Source Models Are A Lot of Fun!](https://blog.character.ai/breaking-news-our-open-source-models-are-a-lot-of-fun/?utm_source=manish.sh&utm_medium=blog&utm_campaign=inside-character-ai-the-technical-story-of-what-keeps-users-hooked&ref=manish.sh)— shift to open-source foundations; fine-tuning for entertainment; engagement metrics.[Smarter Memory for Smarter Chats](https://blog.character.ai/memory/?utm_source=manish.sh&utm_medium=blog&utm_campaign=inside-character-ai-the-technical-story-of-what-keeps-users-hooked&ref=manish.sh)— Story Memory; Facts; Memory Usage; long chats.[April Update: New Model, Memory, and Lorebook](https://blog.character.ai/pipsqueak2-and-more/?utm_source=manish.sh&utm_medium=blog&utm_campaign=inside-character-ai-the-technical-story-of-what-keeps-users-hooked&ref=manish.sh)— PipSqueak 2 / DeepSqueak; memory UX; Lorebook.[Technical Deep Dive: How DigitalOcean and AMD Delivered a 2x Production Inference Performance Increase for Character.ai](https://blog.character.ai/technical-deep-dive-how-digitalocean-and-amd-delivered-a-2x-production-inference-performance-increase-for-character-ai/?utm_source=manish.sh&utm_medium=blog&utm_campaign=inside-character-ai-the-technical-story-of-what-keeps-users-hooked&ref=manish.sh)— GPU / inference serving; production latency; AMD + DigitalOcean.[Slonk: Slurm on Kubernetes for ML Research at Character.ai](https://blog.character.ai/slonk/?utm_source=manish.sh&utm_medium=blog&utm_campaign=inside-character-ai-the-technical-story-of-what-keeps-users-hooked&ref=manish.sh)— Kubernetes; GPU scheduling; research clusters.[Character.AI Open Sources pipeling-sft](https://blog.character.ai/character-ai-open-sources-pipeling-sft-a-scalable-framework-for-fine-tuning-moe-llms-like-deepseek-v3/?utm_source=manish.sh&utm_medium=blog&utm_campaign=inside-character-ai-the-technical-story-of-what-keeps-users-hooked&ref=manish.sh)— fine-tuning MoE models (DeepSeek V3–class); SFT pipeline.[Agent SDK: Teaching Machines to Use Tools with Claude Code](https://blog.character.ai/agent-sdk/?utm_source=manish.sh&utm_medium=blog&utm_campaign=inside-character-ai-the-technical-story-of-what-keeps-users-hooked&ref=manish.sh)— tool-using agents beyond pure conversation ([GitHub: character-ai/claude-agent-sdk-go](https://github.com/character-ai/claude-agent-sdk-go?utm_source=manish.sh&utm_medium=blog&utm_campaign=inside-character-ai-the-technical-story-of-what-keeps-users-hooked&ref=manish.sh)).\n\nRelated context (founders / architecture lineage, not in the list above):\n\n[Attention Is All You Need](https://arxiv.org/abs/1706.03762?utm_source=manish.sh&utm_medium=blog&utm_campaign=inside-character-ai-the-technical-story-of-what-keeps-users-hooked&ref=manish.sh)· Google[LaMDA](https://blog.google/technology/ai/lamda/?utm_source=manish.sh&utm_medium=blog&utm_campaign=inside-character-ai-the-technical-story-of-what-keeps-users-hooked&ref=manish.sh)· Wikipedia[Noam Shazeer](https://en.wikipedia.org/wiki/Noam_Shazeer?utm_source=manish.sh&utm_medium=blog&utm_campaign=inside-character-ai-the-technical-story-of-what-keeps-users-hooked&ref=manish.sh)·[Character.ai](https://en.wikipedia.org/wiki/Character.ai?utm_source=manish.sh&utm_medium=blog&utm_campaign=inside-character-ai-the-technical-story-of-what-keeps-users-hooked&ref=manish.sh)· Wikidata[Daniel De Freitas](https://www.wikidata.org/wiki/Q130363750?utm_source=manish.sh&utm_medium=blog&utm_campaign=inside-character-ai-the-technical-story-of-what-keeps-users-hooked&ref=manish.sh)- Earlier serving write-ups that mention ~20k QPS / sticky KV caching:\n[Optimizing AI Inference](https://blog.character.ai/optimizing-ai-inference-at-character-ai-2/?utm_source=manish.sh&utm_medium=blog&utm_campaign=inside-character-ai-the-technical-story-of-what-keeps-users-hooked&ref=manish.sh)·[Part Deux](https://blog.character.ai/optimizing-ai-inference-at-character-ai-part-deux-2/?utm_source=manish.sh&utm_medium=blog&utm_campaign=inside-character-ai-the-technical-story-of-what-keeps-users-hooked&ref=manish.sh)\n\nAlso used in this post:\n\n- Series:\n[Inside AI Products](/series/inside-ai-products)· sister series:[Inside LLMs](/series/inside-llms) - Site:\n[manish.sh](https://manish.sh) - Images:\n[homepage](/images/blog/character_ai/homepage.png)·[Similarweb traffic snapshot](/images/blog/character_ai/similarweb-traffic-assumption.png)· concept diagrams:[dense vs MoE](/images/blog/character_ai/kaiju-dense-vs-moe.png)·[KV / MQA](/images/blog/character_ai/kaiju-kv-mqa.png)·[safety pipeline](/images/blog/character_ai/character-safety-pipeline.png)·[memory + Lorebook](/images/blog/character_ai/character-memory-lorebook.png)·[habit loop](/images/blog/character_ai/character-habit-loop.png) - Related glossary:\n[Character.ai](/glossary/character-ai)·[Kaiju](/glossary/kaiju)·[LLM](/glossary/large-language-model)·[Transformer](/glossary/transformer)·[MQA](/glossary/multi-query-attention)·[Sliding Window Attention](/glossary/sliding-window-attention)·[KV Cache](/glossary/kv-cache)·[INT8](/glossary/int8)·[QAT](/glossary/quantization-aware-training)·[RMSNorm](/glossary/rmsnorm)·[Squinch](/glossary/squinch)·[DPO](/glossary/dpo)·[SFT](/glossary/sft)·[Beam Search](/glossary/beam-search)·[Lorebook](/glossary/lorebook)·[TTFT](/glossary/ttft)·[TPOT](/glossary/tpot)·[Slurm](/glossary/slurm)·[Kubernetes](/glossary/kubernetes)·[vLLM](/glossary/vllm)·[FlashAttention](/glossary/flashattention)·[FSDP](/glossary/fsdp)·[MoE](/glossary/moe)·[MCP](/glossary/mcp)·[full glossary](/glossary)\n\nPinch or double-tap to zoom · tap outside to close\n\nFAQ\n\n## Frequently asked questions\n\n## How did you end up writing about Character.ai?\n\nI typed c.ai, the domain redirected me to character.ai, and I stayed for the long role-play chats. This post unpacks the public engineering behind that habit. I write from India.\n\n## Is this an interview of the model?\n\nNo. Unlike the Inside LLMs series, this piece is built from Character.ai’s public technical and product blog posts — not from interviewing a chat model about itself.\n\n## What is Kaiju?\n\nKaiju is Character.ai’s own LLM family: dense transformers in Small (13B), Medium (34B), and Large (110B), tuned for fast, fun, safer chats.\n\n## What is Lorebook?\n\nA keyword-triggered world database. Entries for places, side characters, rules, and items only load when matching keywords appear — shared worlds without stuffing every prompt.\n\n## Why do chats feel so sticky?\n\nFast replies (KV cache + INT8 tricks), Story Memory and Facts for continuity, Lorebook for world consistency, and a writing scorecard aimed at good dialogue — not only exam scores.\n\n## Who founded Character.ai?\n\nNoam Shazeer and Daniel De Freitas, both ex-Google. Shazeer co-authored the Transformer paper; together they worked on chat models that became Meena / LaMDA. That is why a full chat stack at this scale makes sense.\n\n## Where do the numbers come from?\n\nFrom Character.ai’s public engineering posts listed in the Research section (Kaiju, pretraining/Squinch, compelling writing, open-source shift, memory, Lorebook, AMD/DigitalOcean inference, Slonk, pipeling-sft, Agent SDK). Treat them as vendor-published, not an independent audit. Traffic charts from Similarweb are third-party estimates.\n\n### Research this topic further\n\n- Click a tool below (ChatGPT, Perplexity, Claude, or Gemini).\n- We copy the\n**full article + companion guide prompt** to your clipboard. - A new chat tab opens — if the message box is empty or only shows a short note, press\n`Ctrl+V`(Windows/Linux) or` Cmd+V`(Mac) to paste. - Send the message. The model already has the article text — it does not need to open this website.\n\n### Related posts\n\n[18 · 07 · 2026Give Your Agents Power to Purchase: Agentcard.sh Go-To GuideAgents fill carts then stall at checkout. Agentcard.sh issues capped one-time Visas so they can finish buying — MCP, Companies wizard, and when not to use it.](/writings/ai-tools/give-your-agents-power-to-purchase-agentcard-sh-go-to-guide)\n\n[22 · 07 · 2026Inside Qwen 3.8-Max-Preview: Reverse Engineering an AI Assistant by Interviewing ItselfI interviewed Qwen 3.8-Max-Preview on memory, tools, context, and hallucinations, then checked published research. Plain English, diagrams, transcript.](/writings/models/inside-qwen-3-8-max-preview-reverse-engineering-an-ai-assistant-by-interviewing-itself)\n\n[20 · 07 · 2026Inside Kimi K2.6: Reverse Engineering an AI Assistant by Interviewing ItselfI interviewed Kimi on memory, tools, and safety, then checked published research on K2.6. Plain English, diagrams, sources.](/writings/models/inside-kimi-k2-6-reverse-engineering-an-ai-assistant-by-interviewing-itself)\n\n### Comments\n\nShare a thought on this post — keep it useful and kind. Comments are moderated before they appear.\n\nLoading comments…", "url": "https://wpnews.pro/news/character-ai-the-technical-story-of-what-keeps-users-hooked", "canonical_source": "https://manish.sh/writings/ai-tools/inside-character-ai-the-technical-story-of-what-keeps-users-hooked", "published_at": "2026-07-22 20:54:51+00:00", "updated_at": "2026-07-22 21:22:59.066231+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-products", "ai-research", "ai-infrastructure"], "entities": ["Character.ai", "Noam Shazeer", "Daniel De Freitas", "Google", "Kaiju", "Transformer", "Meena", "LaMDA"], "alternates": {"html": "https://wpnews.pro/news/character-ai-the-technical-story-of-what-keeps-users-hooked", "markdown": "https://wpnews.pro/news/character-ai-the-technical-story-of-what-keeps-users-hooked.md", "text": "https://wpnews.pro/news/character-ai-the-technical-story-of-what-keeps-users-hooked.txt", "jsonld": "https://wpnews.pro/news/character-ai-the-technical-story-of-what-keeps-users-hooked.jsonld"}}