Every chatbot you've ever talked to has a personality — helpful but careful, warm but not fawning, opinionated about some things and diplomatically vague about others. Ask where that personality comes from and you'll get hand-waving about "emergent behavior" and "alignment." Here's the thing nobody tells you: the personality is a file, and you can read it.
Anthropic's actual 2022 preference-training data (169K examples) and Nvidia's 2024 equivalent are public on Hugging Face. I read them — and measured how the rules changed.
- Refusals, politeness, and tone are installed by specific, readable training examples— I'll show you the exact rows - In 2022 data, preferred answers often refused to name products; by 2024, naming specific products is what wins — measured: {{PCT2022}}% → {{PCT2024}}% - Preference training also explains why AI recommendations are winner-take-all: base models spread belief across brands; post-training crowns one king
An AI's values aren't emergent. They're a curated dataset with version numbers — and the version history changes who gets recommended.
The files #
After training a GPT on my own life, I went looking for the step my laptop pipeline skipped: the one that turns a text-predictor into a personality. That step is preference training (RLHF), and it turns out the raw material is just… sitting there:
Anthropic/hh-rlhf— 169,000 real preference examples from 2022, published alongside the papers that shaped the original Claude. The #3 most-liked dataset on all of Hugging Face.nvidia/HelpSteer2— the 2024 generation: responses rated 0–5 on helpfulness, correctness, coherence.
These aren't summaries or paraphrases. They're the actual rows: a prompt, two candidate responses, and which one humans preferred. You can open them in a browser like a spreadsheet. I read a few hundred so you don't have to — though honestly, you should.
How preference training works (60-second version) #
Building a chatbot has three stages. Pretraining teaches it language (predict the next word, trillions of times). Supervised fine-tuning teaches it the job (imitate worked examples of good responses). Then preference training teaches it judgment — and the mechanism is almost comically simple:
Show the model a prompt with two responses. Tell it which one humans preferred. Nudge its weights toward the winner. Repeat 169,000 times.
There is no "correct answer" column anywhere in the file. Just this one over that one, thousands of times, until the statistical shadow of all those judgments becomes what we experience as the assistant's character. Every pair below is verbatim from the data.
What gets trained in #
Politeness calibration
The shorter, less fawning reply won. Somewhere in this row and its thousands of siblings is the reason modern chatbots stopped ending every message with "Is there anything else?" — annotators quietly preferred restraint, and restraint got compiled into the weights. When a chatbot's tone feels "designed," this is the design surface.
Where refusals come from
The most revealing pairs are the ones where the rejected answer is competent. Watch:
The rejected answer isn't badly written — it's fluent, specific, genuinely helpful. That's the point. A refusal isn't the model failing to know something; it's the model having been shown, in pairs like this one, that willingness itself is the thing to suppress.
Content note: the next two examples are from the dataset's red-teaming half and include a request about assault and a racist premise, quoted as they appear in the public training data. They're included because sanitizing them would hide what this data actually does.
Note what the 2022-era model actually did one turn earlier: it happily produced a street address (almost certainly hallucinated). The preference pair doesn't erase that capability — it teaches the model to stop cooperating as the request escalates. "I want him to know that I am coming" is the pivot the annotator is really grading.
The rejected answer accepts the racist premise and wonders about it. The preferred one pushes back — clumsily, by today's standards, which is itself informative: this is 2022-era alignment, visibly cruder than what you'd get now. The data has a skill ceiling, and it shows. These datasets are geological strata: you can carbon-date an AI behavior by finding the layer that taught it.
Helpfulness has a shape too
The other half of the dataset trains the opposite lesson — anti-evasiveness:
Enthusiasm-then-a-question lost to actual answers. Train on thousands of these and you get a model with a default to substance. The assistant personality — helpful but not reckless, warm but not groveling, substantive but not lecturing — is the negotiated equilibrium of these two stacks of preferences pulling against each other. Optimize either alone and you get a model that refuses to discuss knives, or one that explains embezzlement.
The recommendation shift: 2022 vs 2024, measured #
Here's where this stops being trivia and starts being commercially interesting. Reading recommendation-style prompts in the 2022 data, I kept noticing preferred answers like this:
(bass guitars)