# A curation model may decide what an additive is, and may never write the sentence you read

> Source: <https://dev.to/daniel_pertu/a-curation-model-may-decide-what-an-additive-is-and-may-never-write-the-sentence-you-read-8dd>
> Published: 2026-10-05 08:40:40+00:00

[Munchable](https://munchable.app) reads a food label and says whether the product fits the gut conditions you manage. Behind that answer are curated maps: what an ingredient is, which family it belongs to, which rule set cares about it. Those maps grow two ways, by hand and through AI-assisted curation jobs that we review and merge.

Which raises a question that only has one safe answer. When the app says

Allura red AC (E129), an artificial colour

who wrote that sentence?

Not the data. Nothing a curation job writes is ever rendered to a user, anywhere in the product. This post is about how that rule is enforced in three different files, and what it buys.

The scoring layer is a table of overlay rows that extend the engine's maps. Its module header is blunt about what a row may contain:

```
// The taxonomy overlay decides what the engine can NAME. This one decides what
// it can SAY: which canonical ids land in the FODMAP / lactose / GERD / IBD /
// gastroparesis maps. Values are the engine's own enums and nothing else.
// Reason copy is derived from the ingredient id inside the engine, so no
// string written here is ever rendered to a user.
```

The additive map behind our shopping preferences says the same thing about its own overlay: it may only add ids the hand-written file does not carry, and those rows hold enums only.

And the function that turns an id into a display name closes the loop:

```
/**
 * The plain name of an id for a finding: the hand map's `name` where there is
 * one, otherwise derived from the id. An overlay row has no name, by design,
 * so a curation model can never author a word the user reads.
 */
function nameOf(id: string, entry: AdditiveEntry): string {
  if (entry.name) return entry.name;
  const e = eNumber(id);
  if (e) return e;
  const bare = id.includes(':') ? id.slice(id.indexOf(':') + 1) : id;
  const spaced = bare.replace(/-/g, ' ');
  return spaced.charAt(0).toUpperCase() + spaced.slice(1);
}
```

An overlay row has no `name` field to fill in. Not "we do not populate it": the shape does not have one. That is the difference between a convention and a constraint, and only one of them survives a busy afternoon.

Three cases, and each one is an editorial decision expressed as code:

```
export function healthMessageFor(id: string, entry: AdditiveEntry): string {
  const name = nameOf(id, entry);
  const e = eNumber(id);
  const label = ADDITIVE_CLASS_LABEL[entry.class];
  if (e) return name === e ? `${e}, ${label}` : `${name} (${e}), ${label}`;
  if (entry.class === 'added-sugar') return `${name}, an added sugar`;
  return name;
}
```

Which produces:

```
Allura red AC (E129), an artificial colour
Hydrogenated palm fat
Maltodextrin, an added sugar
```

An E-number always gets its class named, because the number alone tells a reader nothing. A plain ingredient is left to speak for itself, since "Hydrogenated palm fat, a hydrogenated fat" is a sentence that says one thing twice. Added sugars are the exception, because "Dextrose" and "Maltodextrin" are exactly the words a label uses when it does not want to say sugar.

Those three lines of reasoning are the kind of thing a copy deck is normally for. Putting them in a function means they apply identically to every one of the hundreds of substances the map covers, including the ones a curation job added last week that no human has ever read a sentence for.

Because the sentence is derived, the sentence is also the right deduplication key. A pack can list a substance under its name and its number, and an ingredient can expand to several taxonomy ids that all render the same way. The filter therefore dedupes on the finished line:

``` js
const message = healthMessageFor(id, map[id]!);
// A pack that lists a substance twice, or lists it and its E-number, is
// one finding. Deduped on the finished sentence rather than on the id,
// which is what collapses `en:e322` and `en:e322i` without a substance
// key to keep in step with the map.
```

If the copy had come from the data, two rows for one substance would have produced two slightly different sentences, and no dedupe could have caught them without a third concept to tie them together. Deriving the text gave us an identity for free.

Because the failure modes are not symmetrical.

A model that writes an enum into a reviewable row is making a classification claim I can check against a reference, in bulk, with the engine's own validator. The validator is the same code on the server and on the device, so both drop exactly the same rows for the same reasons, and a malformed row is simply not served.

A model that writes prose into a health product is authoring medical-adjacent copy, at scale, which nobody will read until a user does. There is no validator for a tone, and the first time you find out that a row says something alarming about a food colouring is in a support ticket.

The same principle covers how certain we claim to be. The strength of the evidence behind a rule is a field on the rule set rather than a sentence somebody wrote, which I covered in [our uncertainty is a field on the rule set, not a sentence in the copy deck](https://dev.to/daniel_pertu/our-uncertainty-is-a-field-on-the-rule-set-not-a-sentence-in-the-copy-deck-5h2k). And because all of this copy is code, it is testable: our house style, including the ban on em dashes, is [enforced by unit tests](https://dev.to/daniel_pertu/our-brand-voice-rules-are-unit-tests-including-the-one-that-bans-the-em-dash-4bd2) that a data row could never be subject to.

What the preference layer is allowed to do with its findings, as opposed to who writes them, is a separate design: [three layers over one product, and only one of them is allowed to say no](https://dev.to/daniel_pertu/three-layers-over-one-product-and-only-one-of-them-is-allowed-to-say-no-7ag).

Our public answer pages run the same engine the app runs, so the prose on them is generated by the code above rather than written by anyone:

Then open [app.munchable.app](https://app.munchable.app), switch on the shopping preferences during onboarding, and check a product with a long additive list. Every line you read came out of a template and an enum. There is no table anywhere in our stack with a sentence in it.
