{"slug": "a-curation-model-may-decide-what-an-additive-is-and-may-never-write-the-sentence", "title": "A curation model may decide what an additive is, and may never write the sentence you read", "summary": "A developer behind the food-label app Munchable described an architecture that prevents AI-assisted curation jobs from ever writing text users read. Overlay rows in the scoring and additive maps carry only engine enums and have no name field by design, so display sentences are derived inside the engine by a nameOf function and a healthMessageFor function that also serves as the deduplication key for findings.", "body_md": "[Munchable](https://munchable.app) reads a food label and says whether the product fits the gut conditions you manage. Behind that answer are curated maps: what an ingredient is, which family it belongs to, which rule set cares about it. Those maps grow two ways, by hand and through AI-assisted curation jobs that we review and merge.\n\nWhich raises a question that only has one safe answer. When the app says\n\nAllura red AC (E129), an artificial colour\n\nwho wrote that sentence?\n\nNot the data. Nothing a curation job writes is ever rendered to a user, anywhere in the product. This post is about how that rule is enforced in three different files, and what it buys.\n\nThe scoring layer is a table of overlay rows that extend the engine's maps. Its module header is blunt about what a row may contain:\n\n```\n// The taxonomy overlay decides what the engine can NAME. This one decides what\n// it can SAY: which canonical ids land in the FODMAP / lactose / GERD / IBD /\n// gastroparesis maps. Values are the engine's own enums and nothing else.\n// Reason copy is derived from the ingredient id inside the engine, so no\n// string written here is ever rendered to a user.\n```\n\nThe additive map behind our shopping preferences says the same thing about its own overlay: it may only add ids the hand-written file does not carry, and those rows hold enums only.\n\nAnd the function that turns an id into a display name closes the loop:\n\n```\n/**\n * The plain name of an id for a finding: the hand map's `name` where there is\n * one, otherwise derived from the id. An overlay row has no name, by design,\n * so a curation model can never author a word the user reads.\n */\nfunction nameOf(id: string, entry: AdditiveEntry): string {\n  if (entry.name) return entry.name;\n  const e = eNumber(id);\n  if (e) return e;\n  const bare = id.includes(':') ? id.slice(id.indexOf(':') + 1) : id;\n  const spaced = bare.replace(/-/g, ' ');\n  return spaced.charAt(0).toUpperCase() + spaced.slice(1);\n}\n```\n\nAn overlay row has no `name` field to fill in. Not \"we do not populate it\": the shape does not have one. That is the difference between a convention and a constraint, and only one of them survives a busy afternoon.\n\nThree cases, and each one is an editorial decision expressed as code:\n\n```\nexport function healthMessageFor(id: string, entry: AdditiveEntry): string {\n  const name = nameOf(id, entry);\n  const e = eNumber(id);\n  const label = ADDITIVE_CLASS_LABEL[entry.class];\n  if (e) return name === e ? `${e}, ${label}` : `${name} (${e}), ${label}`;\n  if (entry.class === 'added-sugar') return `${name}, an added sugar`;\n  return name;\n}\n```\n\nWhich produces:\n\n```\nAllura red AC (E129), an artificial colour\nHydrogenated palm fat\nMaltodextrin, an added sugar\n```\n\nAn E-number always gets its class named, because the number alone tells a reader nothing. A plain ingredient is left to speak for itself, since \"Hydrogenated palm fat, a hydrogenated fat\" is a sentence that says one thing twice. Added sugars are the exception, because \"Dextrose\" and \"Maltodextrin\" are exactly the words a label uses when it does not want to say sugar.\n\nThose three lines of reasoning are the kind of thing a copy deck is normally for. Putting them in a function means they apply identically to every one of the hundreds of substances the map covers, including the ones a curation job added last week that no human has ever read a sentence for.\n\nBecause the sentence is derived, the sentence is also the right deduplication key. A pack can list a substance under its name and its number, and an ingredient can expand to several taxonomy ids that all render the same way. The filter therefore dedupes on the finished line:\n\n``` js\nconst message = healthMessageFor(id, map[id]!);\n// A pack that lists a substance twice, or lists it and its E-number, is\n// one finding. Deduped on the finished sentence rather than on the id,\n// which is what collapses `en:e322` and `en:e322i` without a substance\n// key to keep in step with the map.\n```\n\nIf the copy had come from the data, two rows for one substance would have produced two slightly different sentences, and no dedupe could have caught them without a third concept to tie them together. Deriving the text gave us an identity for free.\n\nBecause the failure modes are not symmetrical.\n\nA model that writes an enum into a reviewable row is making a classification claim I can check against a reference, in bulk, with the engine's own validator. The validator is the same code on the server and on the device, so both drop exactly the same rows for the same reasons, and a malformed row is simply not served.\n\nA model that writes prose into a health product is authoring medical-adjacent copy, at scale, which nobody will read until a user does. There is no validator for a tone, and the first time you find out that a row says something alarming about a food colouring is in a support ticket.\n\nThe same principle covers how certain we claim to be. The strength of the evidence behind a rule is a field on the rule set rather than a sentence somebody wrote, which I covered in [our uncertainty is a field on the rule set, not a sentence in the copy deck](https://dev.to/daniel_pertu/our-uncertainty-is-a-field-on-the-rule-set-not-a-sentence-in-the-copy-deck-5h2k). And because all of this copy is code, it is testable: our house style, including the ban on em dashes, is [enforced by unit tests](https://dev.to/daniel_pertu/our-brand-voice-rules-are-unit-tests-including-the-one-that-bans-the-em-dash-4bd2) that a data row could never be subject to.\n\nWhat the preference layer is allowed to do with its findings, as opposed to who writes them, is a separate design: [three layers over one product, and only one of them is allowed to say no](https://dev.to/daniel_pertu/three-layers-over-one-product-and-only-one-of-them-is-allowed-to-say-no-7ag).\n\nOur public answer pages run the same engine the app runs, so the prose on them is generated by the code above rather than written by anyone:\n\nThen open [app.munchable.app](https://app.munchable.app), switch on the shopping preferences during onboarding, and check a product with a long additive list. Every line you read came out of a template and an enum. There is no table anywhere in our stack with a sentence in it.", "url": "https://wpnews.pro/news/a-curation-model-may-decide-what-an-additive-is-and-may-never-write-the-sentence", "canonical_source": "https://dev.to/daniel_pertu/a-curation-model-may-decide-what-an-additive-is-and-may-never-write-the-sentence-you-read-8dd", "published_at": "2026-10-05 08:40:40+00:00", "updated_at": "2026-10-05 08:49:40.415542+00:00", "lang": "en", "topics": ["ai-safety", "artificial-intelligence", "ai-products"], "entities": ["Munchable"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/a-curation-model-may-decide-what-an-additive-is-and-may-never-write-the-sentence", "markdown": "https://wpnews.pro/news/a-curation-model-may-decide-what-an-additive-is-and-may-never-write-the-sentence.md", "text": "https://wpnews.pro/news/a-curation-model-may-decide-what-an-additive-is-and-may-never-write-the-sentence.txt", "jsonld": "https://wpnews.pro/news/a-curation-model-may-decide-what-an-additive-is-and-may-never-write-the-sentence.jsonld"}}