{"slug": "building-a-24-language-journal-in-plain-php-ai-translation-hash-based-staleness", "title": "Building a 24-Language Journal in Plain PHP: AI Translation, Hash-Based Staleness, and a Points-Gated Community", "summary": "A developer built a 24-language Journal feature for the GDPR-first metasearch engine findnix.eu using plain PHP 8.4 and MariaDB, translating German source content into 23 EU languages with gpt-4o-mini chunked at heading boundaries and sanitized through a DOMDocument whitelist before storage. Staleness is tracked with an MD5 hash of the source title, excerpt and body rather than updated_at, and a GROUP BY title duplicate check caught eight mistranslated Slovak titles that an LLM self-grading pass missed. \"For batch-generated content, cheap invariants beat asking an LLM to grade itself,\" the developer wrote.", "body_md": "findnix.eu is a GDPR-first metasearch engine with its own index. Search was never the whole story, so we added a **Journal**: articles, a weekly digest, \"finds\" (sites worth a look), user questions and answers, and a profile page for every domain in the index. It's plain PHP 8.4 and MariaDB, no framework, in the same codebase as everything else. These are the decisions that mattered.\n\nArticles, digests, finds and questions share one table with an `ENUM` type column. Translations live in their own table keyed by `(post_id, lang)`, so the original post never changes shape and a missing language is just a missing row. German is the source language; the other 23 EU languages are translations. Each language version gets its own URL (`/journal/fr/some-slug`) with `hreflang` links, and the sitemap lists every version with its translated title and description, so the translated pages end up in our own search index too.\n\nThe renderer is about 100 lines. The one rule that keeps it safe: escape everything first, then generate only the tags we allow.\n\n``` php\n$t = htmlspecialchars($t, ENT_QUOTES | ENT_SUBSTITUTE, 'UTF-8');\n// only now turn [text](url) into <a>, and only for http(s), mailto or local paths\n```\n\nRaw HTML in a post simply shows up as text. Links from user content get `rel=\"nofollow ugc\"`.\n\nPosts are rendered to HTML once, and that HTML is what we translate: split at `<h2>`–`<h4>` boundaries into chunks of a few thousand characters, one `gpt-4o-mini` call per chunk, with the instruction to keep every tag and attribute and translate only visible text. Title and teaser go in a separate call that returns JSON.\n\nThe important part is what happens afterwards. An article is attacker-controlled text as far as the model is concerned: a user post could contain instructions, and the model might obey them and emit markup. So the translated HTML goes through a `DOMDocument` whitelist (a short list of tags and attributes, URL schemes checked) before it is stored. If the model produces a `<script>`, it never reaches a page.\n\n`updated_at` question\nFirst version: compare the translation's timestamp with the post's `updated_at`. It broke immediately, because `updated_at` moves when you pin a post, change its status, or bump the view counter. Every pin would have re-translated 23 languages.\n\nNow each translation stores a hash of the source text it was made from:\n\n```\nMD5(CONCAT(title, CHAR(10), COALESCE(excerpt, ''), CHAR(10), body_md))\n```\n\nA translation is current if its stored hash equals the post's current hash. Pinning, status changes and view counts don't matter any more. (View counters are updated with `updated_at = updated_at` anyway, so they don't touch the timestamp.)\n\nTranslating all languages for a post takes a minute or more, which is too long for a save request. We use two triggers:\n\nBoth call the same function and write with an upsert, so running them twice is harmless.\n\nOur first batch of translations looked fine until we noticed that the Slovak and Slovenian titles were identical, for 8 of 9 posts. The prompt said \"Slovenčina (sk)\". The model read the name, not the code, and answered in Slovenian.\n\nThe fix was to use English language names plus an explicit contrast in the prompt: *\"Slovak (the language of Slovakia; NOT Slovenian and NOT Czech)\"*, and the mirror image for Slovenian, Czech, Croatian, Danish and Swedish. Re-running the nine posts produced correct Slovak.\n\nHow we found it matters more than the fix. Asking the model to identify the language of each translation caught only 2 of the 8 wrong ones. A plain `GROUP BY title HAVING COUNT(*) > 1` across languages caught all of them. For batch-generated content, cheap invariants beat asking an LLM to grade itself.\n\nEvery user action (a comment, an answer, a question, a submitted article, a domain review) goes through the same two helpers: one that grants points with a daily cap, and one that revokes them. The rules we settled on:\n\nThe amounts live in a settings table, so we can tune them in the admin without a deploy.\n\nEvery domain in the index gets a page with its page count, sitemap count, first-seen date and a few sample pages. Counting rows per domain in a 37-million-row table on every request is a non-starter, so the count is capped inside a subquery and the result is cached:\n\n```\nSELECT COUNT(*) FROM (\n  SELECT 1 FROM fnx_sitemap_urls\n  WHERE domain = ? AND status = 'aktiv' AND adult = 0\n  LIMIT 100000\n) t\n```\n\nLarge domains show \"100,000+\", which is honest and costs one index range scan. A batch job refreshes the numbers slowly, with a pause between domains, rather than all at once.\n\nOur spam blacklist matched keywords as substrings. The keyword `anal` blocked `analyst.nl`, `analytics-agentur.ch` and `psykoanalyse.no`; `sex` would have blocked `sussex.ac.uk`. We added a `=word` syntax that matches only at word boundaries (no letter or digit directly before or after) and left long, unambiguous terms as substrings so that concatenated spam domains still get caught. Of the 27 indexed domains the old filter had hidden, all 27 were false positives.\n\nThe Journal is live at [findnix.eu/journal](https://findnix.eu/journal), and the questions section is at [findnix.eu/journal/fragen](https://findnix.eu/journal/fragen).", "url": "https://wpnews.pro/news/building-a-24-language-journal-in-plain-php-ai-translation-hash-based-staleness", "canonical_source": "https://dev.to/findnix/building-a-24-language-journal-in-plain-php-ai-translation-hash-based-staleness-and-a-3a82", "published_at": "2026-10-05 19:08:18+00:00", "updated_at": "2026-10-05 19:18:20.403021+00:00", "lang": "en", "topics": ["ai-tools", "large-language-models", "generative-ai", "developer-tools"], "entities": ["findnix.eu", "PHP", "MariaDB", "gpt-4o-mini", "OpenAI"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/building-a-24-language-journal-in-plain-php-ai-translation-hash-based-staleness", "markdown": "https://wpnews.pro/news/building-a-24-language-journal-in-plain-php-ai-translation-hash-based-staleness.md", "text": "https://wpnews.pro/news/building-a-24-language-journal-in-plain-php-ai-translation-hash-based-staleness.txt", "jsonld": "https://wpnews.pro/news/building-a-24-language-journal-in-plain-php-ai-translation-hash-based-staleness.jsonld"}}