{"slug": "making-llm-ratings-auditable-quotes-limitations-and-can-t-judge-instead-of-zero", "title": "Making LLM ratings auditable: quotes, limitations, and \"can't judge\" instead of zero", "summary": "A developer building a Chinese-language WeChat Official Account directory has published a set of rules for making LLM-generated ratings auditable, including verifying every model quote against stored article text and displaying \"can't judge\" rather than a zero when evidence fails verification. Accounts are scored on up to 20 recent articles using at most the first 6,000 characters of each, and any account with fewer than 20 usable articles receives no overall score. Each profile discloses the article count, model name, rubric version and analysis date, and the site labels its bottom-of-page recommendations as rule-based category matches rather than quality endorsements.", "body_md": "I'm building a side project, [WeChat Source](https://wechat.aim888888.xyz/wechat). It's a directory that helps Chinese readers find WeChat Official Accounts (the newsletter-style publications inside WeChat). The site is in Chinese, but the design problem behind it isn't: **how do you show LLM-generated ratings without them turning into confident-sounding noise?**\n\nHere are the rules I ended up with. All of them are visible on every account page.\n\nEach account is rated on up to its **20 most recent articles with readable text**, using at most the **first 6,000 characters** of each. If an account has fewer than 20 usable articles, it gets no overall score. A small sample produces a confident-looking number with little behind it, so I'd rather show nothing.\n\nThere are five dimensions: depth of argument, source transparency, information gain, clarity, and caution in stating opinions. Each gets a 1–5 score plus:\n\nThe page also states the scale: 3 is acceptable, 4 is good, and 5 requires strong evidence.\n\nLLMs sometimes \"quote\" text that doesn't exist. Quotes are checked against the stored article text, and any that don't match are removed. If that leaves a dimension without solid evidence, it's displayed as **\"can't judge\"**, with a note that the quotes failed verification and the item needs re-analysis or a human review.\n\nImportantly, **\"can't judge\" is not counted as zero.** Treating \"unknown\" as \"bad\" would quietly punish accounts for the model's mistakes.\n\nEach profile shows how many articles it's based on, the model name, a rubric version string, and the analysis date. There's also a line saying the score is the platform's AI analysis, not an official evaluation. If the rubric changes, older profiles are at least identifiable.\n\nAt the bottom of each account page there's a \"keep exploring\" list. It's a rule-based match on shared categories and tags, and the page says exactly that, including the similarity percentage and the shared category. It would be easy to let people assume these are quality recommendations. They aren't, so the label says so.\n\nNext.js front end (served under a `/wechat` base path), a Python back end, and Cloudflare in front.\n\nIf you've shipped LLM-generated ratings or summaries, I'm curious how you handle evidence and \"unknown\" states. And if you read Chinese, the site is [here](https://wechat.aim888888.xyz/wechat). Feedback is very welcome.", "url": "https://wpnews.pro/news/making-llm-ratings-auditable-quotes-limitations-and-can-t-judge-instead-of-zero", "canonical_source": "https://dev.to/stiamora/making-llm-ratings-auditable-quotes-limitations-and-cant-judge-instead-of-zero-27in", "published_at": "2026-10-08 12:46:13+00:00", "updated_at": "2026-10-08 12:49:53.956323+00:00", "lang": "en", "topics": ["large-language-models", "ai-products", "ai-tools", "ai-ethics"], "entities": ["WeChat", "Next.js", "Python", "Cloudflare"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/making-llm-ratings-auditable-quotes-limitations-and-can-t-judge-instead-of-zero", "markdown": "https://wpnews.pro/news/making-llm-ratings-auditable-quotes-limitations-and-can-t-judge-instead-of-zero.md", "text": "https://wpnews.pro/news/making-llm-ratings-auditable-quotes-limitations-and-can-t-judge-instead-of-zero.txt", "jsonld": "https://wpnews.pro/news/making-llm-ratings-auditable-quotes-limitations-and-can-t-judge-instead-of-zero.jsonld"}}