Making LLM ratings auditable: quotes, limitations, and "can't judge" instead of zero A developer building a Chinese-language WeChat Official Account directory has published a set of rules for making LLM-generated ratings auditable, including verifying every model quote against stored article text and displaying "can't judge" rather than a zero when evidence fails verification. Accounts are scored on up to 20 recent articles using at most the first 6,000 characters of each, and any account with fewer than 20 usable articles receives no overall score. Each profile discloses the article count, model name, rubric version and analysis date, and the site labels its bottom-of-page recommendations as rule-based category matches rather than quality endorsements. I'm building a side project, WeChat Source https://wechat.aim888888.xyz/wechat . It's a directory that helps Chinese readers find WeChat Official Accounts the newsletter-style publications inside WeChat . The site is in Chinese, but the design problem behind it isn't: how do you show LLM-generated ratings without them turning into confident-sounding noise? Here are the rules I ended up with. All of them are visible on every account page. Each account is rated on up to its 20 most recent articles with readable text , using at most the first 6,000 characters of each. If an account has fewer than 20 usable articles, it gets no overall score. A small sample produces a confident-looking number with little behind it, so I'd rather show nothing. There are five dimensions: depth of argument, source transparency, information gain, clarity, and caution in stating opinions. Each gets a 1–5 score plus: The page also states the scale: 3 is acceptable, 4 is good, and 5 requires strong evidence. LLMs sometimes "quote" text that doesn't exist. Quotes are checked against the stored article text, and any that don't match are removed. If that leaves a dimension without solid evidence, it's displayed as "can't judge" , with a note that the quotes failed verification and the item needs re-analysis or a human review. Importantly, "can't judge" is not counted as zero. Treating "unknown" as "bad" would quietly punish accounts for the model's mistakes. Each profile shows how many articles it's based on, the model name, a rubric version string, and the analysis date. There's also a line saying the score is the platform's AI analysis, not an official evaluation. If the rubric changes, older profiles are at least identifiable. At the bottom of each account page there's a "keep exploring" list. It's a rule-based match on shared categories and tags, and the page says exactly that, including the similarity percentage and the shared category. It would be easy to let people assume these are quality recommendations. They aren't, so the label says so. Next.js front end served under a /wechat base path , a Python back end, and Cloudflare in front. If you've shipped LLM-generated ratings or summaries, I'm curious how you handle evidence and "unknown" states. And if you read Chinese, the site is here https://wechat.aim888888.xyz/wechat . Feedback is very welcome.