# Making LLM ratings auditable: quotes, limitations, and "can't judge" instead of zero

> Source: <https://dev.to/stiamora/making-llm-ratings-auditable-quotes-limitations-and-cant-judge-instead-of-zero-27in>
> Published: 2026-10-08 12:46:13+00:00

I'm building a side project, [WeChat Source](https://wechat.aim888888.xyz/wechat). It's a directory that helps Chinese readers find WeChat Official Accounts (the newsletter-style publications inside WeChat). The site is in Chinese, but the design problem behind it isn't: **how do you show LLM-generated ratings without them turning into confident-sounding noise?**

Here are the rules I ended up with. All of them are visible on every account page.

Each account is rated on up to its **20 most recent articles with readable text**, using at most the **first 6,000 characters** of each. If an account has fewer than 20 usable articles, it gets no overall score. A small sample produces a confident-looking number with little behind it, so I'd rather show nothing.

There are five dimensions: depth of argument, source transparency, information gain, clarity, and caution in stating opinions. Each gets a 1–5 score plus:

The page also states the scale: 3 is acceptable, 4 is good, and 5 requires strong evidence.

LLMs sometimes "quote" text that doesn't exist. Quotes are checked against the stored article text, and any that don't match are removed. If that leaves a dimension without solid evidence, it's displayed as **"can't judge"**, with a note that the quotes failed verification and the item needs re-analysis or a human review.

Importantly, **"can't judge" is not counted as zero.** Treating "unknown" as "bad" would quietly punish accounts for the model's mistakes.

Each profile shows how many articles it's based on, the model name, a rubric version string, and the analysis date. There's also a line saying the score is the platform's AI analysis, not an official evaluation. If the rubric changes, older profiles are at least identifiable.

At the bottom of each account page there's a "keep exploring" list. It's a rule-based match on shared categories and tags, and the page says exactly that, including the similarity percentage and the shared category. It would be easy to let people assume these are quality recommendations. They aren't, so the label says so.

Next.js front end (served under a `/wechat` base path), a Python back end, and Cloudflare in front.

If you've shipped LLM-generated ratings or summaries, I'm curious how you handle evidence and "unknown" states. And if you read Chinese, the site is [here](https://wechat.aim888888.xyz/wechat). Feedback is very welcome.
