{"slug": "show-hn-revalvo-local-first-llm-eval-bench-multi-model-byok", "title": "Show HN: Revalvo – local-first LLM eval bench (multi-model, BYOK)", "summary": "Revalvo, a new local-first LLM evaluation workbench, launched on Product Hunt today, enabling users to run prompts against multiple models in parallel with per-response latency, tokens, and cost tracking, while keeping API keys in the browser and data in IndexedDB. The tool offers 40 built-in evaluators—25 rule-based and 14 LLM-judge-based—and supports versioning, batch eval, and GitHub sync for prompt recipes. Founder seeks feedback on evaluator coverage and rule-vs-judge labeling.", "body_md": "Hi HN — I built Revalvo: a local-first workbench for prompt engineering and LLM evaluation.\n\nThe problem I kept hitting: chat playgrounds are fast but leave no receipt, and hosted eval platforms are rigorous but slow and server-side. Revalvo sits in the middle — sub-minute setup, your keys stay in your browser, no Revalvo account.\n\nWhat it does today:\n\n• Run the same prompt against multiple models in parallel (OpenRouter, Groq, direct APIs, Ollama/LM Studio locally) • Per-response latency, tokens, and cost • Version prompts like code (saves, diffs, rollback) • Batch eval on datasets with 40 built-in evaluators • GitHub sync for prompt recipes (YAML push/pull)\n\nOn evaluators: ~25 are deterministic rules (exact match, regex, JSON schema, length, PII patterns, etc. — no extra API spend). ~14 need a model (LLM judge, rubric, faithfulness-style scorers) plus embedding similarity. Each is labeled Rule-based vs LLM judge in the UI — I stack rules first, judges when rules aren’t enough.\n\nPrivacy / cost: BYOK only. We don’t markup API spend. App data lives in IndexedDB in your browser; keys aren’t sent to our servers (CORS-blocked providers go through a same-origin dev proxy when you use the hosted app).\n\nWe launched on Product Hunt today; early feedback is pushing on exactly the evaluator split above, which matches how I think about the product too.\n\nTry it: paste an OpenRouter/Groq key or run fully offline with Ollama. Playground is the fastest path to “does this actually work for me?”\n\nI’d love HN feedback on: 1. Evaluator coverage — what’s missing for your workflows? 2. Whether rule vs judge labeling is clear enough before you attach evals to a dataset 3. GitHub sync for prompt YAML — useful or noise?\n\nHappy to answer technical questions (architecture, storage, provider adapters).\n\nComments URL: [https://news.ycombinator.com/item?id=49477347](https://news.ycombinator.com/item?id=49477347)\n\nPoints: 1\n\n# Comments: 0", "url": "https://wpnews.pro/news/show-hn-revalvo-local-first-llm-eval-bench-multi-model-byok", "canonical_source": "https://revalvo.com", "published_at": "2026-08-28 12:07:15+00:00", "updated_at": "2026-08-28 12:18:30.131075+00:00", "lang": "en", "topics": ["ai-tools", "large-language-models", "ai-products", "developer-tools"], "entities": ["Revalvo", "OpenRouter", "Groq", "Ollama", "LM Studio", "Product Hunt"], "alternates": {"html": "https://wpnews.pro/news/show-hn-revalvo-local-first-llm-eval-bench-multi-model-byok", "markdown": "https://wpnews.pro/news/show-hn-revalvo-local-first-llm-eval-bench-multi-model-byok.md", "text": "https://wpnews.pro/news/show-hn-revalvo-local-first-llm-eval-bench-multi-model-byok.txt", "jsonld": "https://wpnews.pro/news/show-hn-revalvo-local-first-llm-eval-bench-multi-model-byok.jsonld"}}