{"slug": "show-hn-oqoqo-build-evals-and-custom-benchmarks-for-real-world-tasks", "title": "Show HN: Oqoqo – build evals and custom benchmarks for real-world tasks", "summary": "Oqoqo, a new tool launched on Hacker News, enables developers to build realistic evals and custom benchmarks for real-world tasks, measuring agent-friendly product surfaces against Codex, Claude Code, OpenClaw, Hermes, Pi, Opencode, Cursor, and GitHub Copilot. It supports regression testing for MCP, CLI, skills, SDK, and other agent-facing interfaces, and allows creating and sharing custom benchmarks for agent discovery and usage.", "body_md": "Most benchmarks today exist in curated environments and do not translate well to the real world.\n\nWe built Oqoqo to bridge this gap. Oqoqo makes it super simple to build realistic evals and custom benchmarks for tasks users actually care about.\n\nOqoqo can:\n\n- Reliably measure how agent friendly your product surfaces are against Codex, Claude Code, OpenClaw, Hermes, Pi, Opencode, Cursor, GitHub Copilot - Regression test MCP, CLI, skills, SDK, and any agent facing interface (we are continuously using Oqoqo to dogfood and improve our own MCP/CLI) - Create and share custom benchmarks for how agents discover and use your product - Compare models and harnesses for domain specific tasks - See whether new versions improve agent experience\n\nWould love to hear feedback and thoughts on what kind of evals you are running today, if you are benchmarking agent facing interfaces, and whether you have published a custom benchmark.\n\nComments URL: [https://news.ycombinator.com/item?id=49249987](https://news.ycombinator.com/item?id=49249987)\n\nPoints: 1\n\n# Comments: 0", "url": "https://wpnews.pro/news/show-hn-oqoqo-build-evals-and-custom-benchmarks-for-real-world-tasks", "canonical_source": "https://oqoqo.ai", "published_at": "2026-08-10 21:27:04+00:00", "updated_at": "2026-08-10 21:41:45.348100+00:00", "lang": "en", "topics": ["ai-tools", "ai-agents", "developer-tools"], "entities": ["Oqoqo", "Codex", "Claude Code", "OpenClaw", "Hermes", "Pi", "Opencode", "Cursor"], "alternates": {"html": "https://wpnews.pro/news/show-hn-oqoqo-build-evals-and-custom-benchmarks-for-real-world-tasks", "markdown": "https://wpnews.pro/news/show-hn-oqoqo-build-evals-and-custom-benchmarks-for-real-world-tasks.md", "text": "https://wpnews.pro/news/show-hn-oqoqo-build-evals-and-custom-benchmarks-for-real-world-tasks.txt", "jsonld": "https://wpnews.pro/news/show-hn-oqoqo-build-evals-and-custom-benchmarks-for-real-world-tasks.jsonld"}}