{"slug": "measure-an-ai-coding-harness-from-l0-to-l4-with-harness-score", "title": "Measure an AI Coding Harness from L0 to L4 with Harness Score", "summary": "A developer released Harness Score, an open-source scanner that grades an AI coding harness from L0 to L4 by checking repository evidence such as context files, scoped rules, skills, hooks, sensors, and CI. The tool makes zero LLM calls and zero network requests, and a companion GitHub template lab walks users through improving a Meeting Cost CLI; a clean template clone scored 17/108 (L0 - Unharnessed) on October 1, 2026 with package version 1.7.5.", "body_md": "AI coding assistants can edit a repository quickly, but speed does not tell you whether the repository can catch a bad edit. A project with no instructions, tests, or CI may produce a plausible patch and still leave every important check to chance.\n\nThis tutorial shows how to measure that surrounding system with [Harness Score](https://github.com/paladini/harness-score) and use a small open-source lab to improve it step by step. The result is not a quality certificate. It is a deterministic checklist of repository evidence: context files, scoped rules, skills, hooks, sensors, CI, and hygiene.\n\n`npx --yes harness-score .` and record the initial level and score.\nThe tutorial documents Git, Node.js 24 or newer, and an AI coding agent that can edit a local repository. GitHub CLI is optional if you want to create a private or public copy from the template. The scanner package itself declares Node.js `>=18` in its npm metadata, but following the lab's Node.js 24 prerequisite keeps the exercise aligned with its current README.\n\nNo API key is required: Harness Score is designed to make zero LLM calls and zero network requests while scanning a repository.\n\nThe lab is a GitHub template rather than a finished application. That is intentional. You create the small Meeting Cost CLI during the first stage, so the harness improvements remain visible instead of being hidden inside an already-complete project.\n\nUsing GitHub CLI, create a private copy and clone it:\n\n```\ngh repo create my-harness-lab \\\n  --template paladini/harness-score-tutorial \\\n  --private \\\n  --clone\ncd my-harness-lab\n```\n\nIf you do not want to create a remote repository, clone the public template locally and remove its origin:\n\n```\ngit clone https://github.com/paladini/harness-score-tutorial.git my-harness-lab\ncd my-harness-lab\ngit remote remove origin\n```\n\nThe template README is written in Portuguese, but the commands and paths are ordinary Git, Node.js, and Harness Score workflows. The prompts work with different coding agents. Inspect every diff yourself and do not allow an agent to commit automatically.\n\nRun the exact current command documented by the template and the scanner project:\n\n```\nnpx --yes harness-score --version\nnpx --yes harness-score .\n```\n\nThe current published package is `1.7.5` and the current template clone is expected to start at L0, because it contains the tutorial and license but not the application harness. In a clean clone of the template, the scanner reported `17/108` and `L0 - Unharnessed` on October 1, 2026.\n\nThat number belongs to one commit and scanner version, not a permanent promise. The tutorial runs an unpinned command, so record the tool version beside every baseline.\n\nFor machine-readable evidence, add `--json`:\n\n```\nnpx --yes harness-score . --json > baseline.json\n```\n\nKeep that report outside the repository if you do not want it to affect future scans. The JSON report includes the maturity level, earned points, dimensions, checks, evidence, and remediation links.\n\nRun one prompt, inspect the resulting files, execute local checks, scan the repository, and commit a checkpoint only after review.\n\nStart by creating the smallest functional Meeting Cost CLI. The application accepts participants, meeting minutes, and hourly labor cost, then calculates the total. Keep the domain calculation separate from terminal argument handling and use Node.js built-ins only.\n\nAfter confirming that the command works, add a substantive root `AGENTS.md`. It should describe the real files, commands, domain invariants, error handling, and security boundaries. Do not add future-stage artifacts early. A long file is not automatically useful: the scanner can detect that a context file is present and substantive, but it cannot decide whether every rule is true.\n\nRun the scan again and inspect the failed checks. The next-level message is more useful than the headline score because it tells you which dimension is blocking progress.\n\nMove procedural guidance into artifacts that load when needed. The lab asks for a path-scoped rule, a reusable skill, and an explicit workflow. It also adds a `.gitignore` and a lockfile.\n\nThis separation matters. A root context file orients every session; a scoped rule applies to relevant paths; a skill packages a repeatable procedure. The hygiene checks then make local state and credentials less likely to enter a commit or an agent context.\n\nDo not treat a passing hygiene check as proof that a repository is safe. It only means the scanner found the expected structural evidence, such as ignored environment files and no credential signatures in harness files.\n\nAt this stage, add actual feedback: tests for valid and invalid calculations, a formatter, a linter, strict type checking, and a GitHub Actions workflow. The tutorial uses fixed development-tool versions for this stage and asks you to run the complete check command locally.\n\nThe important design choice is independent feedback. A README claim that tests exist is weaker than a test file that runs. A local command is weaker than a CI job that repeats the test, lint, and type checks on every push or pull request. Harness Score checks for the presence and wiring of these sensors; you still need to review whether the tests cover meaningful behavior.\n\nThe final stage adds two Cursor hooks: a gate hook that denies dangerous shell commands and a feedback hook that formats supported files after edits. The tutorial also adds tests for allowed, denied, and malformed hook payloads.\n\nHooks are useful because they execute at a boundary where prose can be ignored. They are not a universal security boundary. Keep scripts local, review their inputs, fail safely on malformed payloads, and retain CI as an independent check. The scanner's L4 result means the expected hook configuration and evidence were detected. It does not prove that every possible destructive command is blocked.\n\nFinally, add a separate GitHub Actions workflow that runs `npx --yes harness-score . --min-level 4` or the equivalent project action. The gate turns the maturity level into a regression check, so removing a hook or sensor becomes visible in review.\n\nUse the same small loop after every stage:\n\n```\ngit diff --stat\nnpm run check\nnpx --yes harness-score . --json\ngit status --short\n```\n\nThe first command checks scope. The second checks product behavior and sensors once they exist. The third checks harness evidence. The last catches generated files or local state that should not be committed.\n\nIf the score differs from the README's approximate table, compare the scanner version and inspect the JSON checks rather than adding decorative files to chase points.\n\nThe scanner measures filesystem and configuration facts. It does not establish that the Meeting Cost CLI is commercially useful, that tests are comprehensive, that rules are correct, or that a team actually reviews pull requests. A high score means that more feedback and guardrails are present. It does not mean an agent can be trusted without human review.\n\nThe score can change when the scanner's maturity model changes. Pin a version for release gates when stable comparisons matter. Do not compare scores from different versions without recording that difference.\n\nThe tutorial itself is a template. Its prompts are guidance for an agent, not an automatic migration. Review changes, preserve the stated file boundaries, and keep credentials out of the repository. Read the [Harness Score license](https://github.com/paladini/harness-score/blob/main/LICENSE) and the [tutorial license](https://github.com/paladini/harness-score-tutorial/blob/main/LICENSE) before redistributing either project.\n\nNo. The project describes its checks as deterministic filesystem facts and says the scanner makes zero LLM calls and zero network requests during a scan.\n\nNo. The README says the prompts work with different coding agents. Cursor is used for the runtime hook example because its hook format is recognized by the scanner.\n\nNo. Use the failed checks to choose controls that fit your repository. A hook that is irrelevant to your threat model may add noise, while an untested deployment script may deserve attention even if it does not change the score.\n\nAn AI coding harness is an engineering system around the model. Measure it from a known commit, improve one feedback layer at a time, and preserve evidence in code, tests, hooks, and CI. The useful outcome is not a magic number. It is a repository where an agent has clearer context, faster feedback, and fewer ways to make an unsafe change silently.\n\nWhat is the first missing harness layer in your repository today: durable context, scoped procedures, sensors, CI, or runtime guardrails?\n\n*AI assistance disclosure: I used AI assistance to organize and edit this tutorial. The repository behavior, package metadata, current version, commands, and baseline scan were checked against the linked primary sources and a clean clone on October 1, 2026.*", "url": "https://wpnews.pro/news/measure-an-ai-coding-harness-from-l0-to-l4-with-harness-score", "canonical_source": "https://dev.to/paladini/measure-an-ai-coding-harness-from-l0-to-l4-with-harness-score-2g89", "published_at": "2026-10-01 12:36:22+00:00", "updated_at": "2026-10-01 12:44:26.405242+00:00", "lang": "en", "topics": ["ai-agents", "developer-tools", "ai-tools", "mlops"], "entities": ["Harness Score", "GitHub", "Node.js", "Meeting Cost CLI", "paladini"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/measure-an-ai-coding-harness-from-l0-to-l4-with-harness-score", "markdown": "https://wpnews.pro/news/measure-an-ai-coding-harness-from-l0-to-l4-with-harness-score.md", "text": "https://wpnews.pro/news/measure-an-ai-coding-harness-from-l0-to-l4-with-harness-score.txt", "jsonld": "https://wpnews.pro/news/measure-an-ai-coding-harness-from-l0-to-l4-with-harness-score.jsonld"}}