cd /news/developer-tools/senior-agent-skills-make-your-agent-… · home topics developer-tools article
[ARTICLE · art-68202] src=adityaarakeri.github.io ↗ pub= topic=developer-tools verified=true sentiment=↑ positive

Senior-agent-skills, make your agent act as a senior

Aditya Arakeri released senior-agent-skills, a set of eight portable playbooks that make coding agents behave like careful senior engineers, working identically across Claude Code, Codex CLI, Google Antigravity, OpenCode, GitHub Copilot, and Cursor. The skills cover the full development loop from repo reconnaissance to committing, with a single SKILL.md format and zero dependencies.

read10 min views1 publishedJul 22, 2026

Agent skills · showcase and setup

Eight portable playbooks that cover the whole loop, from understanding a repo to committing the diff. One SKILL.md

format, working identically across Claude Code, Codex CLI, Google Antigravity, OpenCode, GitHub Copilot, and Cursor.

8 skills

6 harnesses

1 shared format

0 dependencies

Install once, inherit everywhere #

The installer symlinks each skill into every harness's directory, so you edit one canonical copy and all six pick it up. Restart your agent afterward so it rescans.

git clone git@github.com:adityaarakeri/senior-agent-skills.git; cd senior-agent-skills
./install.sh

./install.sh

User-level symlinks into the Claude, Codex, Gemini, and agents paths. The recommended default.

./install.sh --copy

Independent copies instead of symlinks, for when you want each harness to drift on its own.

./install.sh --project

Installs into the current repo, writing .github/skills

for Copilot so the team gets them on clone.

The eight skills #

Each one names a failure mode agents fall into, then hands the model a playbook to avoid it. They read in the order work actually happens.

> They form one loop: understand, plan, build, then review, prove, commit, and back around.

01· Understand

repo-recon

The problem. Dropped into unfamiliar code, an agent guesses the structure, edits the wrong file, and reinvents a helper that already existed.

Forces a mapping pass first: read the manifest, learn the build and test commands, trace one existing flow end to end, and match the house conventions before changing anything.

"I'm new to this repo. Map it before we add the webhook handler."

02· Plan

plan-first

The problem. On a wide or risky change, an agent sprints confidently in the wrong direction for twenty minutes before you get a chance to redirect it.

Produces a one-page plan (goal, non-goals, ordered steps, rollback story) you approve before any code is written. The non-goals section alone kills most scope creep.

"Use plan-first for adding rate limiting, then wait for my go-ahead."

03· Build

tdd-loop

The problem. "Looks correct" quietly stands in for "is correct," and the gap ships straight to production.

Strict red, green, refactor, with anti-cheating rules: never weaken an assertion to reach green, never mock the unit under test, never skip a failing test to come back later.

"Implement the discount rule test-first."

04· Diagnose

debug-protocol

The problem. Guess-and-check editing burns context, patches the symptom, and teaches the agent nothing about the real fault.

Reproduce first, form one falsifiable hypothesis at a time, print what the value actually is, bisect when lost, and fix the root cause rather than the crash site.

"This is failing in prod: <stack trace>. Use debug-protocol."

05· Restructure

safe-refactor

The problem. A refactor that silently changes behavior is the worst kind of bug: invisible in review, because everyone trusts the "no behavior change" promise.

Characterization tests first, structure and behavior changes in separate commits, one small mechanical step at a time with the suite run between each.

"Extract the auth logic with safe-refactor."

06· Review

self-review

The problem. The author's context is a blindfold. You see what the code was meant to do, which is exactly not what a reviewer needs to see.

Reads the full diff as a skeptical stranger, walks the unhappy paths, sweeps the changed lines for injection and secrets, and strips leftover debug code.

"Run self-review on the whole diff before we commit."

07· Prove

verify-done

The problem. "Should work" converts your one review pass into a debugging session, and trust erodes fast after the second time it happens.

An evidence gate before anything is called done: it builds, tests pass with real numbers reported, and the actual thing was run once on real input. Evidence or it did not happen.

"Use verify-done before you say it's finished."

08· Commit

git-hygiene

The problem. Messy history breaks bisect and blame, and a single unasked-for destructive command torches trust instantly.

Atomic commits with meaningful messages, plus a hard rule that force-push, reset, and clean require an explicit instruction naming the command before they run.

"Commit this with git-hygiene, root cause in the message."

Install and trigger, per harness #

The format is identical everywhere. Only the directory and the explicit-invoke syntax differ. Pick your tool.

reads from ~/.claude/skills/<name>/SKILL.md

Install

Personal in ~/.claude/skills/

, or per-repo in .claude/skills/

. A project skill shadows a personal one of the same name, so one repo can carry stricter rules.

Trigger

Describe the task and the matching skill loads on its own, or name it directly to be sure: "use the plan-first skill, then implement it."

Verify

Ask "which skills do you currently have available?" and confirm all eight are listed. Restart the session after any edit.

Install

Installer targets ~/.codex/skills/

. Codex also reads .agents/skills/

in a repo and ~/.agents/skills/

for a personal library.

Trigger

Type $

plus the name for an explicit call, $debug-protocol this test is flaky

, or let it match on the description. Browse everything with /skills

.

Verify

Run /skills

in the CLI. Disable one without uninstalling via an [[skills.config]]

block in ~/.codex/config.toml

, then restart.

Install

Replaced Gemini CLI, which stopped serving consumer tiers on June 18, 2026 (enterprise Code Assist keeps legacy access). Global skills: ~/.gemini/config/skills/

, still under ~/.gemini

. Project scope: .agents/skills/

, which --project

already writes.

Trigger

Same description matching; Antigravity kept Agent Skills through the transition, so the files carry over unchanged. A permission prompt appears when a skill activates. Approve it; that is expected, not an error.

Verify

Run /skills

in Antigravity CLI, or ask "what skills are available?" in the desktop app. Migration docs live at antigravity.google/docs/gcli-migration.

Install

No directory of its own needed. OpenCode natively searches the .claude/skills

and .agents/skills

paths (project and global), so the default install covers it twice.

Trigger

Loads through the native skill tool: the agent sees names and descriptions and pulls the full body only on a match. Describe the task or name the skill.

Verify

Discovery is cached at startup, so restart OpenCode after editing a skill. Per-agent permissions support wildcards like internal-*

.

Install

Repo-scoped. Canonical path is .github/skills/

, committed so the team gets it on clone. Run ./install.sh --project

from the repo root. Copilot also auto-reads a repo's .claude/skills/

.

Trigger

Description matching at session startup, across the coding agent, Copilot CLI, and VS Code agent mode. The most specific description wins when several could apply.

Verify

Type /skills

in the VS Code Chat panel. Copilot reads only name

and description

from frontmatter, which this pack respects.

Install

Project skills live in .cursor/skills/

or the shared .agents/skills/

path. The installer also targets ~/.cursor/skills/

for Cursor builds with personal-skill support.

Trigger

Cursor discovers skills from their descriptions and loads the full instructions when the task matches. Name a skill directly when you want to force an explicit workflow.

Verify

Start a fresh agent session after installation, then ask which skills are available and confirm all eight appear.

Preferred model and effort #

A common question: do the skills pick the model? They do not, on purpose. Here is how the layers divide, and what to set for the best results.

The harness owns model and effort. The skill inherits it. No SKILL.md names a model or sets a reasoning budget, so the same file runs identically on all six harnesses. A skill is instructions handed to whatever engine is already running, the way a recipe cannot choose your oven. Set model and effort once at the harness level and every skill picks it up.

That said, the skills split cleanly into two kinds of work, and matching effort to the kind is where quality comes from.

Higher reasoning

Thinking-heavy skills

  • plan-first
  • debug-protocol
  • safe-refactor

Design decisions, hypothesis chains, and blast-radius analysis reward a stronger model and a raised thinking budget. Give these the headroom.

Standard reasoning

Procedural skills

  • repo-recon
  • tdd-loop
  • self-review
  • verify-done · git-hygiene

These are follow-the-checklist work. A default model at normal effort executes them faithfully, so save the budget for where it moves the needle.

Where each harness sets it:

Model via /model

or config; thinking effort is exposed on Opus-class models.

Model plus a reasoning-effort level from minimal to high, set by config or flags.

Model selected at the session level; the skill runs on whatever is active.

Provider-agnostic. Model set per session or pinned to a named agent, Claude, GPT, or local.

Session model selection, with auto-routing in newer builds, entirely outside the skill.

Model selected per agent session; skill instructions run on the active model without choosing one themselves.

One nuance: some harnesses let a named subagent pin a specific model, and a skill's text can point at such an agent. That is still the harness choosing the model, invoked by name, not the skill reaching down to swap engines. The pack keeps all eight free of that so they stay portable.

How they chain in practice #

On a real task you rarely touch all eight. The skills reference each other where the phases connect, so a natural chain forms on its own.

Greenfield feature

New code, unfamiliar repo

  • repo-recon maps the codebase
  • plan-first outlines it, you approve
  • tdd-loop builds it red then green
  • self-review reads the diff
  • verify-done proves it is finished

Production bug

Broken, cause unknown

  • debug-protocol reproduces and diagnoses
  • tdd-loop writes the regression test first
  • verify-done confirms with evidence
  • git-hygiene commits, root cause in the message

Refactor, thin coverage

Restructure without breaking

  • safe-refactor writes characterization tests
  • structure changes in small verified steps
  • self-review confirms zero behavior change
  • verify-done checks old names return no hits

Unattended run

CI or full-auto

  • name skills explicitly in the command
  • implicit matching is weaker on terse prompts
  • always include verify-done
  • the run reports evidence, not a guess

When something misfires #

The fix is almost always in the description

frontmatter, not the skill's body. Descriptions are the entire matching mechanism.

Skill never appears in the available list #

Discovery failed. Check, in order: the file sits at the exact path the harness scans, the folder name matches the name

in frontmatter, the first line is exactly ---

with valid YAML under it, and you restarted the session. OpenCode and Codex both need a restart after changes.

Skill appears but never activates on its own #

The description is not matching your phrasing. Name the skill explicitly for an immediate fix, or edit the description

to include the phrases you actually use, then restart. The body has no effect on triggering.

Skill activates for the wrong tasks #

The description is too broad. Narrow it: drop generic trigger words, add the specific contexts where it should fire, and state plainly what it should not cover.

Edited a skill and nothing changed #

You are mid-session and the harness cached the old version at startup. Restart. This bites hardest in OpenCode and Codex, which cache aggressively.

Antigravity keeps prompting, or Gemini CLI went dead #

The prompt is not a bug: Antigravity asks permission when a skill activates, same as Gemini CLI did. Approve and continue. A dead gemini

command after June 18, 2026 is the deprecation, not a skills problem: consumer tiers moved to Antigravity CLI, and your SKILL.md files carry over unchanged once placed in ~/.gemini/config/skills/

.

Copilot's /skills list is empty #

Copilot is repo-scoped. Run the installer with --project

from inside the repo so skills land in .github/skills

. If the extension predates the January 2026 stable rollout, update VS Code and the Copilot extension, then check again.

A destructive git command ran without asking #

git-hygiene

is instructions, not enforcement, and an agent can ignore it under pressure. For a hard guarantee, pair it with your harness's own permission or hook system so the dangerous commands are blocked at the tool layer, not just discouraged in prose.

── more in #developer-tools 4 stories · sorted by recency
── more on @aditya arakeri 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/senior-agent-skills-…] indexed:0 read:10min 2026-07-22 ·