cd /news/developer-tools/agent-skills-that-bring-team-coding-… Β· home β€Ί topics β€Ί developer-tools β€Ί article
[ARTICLE Β· art-86293] src=github.com β†— pub= topic=developer-tools verified=true sentiment=Β· neutral

Agent skills that bring team coding standards to Claude Code and Codex

Tikal released ADLC Team Skills, an open-source team layer for the Twelve-Factor Agentic SDLC, designed to bring team coding standards to AI coding agents such as Claude Code, Codex, OpenCode, Cursor, and GitHub Copilot. The toolkit includes slash commands, session-start event hooks, and a companion repository for version-controlled team directives, aiming to replace chaotic 'vibe coding' with compliant, accountable AI-assisted development.

read16 min views1 publishedAug 4, 2026
Agent skills that bring team coding standards to Claude Code and Codex
Image: source

Stop Vibe Coding in Silos. Build a Shared Cognitive Layer for Your Engineering Team.

Individual prompt hacks create quick wins for solo developers, but when scaled across a team, "vibe coding" leads to chaotic technical debt, context rot, unreviewable PRs, and lost code ownership. Speed is solved; Trust and Verification is the new bottleneck in AI engineering.

ADLC Team Skills (tikalk/adlc-team-skills

) is the open-source Team Layer of the Twelve-Factor Agentic SDLC. It turns AI agents from isolated guessers into compliant, accountable team members that share your team's constitution, product strategy, architectural standards, and evaluation benchmarks.

Team AI Directives ( tikalk/agentic-sdlc-team-ai-directives) is the companion repository that holds your team's version-controlled context modules (constitution, rules, personas, examples), CDR index, and skills manifest.

npx adlc-skills-cli add tikalk/adlc-team-skills -a opencode

npx skills add tikalk/adlc-team-skills -a claude -g

Works out of the box with any agent supporting the Agent Skills standard β€” Claude Code, Codex, OpenCode, Cursor, GitHub Copilot, and others.

Slash commands + events: adlc-skills-cli wraps

npx skills add

and additionally generates /name

slash commands and wires session_start

event hooks (via .events.json

) for 9 coding agents. Skills repos without .events.json

get commands only.Universal orchestration: mission-brief

auto-discovers skills from any source (mattpocock/skills, addy osmani/agent-skills, superpowers, spec-kit, or your own) and dynamically wires them into the mission pipeline. No vendor lock-in.

team-boot

auto-runs at session start via the event hook. On an unconfigured project it outputs a warning telling the user to run /team-setup

. team-setup

is also available on demand:

npx adlc-skills-cli add tikalk/adlc-team-skills -a opencode   # install skills + commands + events

Then choose Mode 3 β€” Scaffold new empty team-ai-directives, or Mode 1 β€” Clone from GitHub to fork tikalk/agentic-sdlc-team-ai-directives:

team-setup            β†’ pick destination (default ./team-ai-directives) + team name
                      β†’ scaffolds README / AGENTS.md / CDR.md / .skills.json /
                        constitution placeholder / OKF index files + git init

team-constitution     β†’ interactively replace the placeholder with your real principles

team-boot (auto)      β†’ assembles constitution + CDR index + PDR/ADR indexes + skills registry
                        into the system prompt at session start

Already have a directives repo? team-setup

offers three other modes:

Mode 1 β€” Clone from GitHub(e.g. fork)tikalk/agentic-sdlc-team-ai-directives

Mode 2 β€” Point to existing local path(wire a repo you already have)** Mode 4 β€” Already configured**(verify an existing setup)

Without ADLC (Vibe Coding) With ADLC Team Skills
Session starts from zero: Agent knows nothing about your architecture, team rules, or deprecated patterns.
Auto-bootstrapped context: team-boot auto-loads your Team Constitution & active decisions on session start.
Prompt wall bloat: Dumping a 10,000-token prompt wall wastes tokens and causes model instruction drift.
Progressive disclosure: team-boot injects a ~100-token index; team-discover fetches only the 1–2 rules relevant to the task.
Ambiguity leads to guessing: Agent invents functions or database schema instead of asking questions.
Contract-first specs: mission-brief defines Goal, Constraints, Non-Goals, and Success Criteria before writing code.
Learnings evaporate: Debugging fixes and newly discovered edge cases disappear when the chat ends.
Closed feedback loop: levelup-specify extracts session execution traces and commits new rules directly to Git.
Silent regressions: Prompt edits or base model updates silently degrade agent output.
Verification-first evals: Automated LLM judges and binary graders test code against business risks before human review.
                             [ THE GREAT FILTER ]
                      (Human Team Lead Macro-Review)
                                     β–²
                                     β”‚
                  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                  β”‚   Pillar 4: Governance & Evals      β”‚
                  β”‚   (Tier 1 Fast Checks + LLM Judges) β”‚
                  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–²β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                     β”‚
                  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                  β”‚   Pillar 3: Spec-Driven Workflow    β”‚
                  β”‚   (Contract-First Mission Pipeline) β”‚
                  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–²β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                     β”‚
                  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                  β”‚   Pillar 2: Product & Architecture  β”‚
                  β”‚   (Product PDRs + Architecture ADRs)β”‚
                  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–²β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                     β”‚
                 β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                 β”‚   Pillar 1: Strategy & Team Directivesβ”‚
                 β”‚   (team-boot / levelup / CDR repository) β”‚
                 β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Put the team at the center of your AI strategy. Instead of individual developers hoarding prompt shortcuts on local machines, team standards live in a version-controlled Git repository ( team-ai-directives).

: Auto-runs at session start via the event hook, assembling the team constitution, CDR index, PDR/ADR indexes, and skill registry into the system prompt.team-boot

Progressive Disclosure (No Token Waste): Instead of dumping massive prompt walls into every session,team-boot

injects a compact ~100-token index.team-discover

loads full rule bodiesonlywhen relevant to the active task.: Interactively define, review, or amend your engineering team's core principles.team-constitution

: Re-index CDR.md, scan for rule conflicts, and verify directive freshness.team-repair

team-boot        β†’ assembles constitution + CDR index + PDR/ADR + skills into system prompt
team-discover    β†’ manual re-scan for structured discovery tables (/team-discover)
team-constitution β†’ create or amend the team constitution interactively
team-repair      β†’ re-index CDR.md, scan for conflicts, verify freshness

Without documented decisions, every implementation session re-derives (or misinterprets) product intent and architectural rules.

Product Decision Records (: Capture product decisions as individual PDR files, resolve ambiguities through an interactive clarification workflow, and compile them into a self-containedproduct-*

)PRD.md

.Architectural Decision Records (: Reverse-engineer or define architectural decisions using Rozanski & Woods viewpoints (Functional, Security, Deployment, Performance) and compose them into a unifiedarchitect-*

)AD.md

.: Track milestone progress across four layers of truth β€” decisions (PDRs), execution (live issues via MCP), code evidence, and milestone gates.product-roadmap

Product:      product-init β†’ product-clarify β†’ product-implement β†’ product-analyze
Architecture: architect-init β†’ architect-clarify β†’ architect-implement β†’ architect-analyze
Roadmap:      product-roadmap (tracks PDRs + issues + code + gates)

AI is an obsessive guesser β€” when faced with ambiguity, it invents solutions instead of asking questions. Move from a Conversational model to a Contract model.

: The team's autonomous pipeline runner. Takes a feature prompt, derives a formal contract (Goal, Constraints, Non-Goals, Success Criteria), generates an ordered step list, and walks amission-brief

specify β†’ plan β†’ tasks β†’ implement β†Ί converge

loop to completion.Mantra: "Debug the Spec, Not the Code": When an agent makes a mistake, don't just patch the code β€” add the missing constraint to the specification so the mistake is never repeated.

mission-brief "add user profile API with JWT"
  β”œβ”€β”€ Phase 2: Brief (Goal, Constraints, Non-Goals, Success Criteria)
  β”œβ”€β”€ Phase 3: Route Classification (spec | change | quick)
  β”œβ”€β”€ Phase 4: Discovery (auto-wires local installed skills & SDD frameworks)
  └── Phase 5: Execute (specify β†’ plan β†’ tasks β†’ implement β†Ί converge)

Never let the agent that wrote the code decide if the code is good. "Separate the Maker from the Checker."

: Build application-level evaluation suites (PromptFoo or DeepEval) using Eval-Driven Development (EDD). Runs Tier 1 fast checks + Tier 2 LLM judge subagents to test code against defined business risksevals

skillsbeforehuman macro-review inThe Great Filter.: Capture session wins into permanent team memory.levelup

levelup-specify

extracts session execution traces and commits them to Git as reusable rules (Context Directive Records β€” CDRs)."Build to Delete": Prune outdated rules and prompt scaffolding as underlying foundation models improve usingteam-repair --build-to-delete

.

LevelUp: levelup-init β†’ levelup-specify β†’ levelup-clarify β†’ levelup-publish
Evals:   evals-init β†’ evals-specify β†’ evals-clarify β†’ evals-implement β†’ evals-validate β†’ evals-analyze

Factor XII β€” Build to Delete. Factor XIII β€” Loop Engineering.

mission-brief

acts as an open, vendor-agnostic orchestrator across all popular agent skill repositories and Spec-Driven Development (SDD) frameworks:

SDD Framework / Skill Source Supported Workflows
Agentic SDLC Spec-Kit (tikalk/agentic-sdlc-spec-kit )
Twelve-Factor SDD pipeline, native specify CLI discovery, contract verification
Spec-Kit (specify_cli )
Native command discovery and specification templates
OpenSpec
Structured edge-case contracts and verification schemas
/tdd , /grill-me , /grill-with-docs , /code-review , /prototype
Exit criteria checklists, quality gate skills
Developer tooling & workflow skills
ADLC Team Skills (this repo)
product-specify , architect-specify , evals-validate , levelup-specify
Your Custom Skills
Any skill following the SKILL.md standard

Discoveryβ€” At mission start,mission-brief

scans all skills directories (.claude/skills

,.agents/skills

, etc.) and reads everySKILL.md

frontmatter to build a vendor-agnostic inventory of installed skills with their names and descriptions.LLM-decided routingβ€” Each step's delegation prompt includes the full skills inventory. The subagent decides which skill (if any) fits the current phase β€” the LLM matches, not a brittle lookup table.Graceful fallbackβ€” If no skill matches, the subagent executes directly. If a skill matches, it's invoked. Either way, the mission pipeline continues.

npx skills add mattpocock/skills
npx skills add tikalk/adlc-team-skills

mission-brief "add user profile API with JWT"

Skills are organized under the four pillars of the Twelve-Factor Agentic SDLC, flattened directly under the skills/

directory:

skills/
β”œβ”€β”€ architect/             # architect-* (5 skills)
β”œβ”€β”€ product/               # product-* (6 skills) + product-templates/
β”œβ”€β”€ levelup/               # levelup-* (4 skills) + levelup-helpers.{sh,ps1}
β”œβ”€β”€ mission-brief/         # core SDD orchestrator (1 skill)
β”œβ”€β”€ evals/                 # evals-* (6 skills) + evals-templates/
β”œβ”€β”€ tech-radar/            # tech-radar-* (1 skill) + resources/radar.json
β”œβ”€β”€ workspace/             # workspace (1 skill) β€” multi-repo coordination
└── team/                  # team-* (6 skills) + team-helpers.{sh,ps1}

This places every single skill exactly 2 levels deep, fully resolving the default depth limit of the skills

CLI and ensuring all skills install out of the box.

β€” Bootstrap session: assembles constitution, CDR index, PDR/ADR indexes, and skill registry into the system prompt at session start. Outputs a warning to runteam-boot

/team-setup

on unconfigured projects.β€” Manually re-scan team context modules and produce a structured discovery table. Available viateam-discover

/team-discover

.β€” Clone, scaffold, or configure a team AI directives repository. Model-invoked byteam-setup

team-boot

(self-install) and available on demand. Say "Set up team directives for this project."

β€” Create or amend the team constitution interactively. Say "Create our team constitution" or "Amend our team principles."team-constitution

β€” Re-index CDR.md, .skills.json, AGENTS.md; health check; conflict scan; freshness verification. Say "Check our team directives health" (team-repair

--health-only

), "Repair our CDR index," or "Scan for rule conflicts" (--conflicts

).β€” Browse and install team skills from the team AI directives. Say "Show me available team skills."team-skills

All user-invoked. Capture and publish reusable patterns to team-ai-directives, including paired directive compliance evals.

β€” Brownfield CDR discovery from existing codebase, including paired eval CDRs from code patterns. Say "Discover directives from this codebase."levelup-init

β€” Extract CDRs and paired eval CDRs from the current session. Say "Extract lessons from this session."levelup-specify

β€” Review, accept, reject, or defer pending CDRs. Evals regression gate runs by default. Say "Review pending CDRs."levelup-clarify

β€” Compile accepted CDRs into team directives artifacts, evals goldensets, and draft PR. Say "Publish accepted CDRs" or "Build one skill from a CDR" (levelup-publish

--skill CDR-NNN

).

All user-invoked. Document product decisions as individual PDRs and compile into a self-contained PRD.md.

β€” Brownfield PDR discovery from existing codebase and documentation. Say "Discover product decisions from this codebase."product-init

β€” Greenfield PDR creation through interactive product exploration. Say "Let's define our product strategy."product-specify

β€” Refine, validate, and approve PDRs before PRD generation. Say "Review our product decisions."product-clarify

β€” Generate PRD.md from accepted PDRs (multi-agent DAG orchestration). Say "Generate our PRD."product-implement

β€” Read-only PDR↔PRD consistency and quality analysis. Say "Analyze our product docs."product-analyze

β€” Track milestone progress: decision status, live issues via MCP, code evidence, and gates. Say "Show roadmap progress."product-roadmap

All user-invoked. Create and manage Architecture Decision Records using the Rozanski & Woods methodology.

β€” Reverse-engineer ADRs from an existing codebase (brownfield). Say "Reverse-engineer architecture from this codebase."architect-init

β€” Create ADRs from a PRD or feature description (greenfield). Say "Create ADRs from this PRD."architect-specify

β€” Refine and validate existing ADRs. Say "Refine and validate my ADRs."architect-clarify

β€” Generate an Architecture Description (AD.md) from accepted ADRs. Say "Generate AD.md from my ADRs."architect-implement

β€” Check ADR↔AD consistency and architecture quality. Say "Analyze architecture consistency."architect-analyze

All user-invoked. Build and maintain application-level evaluation suites following EDD (Eval-Driven Development) principles (PromptFoo or DeepEval).

β€” Initialize evaluation directory structure (evals-init

evals/{system}/

) with security baseline. Say "Initialize my evaluation harness."β€” Extract eval criteria from specs and production failure traces (bottom-up open coding). Say "Specify evaluation criteria from this failure log."evals-specify

β€” Cluster related patterns, refine criteria, isolate 20% holdout split, and publish goldset. Say "Clarify and accept my draft evaluations."evals-clarify

β€” Generate executable graders and test configs, automatically running unit tests to verify evaluator correctness. Say "Generate graders from the goldset."evals-implement

β€” Run the evaluation pyramid (Tier 1 fast checks + Tier 2 LLM judges) and compute quality metrics (TPR/TNR, SLA headroom). Say "Validate my evaluation suite."evals-validate

β€” Deep-analyze trajectory failure traces, routing spec-level failures toevals-analyze

levelup-specify

(rules) and generalization failures to backlog. Say "Analyze evaluation failures."

User-invoked. Structure a feature description into a Mission Brief and run it end-to-end with any installed SDD skill set.

β€” Takes a description, structures it into a Mission Brief (goal, constraints, success criteria), generates an ordered step list with prompts that trigger installed SDD skills, and walks those steps to converged implementation. Sync (gated) ormission-brief

--async

(ungated, checkpoint across sessions). Say "Build this feature end to end" ormission-brief "add dark mode"

. Resume withmission-brief --resume

.

Model-invoked. Grounds tech stack choices in Tikal's Israeli Tech Radar.

β€” Discovers technologies implied by the prompt, matches them against the Tikal Tech Radar (tech-radar-context

radar.json

), and injects a context table with each technology's adoption ring (Keep

/Start

/Try

/Stop

), quadrant, and Tikal's "Why?" opinion β€” plus Tikal-aligned alternatives for anything onStop

. Auto-triggered whenever a technology, framework, database, library, or cloud tool is being chosen or evaluated. Fetches the live radar best-effort and falls back to a bundled snapshot atresources/radar.json

.

User-invoked. Multi-repo workspace coordination for shared team context.

β€” Discover child repos at depth 1 and optionally link them as Git submodules, creating a multi-repo workspace analogous to VS Code'sworkspace

.code-workspace

. The parent repo holds shared PDRs, ADRs, and CDRs under.adlc/

(created byproduct-specify

,architect-specify

,levelup-specify

); child implementation repos are linked for unified context. Commands:/workspace

(discover),/workspace --link

(register submodules),/workspace --status

(audit: branch, dirty, unpushed, SHA drift). Noworkspace.yml

β€” pure auto-discovery by convention. Say "Set up multi-repo workspace" or/workspace --link

.

Output File Layout

All skills write to .adlc/

(project root) and the team AI directives repo.

Team Directives (inside the team AI directives repository):

AGENTS.md

β€” agent instructions ( order, rules, skills)CDR.md

β€” index of approved context contributions.skills.json

β€” skills manifest (schema v2.0.0).mcp.json.example

β€” MCP servers config examplecontext_modules/constitution.md

β€” team constitution (OKF frontmatter)context_modules/{rules,personas,examples}/**/*.md

β€” context modulescontext_modules/{type}/index.md

β€” progressive disclosure per concept typecontext_modules/{type}/log.md

β€” chronological change log per concept typeskills/{name}/SKILL.md

+.skills-entry.json

β€” published team skillsevals/{directive-id}/goldset.md

+goldset.json

β€” directive compliance goldensets

LevelUp (inside .adlc/

of the target project):

.adlc/drafts/cdr/CDR-{NNN}.md

β€” proposed/discovered CDRs (including eval CDRs).adlc/drafts/cdr/cdr.md

β€” auto-generated CDR index.adlc/init-options.json

β€” team AI directives path config

Product (inside .adlc/

and repo root):

.adlc/drafts/pdr/PDR-{NNN}.md

β€” proposed/discovered PDRs.adlc/drafts/pdr/pdr.md

β€” auto-generated PDR index.adlc/memory/pdr/PDR-{NNN}.md

β€” accepted/completed PDRs.adlc/memory/pdr/pdr.md

β€” accepted PDR index.adlc/product/sections/{feature-area}/{section}.md

β€” PRD section build artifacts.adlc/product/state.json

β€” DAG execution statePRD.md

β€” Product Requirements Document (repo root)

Architecture (inside .adlc/

and repo root):

.adlc/drafts/adr/ADR-{NNN}.md

β€” proposed/discovered ADRs.adlc/drafts/adr/adr.md

β€” auto-generated ADR index.adlc/memory/adr/ADR-{NNN}.md

β€” accepted ADRs.adlc/memory/adr/adr.md

β€” accepted ADR indexAD.md

β€” Architecture Description (repo root).adlc/architect/

β€” per-view DAG artifacts

Missions (inside .adlc/

of the target project):

.adlc/workflow/workflow-config.yml

β€” mission execution/supervision/budgets config.adlc/workflow/.mission-state.json

β€” step list, completed steps, brief, discovery results.adlc/workflow/runs/<feature>/mission-log.json

β€” final audit trail.adlc/workflow/runs/<feature>/iterations.md

β€” per-implement audit entries

Governance (inside target project and repo root):

.adlc/drafts/evals/EVAL-{NNN}.md

β€” proposed/discovered eval criteria drafts.adlc/drafts/evals/evals.md

β€” draft evals index.adlc/memory/evals/EVAL-{NNN}.md

β€” accepted/completed eval criteria.adlc/memory/evals/evals.md

β€” accepted evals index.adlc/memory/evals/holdout.json

β€” isolated/reserved holdout test datasetevals/{system}/goldset.md

β€” published goldset (human-readable)evals/{system}/goldset.json

β€” published goldset (machine-readable)evals/{system}/config.yml

β€” evaluation framework configurationevals/{system}/config.{js,py}

β€” framework test configevals/{system}/graders/check_*.py

β€” generated binary Python graders / metricsevals/{system}/tests/test_check_*.py

β€” generated unit tests verifying grader correctnessevals/results/validation_report.md

β€” statistical validation results report

Workspace (inside parent repo root):

.gitmodules

β€” Git submodule registrations for child repos (created by--link

).adlc/

β€” shared team context (PDRs, ADRs, CDRs); parent is the single source of truth- Child repos discovered at depth 1; each child's .adlc/

presence is reported (informational)

OKF Compliance

Generated context modules include Open Knowledge Format (OKF) v0.1 compliant frontmatter alongside custom fields.

OKF field Status Source
type
βœ… CDR context type
title
βœ… CDR title
description
βœ… CDR descriptor
resource
βœ… Relative path to artifact
tags
βœ… Context type tag
timestamp
βœ… ISO 8601 datetime

Custom fields co-exist with OKF frontmatter: id

, cdr_ref

, created

, modified

, verified

, age_days

, evidence

.

Directory structure: context_modules/{type}/index.md

(progressive disclosure), context_modules/{type}/log.md

(change history), cross-links between related concepts.

Workflows

Team Directives setup:

team-setup β†’ team-constitution β†’ team-boot (auto at session start)

Product lifecycle:

Brownfield: product-init β†’ product-clarify β†’ product-implement β†’ product-analyze
Greenfield: product-specify β†’ product-clarify β†’ product-implement β†’ product-analyze
Roadmap:    product-roadmap (anytime)

Architecture lifecycle:

Brownfield: architect-init β†’ architect-clarify β†’ architect-implement β†’ architect-analyze
Greenfield: architect-specify β†’ architect-clarify β†’ architect-implement β†’ architect-analyze

LevelUp / CDR lifecycle:

Brownfield: levelup-init β†’ levelup-clarify β†’ levelup-publish β†’ team-repair
Session:    levelup-specify β†’ levelup-clarify β†’ levelup-publish β†’ team-repair
Build to Delete: team-repair --build-to-delete β†’ levelup-clarify (review deletion CDRs)

Mission:

mission-brief "feature" β†’ review brief β†’ execute steps β†’ converge β†’ mission-log.json

Multi-repo workspace:

product-specify / architect-specify β†’ create shared PDRs/ADRs in parent .adlc/
workspace --link β†’ register child repos as submodules
workspace --status β†’ audit branch, dirty, unpushed, SHA drift

Application Evaluation lifecycle:

Greenfield (Spec-Driven): evals-init β†’ evals-specify (from spec) β†’ evals-clarify β†’ evals-implement β†’ evals-validate
Brownfield (Error-Driven): evals-init β†’ evals-specify (from failures) β†’ evals-clarify β†’ evals-implement β†’ evals-validate β†’ evals-analyze

Full product β†’ architecture β†’ team:

Product:     product-specify β†’ product-clarify β†’ product-implement β†’ product-analyze
Architecture: architect-specify β†’ architect-clarify β†’ architect-implement β†’ architect-analyze
Team:        levelup-specify β†’ levelup-clarify β†’ levelup-publish β†’ team-repair

12-Factor Alignment

Factor Skills How
III β€” Mission Definition
Product skills PRD/PDR lifecycle ensures product decisions are documented, reviewed, and traceable before execution
IV β€” Structured Planning
Architecture skills ADRs and AD.md provide structured planning artifacts using Rozanski & Woods viewpoints
VII β€” Verification-First Evals
LevelUp + Evals skills LevelUp creates directive-compliance eval CDRs; evals skills build and run application-level evaluation suites (PromptFoo/DeepEval) with binary graders, holdout splits, and statistical validation
VIII β€” Ratchet Effect
LevelUp + Evals skills Each session extracts eval CDRs alongside directive CDRs; each goldset publication adds criteria that monotonically increase quality β€” evals-clarify publishes, evals-validate enforces
IX β€” Traceability
Product + Architecture Every decision traces from PDR β†’ PRD β†’ feature and from ADR β†’ AD β†’ code
X β€” Context Engineering
Team Directives team-boot assembles constitution, CDR index, and PDR/ADR indexes into the system prompt at session start; team-discover provides manual re-scan
XI β€” Directives as Code
Team + LevelUp + Product + Architecture All directive lifecycles (CDR, PDR, ADR) live in version-controlled repos, each with draft β†’ clarify β†’ accept β†’ publish β†’ analyze stages
XII β€” Build to Delete
team-repair + evals-analyze --build-to-delete runs evals without directives via LLM calls; if model passes, proposes deletion (Harness Decay); evals-analyze routes spec failures to levelup-specify (rules) and generalization failures to the evaluator backlog β€” the feedback loop that makes build-to-delete verifiable

See RELEASE.md for the release runbook, tag naming conventions, and recovery procedures.

MIT β€” see LICENSE.

── more in #developer-tools 4 stories Β· sorted by recency
── more on @tikal 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/agent-skills-that-br…] indexed:0 read:16min 2026-08-04 Β· β€”