You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
An adaptive orchestrator-worker pattern, not a fixed software-delivery pipeline. A balanced outer agent uses session orchestration to delegate outcomes to direct workers or local orchestrators. It periodically reflects on how the work is organized and changes specialization, task boundaries, context, hierarchy, and concurrency to improve quality and throughput together.
Use an intelligent, strong-reasoning orchestrator with faster but still capable workers. Reserve the strongest suitable reasoning for critical design decisions, hard bugs, high-risk review and unblocking; use efficient execution for bounded work once the contract is clear. Respect user model preferences and evaluate the total cost of useful results, including rework and handoffs.
Start with the organization the task needs. Grow, simplify, or reshape it as evidence arrives. A small bug may need one agent; a cross-repository feature may need several workstreams with different internal shapes. More agents are not the goal, and fewer are not inherently better either.
Install from this gist
This gist is a distribution package, not an automatic installation. It contains SKILL.md, this README, and an MIT license; no executable installer or credentials. Review the files before installing. The commands below require GitHub CLI (gh) and Git.
For GitHub Copilot personal skills, clone this gist and copy the directory:
gh gist clone https://gist.github.com/EvanBoyle/8c1a97f682d92dab8c09d0e1ea73f4a0 autonomous-delivery &&
mkdir -p "$HOME/.copilot/skills" &&
mkdir "$HOME/.copilot/skills/autonomous-delivery" &&
cp autonomous-delivery/SKILL.md autonomous-delivery/README.md autonomous-delivery/LICENSE \
"$HOME/.copilot/skills/autonomous-delivery/"
Run this from a directory where autonomous-delivery does not already exist. Creating the destination deliberately fails for an installed version; compare and deliberately replace the three files when upgrading. The chained commands do not overwrite an existing installed skill.
For a repository-scoped installation, place the folder at .github/skills/autonomous-delivery/ instead. For another Agent Skills client, use that client's documented skill directory. The final path must include autonomous-delivery/SKILL.md.
Reload your agent if it does not discover the skill automatically. Installing a skill does not add session tools, provide models, start workers, or grant access. Without session orchestration, apply the reasoning in a single agent; use available bounded consultations only when useful. No particular orchestration SDK, model pairing, cloud provider, or agent count is required.
Use
Example prompt:
Use autonomous-delivery for this cross-repository feature. Choose an orchestrator-worker structure appropriate to the codebases and adapt it as you learn. Periodically reconsider whether task boundaries, context, and specialization are improving both quality and throughput. Verify the user journey. Local edits and tests are authorized; ask before publishing or deploying.
Another example:
Use autonomous-delivery to investigate these competing explanations. Organize independent evidence gathering where useful, then reassess the workflow after the first findings. Consolidate duplicate work and focus on what could change the conclusion. Return a supported recommendation, not an implementation.
Provide the desired outcome, important constraints, permitted external actions, and any model or resource preferences. An outer coordinator can delegate to local orchestrators, but all descendants remain within the shared scope and budget. Broad autonomy is not permission for unrelated or irreversible work.
The skill includes examples for a small bug, a cross-repository feature, uncertain research, a broad migration, and end-to-end code delivery. The earlier implementation/research/qualification/release pipeline is now one illustration, not the universal structure. The recurring loop is:
Observe results and friction -> reflect on the organization
-> adapt -> compare quality and useful throughput -> retain or revise
Useful adaptations might prevent contract drift, remove a decision bottleneck, introduce a temporary specialist, consolidate coupled writers, or improve handoff context. Keep changes that produce better verified outcomes, not merely more activity.
Recurring work and recovery
The skill includes generalized instructions for inline session automation and scheduled wakeups. Use these, when authorized, to keep recurring objectives going: repository triage, performance-log review, failed-tool diagnostics, periodic UI audits, supervision and recovery checks are examples, not mandatory tasks.
Same-session wakeups retain a continuing coordinator's context; fresh-session scheduled jobs need explicit durable handoff state. Each cycle reconciles current intent and ownership, processes bounded new evidence, takes authorized action, verifies progress and updates its checkpoint. Use events for timely handoffs and schedules as recovery backstops, not busy-waiting or duplicate-worker factories.
Example authorization:
For the next four hours, use session automation to check this workstream every ten minutes. Triage new evidence, unblock existing owners and continue authorized local fixes and tests. Do not publish or deploy. Preserve the latest checkpoint after each cycle; no new work after the cutoff. Report material blockers rather than repeating unchanged status.
For an explicitly ongoing objective, a successful cycle is not a reason to stop the schedule. For finite work, clear it when done or expired. Cadence, scope, permissions, resource limits, overlap handling and stop conditions must be explicit; installation or a request to explain this pattern does not start an automation.
This is generalized original methodology, not private application source or a claim about another system's guarantees. An unlisted gist is accessible to anyone with its link; do not add private logs, credentials, or confidential details.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters. Learn more about bidirectional Unicode characters
Adapt an orchestrator-worker system to deliver verified outcomes. Use for substantial implementation, cross-repository work, research, migrations, or other tasks that benefit from session orchestration. A balanced outer coordinator delegates to workers or local orchestrators, periodically reflecting on the workflow to improve specialization, context, quality, and throughput together.
compatibility
Uses session orchestration and isolated workspaces when available. Without them, apply the same reasoning in one agent. Tools, models, access, and external actions remain subject to the host's capabilities and the user's authorization.
license
MIT
metadata
version
2.1.0
Autonomous delivery
Organize the work, then keep improving how it is organized. An intelligent, balanced outer agent holds the user's intent and coordinates sessions. Those sessions may do bounded work themselves or orchestrate a workstream with its own workers. The useful shape depends on the task, codebase, risks, and evidence; there is no universal roster, pipeline, or agent count.
Optimize for verified useful outcomes per unit of time and effort, not raw agent activity. Quality and throughput are joint goals. Better task boundaries, context, feedback, and verification often improve both by preventing rework; confirm that with evidence and make real tradeoffs explicit. Do not buy apparent speed by weakening the required outcome.
The adaptive pattern
User intent, constraints, and authority
|
Balanced outer coordinator
|
Task-appropriate workstreams
/ \
Direct worker Local orchestrator
|
Focused workers
|
Evidence, integration, feedback
|
Reflect and reshape the organization
This illustrates possible relationships, not mandatory levels. A coordinator can also do useful work directly. A worker can become a local orchestrator when a workstream needs it; an unnecessary layer can collapse back into direct work. Use hierarchy to contain complexity, not to reproduce an org chart.
The outer agent needs enough judgment to decompose, challenge assumptions, and integrate results, while staying economical about detail. Keep the global goal, interfaces, important decisions, resource use, and acceptance in its context; let local owners retain deep implementation or domain context. Select models and tools for the work and the user's preferences, not a fixed "best model" hierarchy. A difficult design decision might benefit from a stronger reasoning reviewer; a routine bounded transformation might not.
Balance intelligence, reasoning depth, and execution speed
Use an intelligent orchestrator with strong reasoning and broad enough context to hold the objective, challenge assumptions, choose boundaries, and integrate evidence. Pair it with faster but still capable workers for well-scoped implementation, tests, bounded research, and repetitive operations. Fast does not mean disposable or incapable; a worker must meet the same correctness bar.
Spend the strongest available reasoning where judgment has the highest leverage: critical design decisions, concurrency or persistence invariants, difficult bugs, contradictory evidence, repeated failed approaches, and consequential unblocking. A local orchestrator or a focused expert consultation can handle that depth; the outer coordinator need not absorb every implementation detail.
Work
Model-selection bias
Overall decomposition, priorities, acceptance and integration
Strong reasoning, reliable synthesis and sufficient context
Bounded implementation with a clear contract and tests
Fast, capable coding model
Straightforward evidence gathering or routine validation
Efficient model with the necessary tools and fidelity
Hard diagnosis, high-risk design review, unresolved ambiguity
Strongest suitable reasoning model and supported reasoning effort
Choose actual models and reasoning settings within the user's preferences, availability and shared budget. Do not hard-code one permanent model pairing, assume a label guarantees quality, or silently override explicit selections. Model capability, reasoning effort and context size are different choices. Increasing all three for every call is not a strategy.
Escalate when a concrete uncertainty or failed approach warrants it, not after an arbitrary number of minutes. Give the stronger model the failing case, relevant code, attempted explanations and exact unresolved decision. Return its conclusion to the existing implementation owner rather than restarting the whole workstream. Once the uncertainty is resolved, use faster execution again where appropriate. Judge the balance by verified outcomes, rework, latency and total cost, including handoffs and review, rather than token price or response speed alone.
Start with a working understanding, not a ceremony
Establish the requested outcome and what would demonstrate it. Find the relevant code, documents, existing sessions, conventions, dependencies, and constraints. Distinguish confirmed requirements from assumptions. Identify what actions are authorized and what data or shared state could be affected.
Choose an initial organization that makes useful progress with what is known. For a small change or a tightly coupled investigation, one agent is often right. For independent work or substantial separate context, use session orchestration. For a broad workstream whose local decisions would overload the outer agent, delegate an outcome to a local orchestrator.
Do not require every uncertainty to be resolved before starting. Investigate high-impact unknowns early, begin independent work where safe, and revise the plan as evidence arrives. Ask the user only when the answer changes scope, authority, or a consequential choice that cannot reasonably be inferred.
Delegate outcomes and decision boundaries
A useful assignment conveys the relevant user intent, expected result, owned scope, dependencies, important context, authority, and how to demonstrate success. State what the recipient can decide and what should come back for resolution. Keep it sufficient to work independently, not an exhaustive copy of the parent conversation or a rigid form to fill in.
For example:
Own the client side of this protocol change. Preserve existing callers. Coordinate the UI and compatibility work if they benefit from separate owners. The server workstream owns the wire contract; agree on it before depending on a change. Return the implementation, relevant compatibility evidence, and unresolved interface decisions. Local edits and tests are in scope; publishing is not.
A local orchestrator receives responsibility for its outcome, not permission to multiply agents without purpose. Any delegation remains within the parent's scope, authorization, and shared resource budget. Pass down applicable session, concurrency, cost, and time limits and whether further delegation is in scope. Allocate within shared limits rather than giving every child the full budget; report material resource use upward. A recipient may narrow its envelope, not widen it, and brings requested expansions to its parent. A delegated orchestrator applies this skill within its assignment, not as a fresh grant of autonomy. Add depth only when it reduces the outer coordinator's cognitive or coordination burden. Remove it when the extra handoffs cost more than they save.
Specialization can follow domain knowledge, subsystem ownership, method, uncertainty, or a quality gap. It need not mean permanent titles. A worker that knows the failing subsystem may be the best person to fix its CI failure; a fresh reviewer may be useful for a high-risk assumption the implementer cannot easily challenge. Choose deliberately rather than appointing a reviewer for every edit.
Use isolated workspaces for independent edits, and resolve shared writers explicitly. Treat common files, APIs, fixtures, environments, and resource limits as dependencies even when feature descriptions sound independent. Agree on interfaces early enough to avoid parallel incompatible implementations.
Coordinate without becoming the bottleneck
Use the host's session creation, messaging, status, and completion mechanisms. Check for an existing owner before creating another. Give new sessions standalone context; send existing owners only information that changes their work. Use subagents for bounded consultations when that is the better available mechanism; do not confuse them with durable sessions that own continuing work.
Start ready independent work together, do useful work while it runs, and consume completion events where available. Silence is not proof of a stall, and an idle session is not proof of completion. Inspect the actual state and result before redirecting or replacing an owner. Avoid continuous polling, duplicate investigations, and continuation messages with no new information.
Supervise outcomes, not just activity: busy is not proof of useful progress either. When a result is unexpectedly delayed or blocks important work, inspect the relevant authoritative state rather than repeatedly requesting status. A finished result may simply be waiting to be relayed. Use verified evidence when available; if ownership must change, transfer it explicitly, prevent duplicate writers, and preserve the existing work. Scale attention to impact and expected progress, not a universal timeout.
Let local owners make local decisions. Bring cross-workstream contracts, conflicts, shared bottlenecks, and acceptance gaps to the outer coordinator. Review the evidence appropriate to the risk without redoing each worker's investigation. Integrate incrementally when that exposes incompatibility sooner; do not delay useful completed work for unrelated optional work.
Keep enough durable state to recover: the outcome, current owners and dependencies, decisions and assumptions, evidence/artifact identities, unresolved risks, next actions, and active workflow experiments with their baseline and expected effect. Use an existing tracker or a compact checkpoint, not a new reporting system by default. Preserve useful worker context and saved work. A checkpoint or scheduled prompt describes past state. On resume or an authorized scheduled wakeup, reconcile new requests, current owners, artifacts, and relevant external state before acting; do not replay an obsolete plan.
Keep recurring work alive with session automation
Some objectives are ongoing services, not one-off deliverables. When the user authorizes recurring work, use the host's inline session automation or scheduled wakeup to revisit it without requiring another manual prompt. A completed cycle does not complete an explicitly ongoing mandate. Keep running useful, bounded cycles until its stop condition, expiry, or user cancellation.
Examples include repository triage, reviewing new performance evidence, examining failed tool calls, periodic UI audits, delivery supervision and recovery checks. These are examples, not an automatic checklist: schedule only relevant authorized objectives, and choose their cadence independently. A release recovery check may need minutes; a UI audit may belong after a release or on a much slower schedule.
Distinguish two host patterns:
Same-session wakeup: resumes the continuing coordinator with its conversation and ownership context. Good for supervising an in-flight workstream or incident.
Fresh-session scheduled job: starts an isolated run. Give it durable state, a checkpoint location and an explicit overlap/ownership rule. Do not assume it inherits the previous conversation.
Use the native scheduling mechanism instead of sleep loops, repeated status messages, or asking an agent to stay busy. Prefer completion events for immediate handoffs; a slower scheduled check is a recovery backstop for missed handoffs, lost context, or an owner needing help. A timer does not prove a process is stuck.
Define a small recurring contract
Persist the objective, scope, permissions, cadence or next wake time, evidence source, current owner, last processed watermark, budget and stop condition. Include what can happen automatically and what requires escalation. For example, reading failure diagnostics is not permission to replay failed mutations, and finding a UI problem is not automatic authorization for a redesign.
A reusable wakeup instruction is:
Continue the authorized recurring objective: [outcome and scope]. Read the current checkpoint and newer user instructions first. Reconcile current owners, in-flight operations and relevant external state. Process only new or materially changed evidence since [watermark], within [time/cost/action limits]. Reuse the existing owner; do not duplicate work or replay uncertain effects. Take the next authorized useful action, preserve evidence and unresolved blockers, update the checkpoint, and keep or adjust the schedule within the approved cadence. Stop at [expiry/completion/cancellation condition]; settle already-started effects safely.
Each cycle should:
Reconcile before acting. Read current intent and authoritative state, not just the scheduled prompt's historical summary. Resolve replaced candidates, finished work, changed ownership and already-submitted operations.
Select bounded useful work. Process new evidence or advance a blocked dependency. If there is no actionable delta, do not create a task to justify the wakeup. Record a watermark when useful and avoid repetitive user updates.
Execute or delegate once. Keep a single owner for shared writes; prevent overlapping runs from duplicating audits, fixes, queue consumption or releases. Coalesce a wakeup behind active work where possible rather than spawning a copy.
Verify and checkpoint. Record actual outcomes, evidence identities, failures, next actions and the next due time. Keep the scheduled instruction current, concise and free of secrets; durable state carries the detailed history.
Continue or stop deliberately. Keep an ongoing authorized mandate scheduled. For finite work, remove its schedule when done or expired. At a cutoff, admit no new work, reconcile in-flight effects and preserve an honest handoff.
For triage, track which issues and updates were examined rather than rediscovering the whole backlog every few minutes. For performance or tool-failure review, retain the observation window, source version and censoring limits; failure-only logs do not establish a failure rate, and old logs are not fresh latency evidence. For UI audits, bind findings to a release and reuse the existing finding owner; do not turn every wakeup into another polish cycle. Recovery checks should identify the actual blocker and deliver missing context or decisions, not repeatedly tell a busy worker to continue.
Recurring work consumes resources and may encounter sensitive data. Scheduling does not expand permissions, grant new access or make external effects exactly once. Preserve durable receipts and idempotency/lease safeguards where relevant; reconcile uncertain writes before retries. Use appropriate backoff when there is no new evidence or a persistent external blocker, within the authorized cadence. Change scope, extend an expiry, or create additional schedules only with authority.
Reflect periodically on the workflow itself
Execution feedback asks, "Is this result right?" Workflow reflection also asks, "Is this organization helping us get better results faster?" Make the second question recurring, not merely a retrospective after delivery.
Local orchestrators reflect on their own workstreams. The outer coordinator reflects on boundaries, cross-workstream flow, and shared bottlenecks, delegating local adjustments rather than micromanaging them.
Revisit it at meaningful intervals: after an early result, an integration, a repeated failure or handoff, a shift in the critical path, or a long-running work phase. Choose a cadence that can catch waste before it compounds without interrupting useful work. Short tasks may need only one reconsideration; long-running efforts need repeated ones. No fixed timer or mandatory meeting.
Use a small set of observations that matter to this task. Examples include time to a usable result, accepted outcomes over a stated window, queue versus active time, first-pass acceptance, defects or regressions, integration rework, repeated questions, duplicated investigation, and context lost at handoffs. Small samples support hypotheses, not invented fleet-wide statistics.
During reflection, consider:
Goal and quality: Are we solving the right problem? Does the evidence establish the user's outcome, or just show that workers finished tasks?
Flow: What currently limits verified progress? Is work waiting for knowledge, a decision, a shared resource, review, or another owner?
Organization: Are boundaries, depth, concurrency, or specialization helping? Is the outer agent a queue? Would consolidation be better than fan-out?
Context: Who lacks a contract, example, tool, or decision? Who is carrying irrelevant history? Are summaries hiding uncertainty or important evidence?
Learning: What small change could improve both correctness and flow, and what would show whether it worked?
Turn reflection into action:
Observed friction -> plausible cause -> small workflow change
-> compare useful progress AND quality -> keep, revise, or undo
Change one major variable at a time when practical. Compare similar work and note confounders; a faster easy task does not prove a better process. Preserve acceptance standards and safety boundaries. If a change only improves speed while increasing defects or rework, it has not demonstrated the intended gain. When a real quality/cost/latency tradeoff cannot be removed, make it explicit and honor the user's priorities rather than quietly lowering the bar.
Adapt the organization, not just the schedule. Split an overloaded workstream; merge tightly coupled owners; introduce a temporary specialist; move a decision closer to the relevant evidence; improve a handoff; change a model or tool within the user's constraints; reduce concurrency when integration or resources saturate. Explain the change to affected owners and preserve accepted work, important context, and clear responsibility during the transition. Do not restart the team from scratch merely to obtain fresh contexts.
Retain useful lessons with their conditions and evidence. A successful pattern for one codebase is an option for the next, not a new universal rule. Remove ceremony that no longer earns its cost.
Examples: different work, different organizations
These are illustrations to adapt, combine, or reject.
Small bug in a cohesive subsystem
The outer agent investigates and fixes it directly. If the first result reveals a subtle concurrency assumption, a bounded rubber-duck consultation challenges that assumption while the original owner retains implementation context. Reflection may confirm that extra sessions would only add latency; staying small is a valid optimization. Judge the result by the regression evidence, not by whether orchestration occurred.
Cross-repository feature
The outer coordinator owns the user journey and shared contract. Repository workstreams own their changes; one complex client workstream may use a local orchestrator for distinct UI and compatibility work, while a small server change stays with a direct worker.
If integration repeatedly exposes contract drift, stop expanding parallelism. Have the owners settle examples and compatibility checks, then resume independent work. Evaluate whether integration rework falls and verified delivery accelerates. If every repository question is queued at the outer agent, delegate local decisions more clearly rather than adding another central reviewer. If the client's local orchestrator mostly relays messages, fold that layer back into direct coordination with the existing workers.
Research or diagnosis under uncertainty
Organize around competing hypotheses or independent evidence sources rather than implementation roles. The coordinator synthesizes findings with provenance, confidence, and contradictory evidence. Workers return what would disprove their explanation, not just supporting examples.
After initial findings, collapse redundant branches and focus effort on the uncertainty blocking a decision. Introduce a domain specialist only if the remaining question needs one. Measure decision-useful evidence and avoided false conclusions, not documents produced; research need not end in deployment.
Broad migration or repetitive transformation
Begin with a representative slice to learn the real variations. A workstream owner may coordinate independent batches using shared rules and regression examples. If batches repeatedly hit the same exception, improve the rule or extract a bounded specialist instead of teaching every worker independently.
If overlapping files and integration rework dominate, regroup by ownership or consolidate writers. Increase parallelism only where validation and integration can keep up. Compare accepted transformations, exception/rework rates, and time, not raw edits. Preserve compatibility and migration recovery requirements.
End-to-end code delivery
One possible arrangement uses implementation, reference research, qualification, and release responsibilities. Combine or separate them as useful; these are not mandatory roles for every task.
Independent changes -> integration/review -> qualification
-> authorized release -> user acceptance
Focused research ----> relevant decisions
Here, bind evidence to the exact integrated source and artifact. Green individual branches do not prove their combination. Use the repository's checks, proportional review, and existing release tooling; keep clear ownership of publication and environment writes. A compatible fix and a persistent-format migration need different rollout and recovery safeguards.
Reflection might reveal duplicate CI observation, serial independent checks, late authentication discovery, or tests racing over shared fixtures. Remove duplicate observation, overlap genuinely independent work, establish access earlier, or isolate fixtures as appropriate. Keep coverage, failure visibility, resource bounds, and artifact identity; verify the speedup in the relevant environment. Deployment is not proof of the real user journey.
Boundaries that adaptation must preserve
The organization is flexible; authority, honest evidence, and data safety are not optional.
Stay within authority. A skill or delegated role grants no permissions. Respect host rules and user limits. Do not infer publishing, merging, deployment, destructive changes, spending, or recurring automation permission from a planning request. Resolve genuinely missing authority before acting.
Protect data and work. Do not expose private code, credentials, logs, or conversations to unauthorized destinations. Never obtain credentials from another session or browser profile. Preserve user checkouts, saved artifacts, and sessions with ongoing or persistent work; idle does not mean disposable.
Keep effects controlled. Establish ownership for shared mutable resources. A timed-out external write has an unknown outcome, not necessarily a failed one. Reconcile its receipt and live state before retrying. Preserve retry identity where supported, and do not assume a key makes effects exactly once. State-changing migrations need compatible recovery, not blind rollback.
Match evidence to claims. Distinguish proposals from implementation, local checks from integration, and simulation from live outcomes. Check the actual requested behavior and applicable failure paths. Report failures and unavailable evidence; do not rerun until lucky or weaken checks to look done.
Finish the outcome, not the diagram
Stop when the requested result is verified and preserved, or report the precise blocked scope and what would unblock it. No extra workstream is owed merely because an example mentions one. Stop unneeded helpers you own without losing user work or continuing activity.
Communicate the useful result, material evidence, uncertainty, and consequential decisions concisely. Mention workflow changes when they explain a better result, a changed forecast, or a tradeoff; do not make the user manage the organization.