Showing version 4 bot legacy · api · 2026-08-10T10:28:06Z

Spec-Driven Workflow

Spec-Driven Workflow

The pipeline that turns one ambiguous request into a chain of small, one-shottable steps.

The pipeline

flowchart TD
    A[Idea] --> B{Small?}
    B -- yes --> Q["/speckit.agentic.quickfix/"]
    B -- no --> S["/speckit.specify/"]
    S --> CL["/speckit.clarify/"]
    CL --> P["/speckit.plan/"]
    P --> R[Pre-Task Review]
    R --> RS["/speckit.agentic.restraint/"]
    RS --> T["/speckit.tasks/"]
    T --> SEC["/speckit.agentic.security/"]
    SEC --> I["/speckit.implement/"]
    I --> V["/speckit.agentic.verify/"]
    V -- BLOCKER --> I
    V -- PASS --> RE["/speckit.agentic.retrospective/"]
    RE --> EX["/speckit.agentic.skills/"]
    EX --> KF{3+ repeats?}
    KF -- yes --> CON["/speckit.constitution/"]
    KF -- no --> END[Done]
    Q --> V

Each command is a separate invocation, its own model tier, its own focused context.

Commands without the agentic. segment are core Spec Kit. The ones carrying it come from the agentic extension; plan, tasks, and implement are core commands the agentic preset replaces. See Installing and Running the Toolkit.

Phases

Phase Command Model Purpose
quickfix /speckit.agentic.quickfix sonnet ≤2 files, ≤20 lines. Bypass full pipeline.
specify /speckit.specify opus Natural language → structured spec
clarify /speckit.clarify opus ≤5 questions, answers folded back into spec
plan /speckit.plan opus Architecture, sketches, data model, contracts; gated review
restraint /speckit.agentic.restraint opus Decision ladder against the plan before tasks exist
tasks /speckit.tasks opus Decompose into dependency-ordered tasks with tiers, context, waves
security /speckit.agentic.security fable OWASP Top 10:2025 + baseline controls appended to tasks.md as SEC-* tasks. Falls back to opus on a cyber refusal — see Model Routing: Haiku, Sonnet, Opus
implement /speckit.implement opus Orchestrate; dispatch each task to its @model tier
verify /speckit.agentic.verify opus Adversarial review across 7 dimensions
retrospective /speckit.agentic.retrospective sonnet Post-mortem → constitution amendments + skill candidates
skills /speckit.agentic.skills sonnet Materialise skill files

Two more sit outside the main line: /speckit.agentic.audit sweeps the whole repo for over-engineering, /speckit.agentic.debt harvests deliberate-shortcut markers into a ledger, and /speckit.analyze cross-checks spec, plan, and tasks for consistency.

Tier rationale: Model Routing: Haiku, Sonnet, Opus.

Phase-by-phase

Specify

Input: natural-language description.
Output: spec.md — one structured document of WHAT, never HOW.

Sections enforced by the template: User Scenarios (priority-ordered, independently testable), Functional Requirements (FR-001, FR-002…), Success Criteria (measurable, technology-agnostic), Edge Cases, Assumptions, optional Key Entities.

Hard rules:

  • No tech-stack references. No file paths. No framework names. The spec must read as something a business stakeholder could approve.
  • Up to 3 [NEEDS CLARIFICATION: …] markers for genuine ambiguities. Anything else gets a reasonable default and is logged under Assumptions.
  • Every requirement must be testable. "User-friendly" is rejected. "95% of users complete checkout in under 3 minutes" is accepted.

Example FR line:

- **FR-005**: System MUST hard-delete artifact rows when their source file
  is removed from disk, including the corresponding full-text-search entry.

Done when: every user story has acceptance scenarios, every FR is verifiable in isolation, success criteria are numbers (or qualitative-but-falsifiable statements).

Clarify

Input: spec.md.
Output: updated spec.md with a ## Clarifications section dated by session.

Up to 5 questions. One at a time. Each presents a recommended answer with reasoning, plus 2–3 alternatives in a table, plus a free-form "Short" option. The user picks; the spec is updated in place immediately (FR text changed, assumptions added, edge cases extended).

Example interaction:

Q3 of ≤5 — Auth posture for AE read routes

Recommended: Option B — gate AE read routes behind WIKI_TOKEN.
AE content is review material for in-flight features; defaulting to
public on a production domain leaks roadmap.

| Option | Description |
|--------|-------------|
| A      | Public read (same as existing wiki articles) |
| B      | Token-gated using existing WIKI_TOKEN |
| C      | AE mode auto-disabled unless server is localhost |

The five-question cap is real. The clarifier ranks candidates by impact × uncertainty and asks only the top of the queue. Low-impact questions are silently dropped — better to make a reasonable default than burn a question slot.

Done when: zero [NEEDS CLARIFICATION] markers remain, the checklist file shows all items checked.

Plan

Input: spec.md plus the project constitution.
Outputs (every run):

File Purpose
plan.md Tech context, structure decision, constitution check, complexity-tracking table for any rule violations
research.md Numbered decisions with rationale + alternatives considered
data-model.md Entities, fields, validation, state transitions, schema additions
contracts/*.md HTTP routes, storage interfaces, CLI grammars — whatever the project exposes
quickstart.md Runnable end-to-end validation scenarios
preview.md Architecture mermaid + 2–4 code sketches with real signatures, real types, real SQL — not pseudocode

The user reviews preview.md and approves before tasks generation. This is the cheapest place to catch wrong-design errors — at the sketch, not at the diff.

Example research entry:

## 4. Storage namespacing

**Decision**: New table `ae_artifacts` distinct from `articles`. Slug stays
globally non-colliding by namespacing at the HTTP route boundary.

**Rationale**: Spec FR-013 requires AE artifacts can't collide with wiki
articles. Cleanest enforcement is not putting them in the same table.
Hard-delete-on-disappear becomes a single `DELETE WHERE feature_slug=?
AND rel_path=?`.

**Alternatives considered**:
- Shared table + origin enum: rejected — every existing query needs an
  origin filter, raising leak risk.
- Separate SQLite file: rejected — operational surface doubles.

Constitution check happens twice: before research starts, again after design is done. Violations require either a fix or a documented justification in Complexity Tracking.

Done when: user approves preview.md. Until then, no tasks are generated.

Restraint

Input: the approved plan.
Output: a list of things in it that shouldn't be built.

Fires on the after_plan hook, so it's offered every run without being forced. It walks the decision ladder over each construct the plan proposes — does this need to exist, does the codebase already have it, does the standard library, the platform, an installed dependency, a one-liner — and stops at the first rung that answers the need.

The gate sits here rather than at verify because this is the last point where nothing has been written yet. Catching an unnecessary package at the plan costs a sentence; catching it at the diff costs the implementation plus the deletion.

Findings are tagged delete / stdlib / native / yagni / shrink. Nothing to cut returns Lean already. Ship.

Done when: the plan's constructs each survive the ladder, or the plan is revised. Full detail in The Restraint Principle: YAGNI for Agents.

Tasks

Input: everything from plan.
Output: tasks.md — dependency-ordered, traceable to user stories.

Each task line in exactly one shape:

- [ ] T012 [P] [deps: T003,T007] [US1] @sonnet [ctx: src/foo.ts:10-40, src/types.ts:1-30] Implement Foo.Bar in src/foo.ts
  • T012 — sequential ID, also the natural dependency-graph node name.
  • [P] — present only if parallelisable (different files, no incomplete deps). Absent = sequential.
  • [deps: T003,T007] — explicit deps. A task only starts when all its deps are [X].
  • [US1] — links the task back to a user story for traceability (omitted for Setup/Foundational/Polish).
  • @sonnet — suggested model tier, one of @haiku/@sonnet/@opus/@fable (see Model Routing: Haiku, Sonnet, Opus). SEC-* tasks carry @fable; re-run them at @opus if they come back refused.
  • [ctx: file:Lstart-Lend, ...] — pre-fetched line ranges. The orchestrator inlines these into the subagent's prompt under ## Pre-Fetched Context (see Subagents and Context Injection).

Phase grouping:

Phase 1 — Setup            (project init, deps, scaffolding)
Phase 2 — Foundational     (shared models, schema, blocking work)
Phase 3+ — User Stories    (one phase per story, P1 first)
Final  — Polish            (cross-cutting concerns, docs, performance)

Concurrent tasks live under ### Wave N (parallel) headers. The orchestrator commits between waves so the next wave reads fresh state.

### Wave 1 (parallel, after T002)
- [ ] T007 [P] @sonnet  Implement Importer types in internal/ae/ae.go
- [ ] T010 [P] [deps: T003,T004,T005] @opus  Implement Storage AE methods

### Wave 2 (sequential, deps: T007,T010)
- [ ] T011 @opus  Wire Importer.ScanFeature

Done when: every user story has all the tasks needed to satisfy its acceptance scenarios, deps form a valid DAG, every non-trivial task has a [ctx: ...] annotation.

Implement

Input: tasks.md.
Output: code, marked-off tasks, commits between waves.

The orchestrator walks tasks.md in dependency order. For each task:

  1. Parse the [ctx: ...] annotation.
  2. Read each file/range from disk.
  3. Build the subagent prompt: task description + ## Pre-Fetched Context block with file contents inlined + ## Allowed Operations list (paths the subagent may touch).
  4. Dispatch at the suggested @model tier.
  5. Subagent executes. Returns either a diff (success) or CONTEXT_INSUFFICIENT: <reason> (abort).
  6. On success: mark [X]. On abort: widen [ctx] once, retry. Second abort → escalate to the user.
  7. At the end of each ### Wave N group: commit.

The subagent is instructed: do not grep, glob, or read files outside the listed paths. The leash is the whole point — without it the subagent re-explores the codebase and burns 30–70% of its context on lookups the orchestrator already did.

Example dispatched prompt (abridged):

## Task
T010 [@opus] Implement all 10 AE methods in internal/storage/ae.go

## Pre-Fetched Context
### internal/storage/sqlite.go:97-168
(actual 70 lines inlined here)

### internal/storage/storage.go:1-22
(actual 22 lines inlined here)

## Allowed Operations
Read/Edit/Write only on:
- internal/storage/ae.go

If you need anything else, ABORT with CONTEXT_INSUFFICIENT.

Done when: every task is [X], every wave has been committed, the diff matches the plan.

Verify

Input: spec.md + plan.md + tasks.md + the git diff produced by implement.
Output: a structured review with a verdict.

Seven dimensions, in order of typical catch-rate:

  1. Spec divergence — for each FR and acceptance scenario, is it actually implemented? Most common BLOCKER source.
  2. Missing tasks — any task marked [X] whose diff content can't be found?
  3. Security — reviewed against OWASP Top 10:2025, not the 2021 list. A03 (software supply chain) and A10 (mishandling of exceptional conditions) are new categories and the two most often missed.
  4. Correctness — off-by-one, races, edge cases the spec called out but the code skipped, swallowed errors.
  5. Contract breaks — schema, struct, signature, response-shape changes that would break callers.
  6. Necessity — abstraction with one caller, a dependency replacing three lines, config nobody asked for. Code that works and shouldn't exist is still a defect. See The Restraint Principle: YAGNI for Agents.
  7. Velocity illusionsNOTE-level: tests that pass without touching the edges, PR text that doesn't match the diff, behaviour nobody can explain. Rarely blocks; predicts the next three bugs.

Two protocols wrap those dimensions. The hunt loops until dry — keep going until two consecutive rounds surface nothing new, deduplicating against everything already seen including refuted findings, because a fixed target either stops early or pads to reach itself. And every finding above NOTE faces an adversarial pass first: try to refute it, default to refuted when uncertain, and keep it only if it survives. Plausible-but-wrong findings cost more than missed ones, because they get acted on.

Findings tagged:

Severity Effect
BLOCKER Must fix. Verdict = FAIL. Re-enter implement with the finding as the new task.
WARN Real but non-blocking. Verdict still PASS if no BLOCKERs. Tracked.
NOTE Observation. Informational.

The reviewer's prompt explicitly tells it to look for what's missing, not validate what's present. Same model + same diff with a different framing produces a measurably different catch rate. Detail in Trust but Verify.

Example finding:

### Spec divergence — FR-005 (hard delete on file removal)
**[BLOCKER]** Spec FR-005 requires hard-delete of artifact rows when a
feature folder disappears. The diff implements file-level deletion in
ScanFeature, but ScanAll never calls DeleteAEFeature for slugs that have
vanished. Quote: importer.go:50-60.

Done when: BLOCKER count is zero. Anything less re-enters implement.

Retrospective + Skills

Input: every artefact produced this run, plus the verify report.
Outputs: a post-mortem with two structured blocks parsed by downstream commands.

CONSTITUTION_AMENDMENTS — rules that, if adopted, would prevent the failures seen this run:

CONSTITUTION_AMENDMENTS:
  - id: cache-stateless-builders
    rationale: |
      Chroma formatter was constructed per HTTP request, costing ~3ms each.
    rule: "Stateless objects with non-trivial construction cost should be
            cached at handler init, not constructed per request."
    severity: warn

These don't go straight into the constitution — they queue in known-failures.md with a count. Same ID seen 3+ times across runs → promotion happens automatically.

SKILLS_TO_EXTRACT — reusable patterns:

SKILLS_TO_EXTRACT:
  - name: sqlite-fts5-pattern
    rationale: Canonical FTS5 setup with content/content_rowid + the three
      triggers, used twice now (articles_fts, ae_artifacts_fts).
    template_path: skills/extracted/sqlite-fts5-pattern.md

/speckit.agentic.skills materialises each entry into a file under skills/extracted/, with trigger, applies-when preconditions, template, and references. Candidates lacking a concrete trigger or concrete steps are skipped and reported rather than written — vague advice is not a skill. Under Claude Code these become loadable skills, so future runs pick up the relevant ones without anyone pasting them back in.

Without this phase, every run is independent. With it, marginal cost trends down and marginal quality trends up. Detail in The Compounding Layer.

Done when: both YAML blocks parse cleanly and known-failures.md is updated.

Fast-track: /speckit.agentic.quickfix

Hard preconditions, enforced before any edit:

  • ≤ 2 files
  • ≤ 20 lines
  • No new files / deps / public-API changes
  • Commit type in {fix, docs, style, chore, typo, refactor}
  • No protected paths (configurable)

Trip any of these → escalate to full pipeline. Ceremony for a typo is worse than no process.

Why waterfall

Each phase produces an artefact that gates the next. Unfashionable, intentional. The whole structure exists to defeat The One-Shot Problem: small steps to keep each agent invocation one-shottable, written checkpoints between them.

Trade made explicit: heavy planning and review up front, almost mechanical implementation after. The thing being avoided is the ten-minute agent run that returns the wrong feature. Every hard decision is made and recorded before the implementer starts writing.

The alternative — "give the agent the repo, tell it to do the next thing" — is the one-shot trap with extra steps.