Spec-Driven Workflow
Spec-Driven Workflow
The pipeline that turns one ambiguous request into a chain of small, one-shottable steps.
The pipeline
flowchart TD
A[Idea] --> B{Small?}
B -- yes --> Q["/speckit.agentic.quickfix/"]
B -- no --> S["/speckit.specify/"]
S --> CL["/speckit.clarify/"]
CL --> P["/speckit.plan/"]
P --> R[Pre-Task Review]
R --> RS["/speckit.agentic.restraint/"]
RS --> T["/speckit.tasks/"]
T --> SEC["/speckit.agentic.security/"]
SEC --> I["/speckit.implement/"]
I --> V["/speckit.agentic.verify/"]
V -- BLOCKER --> I
V -- PASS --> RE["/speckit.agentic.retrospective/"]
RE --> EX["/speckit.agentic.skills/"]
EX --> KF{3+ repeats?}
KF -- yes --> CON["/speckit.constitution/"]
KF -- no --> END[Done]
Q --> V
Each command is a separate invocation, its own model tier, its own focused context.
Commands without the agentic. segment are core Spec Kit. The ones carrying it come from the agentic extension; plan, tasks, and implement are core commands the agentic preset replaces. See Installing and Running the Toolkit.
Phases
| Phase | Command | Model | Purpose |
|---|---|---|---|
| quickfix | /speckit.agentic.quickfix |
sonnet | ≤2 files, ≤20 lines. Bypass full pipeline. |
| specify | /speckit.specify |
opus | Natural language → structured spec |
| clarify | /speckit.clarify |
opus | ≤5 questions, answers folded back into spec |
| plan | /speckit.plan |
opus | Architecture, sketches, data model, contracts; gated review |
| restraint | /speckit.agentic.restraint |
opus | Decision ladder against the plan before tasks exist |
| tasks | /speckit.tasks |
opus | Decompose into dependency-ordered tasks with tiers, context, waves |
| security | /speckit.agentic.security |
opus | OWASP Top 10:2025 + baseline controls appended to tasks.md as SEC-* tasks |
| implement | /speckit.implement |
opus | Orchestrate; dispatch each task to its @model tier |
| verify | /speckit.agentic.verify |
opus | Adversarial review across 7 dimensions |
| retrospective | /speckit.agentic.retrospective |
sonnet | Post-mortem → constitution amendments + skill candidates |
| skills | /speckit.agentic.skills |
sonnet | Materialise skill files |
Two more sit outside the main line: /speckit.agentic.audit sweeps the whole repo for over-engineering, /speckit.agentic.debt harvests deliberate-shortcut markers into a ledger, and /speckit.analyze cross-checks spec, plan, and tasks for consistency.
Tier rationale: Model Routing: Haiku, Sonnet, Opus.
Phase-by-phase
Specify
Input: natural-language description.
Output: spec.md — one structured document of WHAT, never HOW.
Sections enforced by the template: User Scenarios (priority-ordered, independently testable), Functional Requirements (FR-001, FR-002…), Success Criteria (measurable, technology-agnostic), Edge Cases, Assumptions, optional Key Entities.
Hard rules:
- No tech-stack references. No file paths. No framework names. The spec must read as something a business stakeholder could approve.
- Up to 3
[NEEDS CLARIFICATION: …]markers for genuine ambiguities. Anything else gets a reasonable default and is logged under Assumptions. - Every requirement must be testable. "User-friendly" is rejected. "95% of users complete checkout in under 3 minutes" is accepted.
Example FR line:
- **FR-005**: System MUST hard-delete artifact rows when their source file
is removed from disk, including the corresponding full-text-search entry.
Done when: every user story has acceptance scenarios, every FR is verifiable in isolation, success criteria are numbers (or qualitative-but-falsifiable statements).
Clarify
Input: spec.md.
Output: updated spec.md with a ## Clarifications section dated by session.
Up to 5 questions. One at a time. Each presents a recommended answer with reasoning, plus 2–3 alternatives in a table, plus a free-form "Short" option. The user picks; the spec is updated in place immediately (FR text changed, assumptions added, edge cases extended).
Example interaction:
Q3 of ≤5 — Auth posture for AE read routes
Recommended: Option B — gate AE read routes behind WIKI_TOKEN.
AE content is review material for in-flight features; defaulting to
public on a production domain leaks roadmap.
| Option | Description |
|--------|-------------|
| A | Public read (same as existing wiki articles) |
| B | Token-gated using existing WIKI_TOKEN |
| C | AE mode auto-disabled unless server is localhost |
The five-question cap is real. The clarifier ranks candidates by impact × uncertainty and asks only the top of the queue. Low-impact questions are silently dropped — better to make a reasonable default than burn a question slot.
Done when: zero [NEEDS CLARIFICATION] markers remain, the checklist file shows all items checked.
Plan
Input: spec.md plus the project constitution.
Outputs (every run):
| File | Purpose |
|---|---|
plan.md |
Tech context, structure decision, constitution check, complexity-tracking table for any rule violations |
research.md |
Numbered decisions with rationale + alternatives considered |
data-model.md |
Entities, fields, validation, state transitions, schema additions |
contracts/*.md |
HTTP routes, storage interfaces, CLI grammars — whatever the project exposes |
quickstart.md |
Runnable end-to-end validation scenarios |
preview.md |
Architecture mermaid + 2–4 code sketches with real signatures, real types, real SQL — not pseudocode |
The user reviews preview.md and approves before tasks generation. This is the cheapest place to catch wrong-design errors — at the sketch, not at the diff.
Example research entry:
## 4. Storage namespacing
**Decision**: New table `ae_artifacts` distinct from `articles`. Slug stays
globally non-colliding by namespacing at the HTTP route boundary.
**Rationale**: Spec FR-013 requires AE artifacts can't collide with wiki
articles. Cleanest enforcement is not putting them in the same table.
Hard-delete-on-disappear becomes a single `DELETE WHERE feature_slug=?
AND rel_path=?`.
**Alternatives considered**:
- Shared table + origin enum: rejected — every existing query needs an
origin filter, raising leak risk.
- Separate SQLite file: rejected — operational surface doubles.
Constitution check happens twice: before research starts, again after design is done. Violations require either a fix or a documented justification in Complexity Tracking.
Done when: user approves preview.md. Until then, no tasks are generated.
Restraint
Input: the approved plan.
Output: a list of things in it that shouldn't be built.
Fires on the after_plan hook, so it's offered every run without being forced. It walks the decision ladder over each construct the plan proposes — does this need to exist, does the codebase already have it, does the standard library, the platform, an installed dependency, a one-liner — and stops at the first rung that answers the need.
The gate sits here rather than at verify because this is the last point where nothing has been written yet. Catching an unnecessary package at the plan costs a sentence; catching it at the diff costs the implementation plus the deletion.
Findings are tagged delete / stdlib / native / yagni / shrink. Nothing to cut returns Lean already. Ship.
Done when: the plan's constructs each survive the ladder, or the plan is revised. Full detail in The Restraint Principle: YAGNI for Agents.
Tasks
Input: everything from plan.
Output: tasks.md — dependency-ordered, traceable to user stories.
Each task line in exactly one shape:
- [ ] T012 [P] [deps: T003,T007] [US1] @sonnet [ctx: src/foo.ts:10-40, src/types.ts:1-30] Implement Foo.Bar in src/foo.ts
T012— sequential ID, also the natural dependency-graph node name.[P]— present only if parallelisable (different files, no incomplete deps). Absent = sequential.[deps: T003,T007]— explicit deps. A task only starts when all its deps are[X].[US1]— links the task back to a user story for traceability (omitted for Setup/Foundational/Polish).@sonnet— suggested model tier (see Model Routing: Haiku, Sonnet, Opus).[ctx: file:Lstart-Lend, ...]— pre-fetched line ranges. The orchestrator inlines these into the subagent's prompt under## Pre-Fetched Context(see Subagents and Context Injection).
Phase grouping:
Phase 1 — Setup (project init, deps, scaffolding)
Phase 2 — Foundational (shared models, schema, blocking work)
Phase 3+ — User Stories (one phase per story, P1 first)
Final — Polish (cross-cutting concerns, docs, performance)
Concurrent tasks live under ### Wave N (parallel) headers. The orchestrator commits between waves so the next wave reads fresh state.
### Wave 1 (parallel, after T002)
- [ ] T007 [P] @sonnet Implement Importer types in internal/ae/ae.go
- [ ] T010 [P] [deps: T003,T004,T005] @opus Implement Storage AE methods
### Wave 2 (sequential, deps: T007,T010)
- [ ] T011 @opus Wire Importer.ScanFeature
Done when: every user story has all the tasks needed to satisfy its acceptance scenarios, deps form a valid DAG, every non-trivial task has a [ctx: ...] annotation.
Implement
Input: tasks.md.
Output: code, marked-off tasks, commits between waves.
The orchestrator walks tasks.md in dependency order. For each task:
- Parse the
[ctx: ...]annotation. - Read each file/range from disk.
- Build the subagent prompt: task description +
## Pre-Fetched Contextblock with file contents inlined +## Allowed Operationslist (paths the subagent may touch). - Dispatch at the suggested
@modeltier. - Subagent executes. Returns either a diff (success) or
CONTEXT_INSUFFICIENT: <reason>(abort). - On success: mark
[X]. On abort: widen[ctx]once, retry. Second abort → escalate to the user. - At the end of each
### Wave Ngroup: commit.
The subagent is instructed: do not grep, glob, or read files outside the listed paths. The leash is the whole point — without it the subagent re-explores the codebase and burns 30–70% of its context on lookups the orchestrator already did.
Example dispatched prompt (abridged):
## Task
T010 [@opus] Implement all 10 AE methods in internal/storage/ae.go
## Pre-Fetched Context
### internal/storage/sqlite.go:97-168
(actual 70 lines inlined here)
### internal/storage/storage.go:1-22
(actual 22 lines inlined here)
## Allowed Operations
Read/Edit/Write only on:
- internal/storage/ae.go
If you need anything else, ABORT with CONTEXT_INSUFFICIENT.
Done when: every task is [X], every wave has been committed, the diff matches the plan.
Verify
Input: spec.md + plan.md + tasks.md + the git diff produced by implement.
Output: a structured review with a verdict.
Seven dimensions, in order of typical catch-rate:
- Spec divergence — for each FR and acceptance scenario, is it actually implemented? Most common BLOCKER source.
- Missing tasks — any task marked
[X]whose diff content can't be found? - Security — reviewed against OWASP Top 10:2025, not the 2021 list. A03 (software supply chain) and A10 (mishandling of exceptional conditions) are new categories and the two most often missed.
- Correctness — off-by-one, races, edge cases the spec called out but the code skipped, swallowed errors.
- Contract breaks — schema, struct, signature, response-shape changes that would break callers.
- Necessity — abstraction with one caller, a dependency replacing three lines, config nobody asked for. Code that works and shouldn't exist is still a defect. See The Restraint Principle: YAGNI for Agents.
- Velocity illusions —
NOTE-level: tests that pass without touching the edges, PR text that doesn't match the diff, behaviour nobody can explain. Rarely blocks; predicts the next three bugs.
Two protocols wrap those dimensions. The hunt loops until dry — keep going until two consecutive rounds surface nothing new, deduplicating against everything already seen including refuted findings, because a fixed target either stops early or pads to reach itself. And every finding above NOTE faces an adversarial pass first: try to refute it, default to refuted when uncertain, and keep it only if it survives. Plausible-but-wrong findings cost more than missed ones, because they get acted on.
Findings tagged:
| Severity | Effect |
|---|---|
BLOCKER |
Must fix. Verdict = FAIL. Re-enter implement with the finding as the new task. |
WARN |
Real but non-blocking. Verdict still PASS if no BLOCKERs. Tracked. |
NOTE |
Observation. Informational. |
The reviewer's prompt explicitly tells it to look for what's missing, not validate what's present. Same model + same diff with a different framing produces a measurably different catch rate. Detail in Trust but Verify.
Example finding:
### Spec divergence — FR-005 (hard delete on file removal)
**[BLOCKER]** Spec FR-005 requires hard-delete of artifact rows when a
feature folder disappears. The diff implements file-level deletion in
ScanFeature, but ScanAll never calls DeleteAEFeature for slugs that have
vanished. Quote: importer.go:50-60.
Done when: BLOCKER count is zero. Anything less re-enters implement.
Retrospective + Skills
Input: every artefact produced this run, plus the verify report.
Outputs: a post-mortem with two structured blocks parsed by downstream commands.
CONSTITUTION_AMENDMENTS — rules that, if adopted, would prevent the failures seen this run:
CONSTITUTION_AMENDMENTS:
- id: cache-stateless-builders
rationale: |
Chroma formatter was constructed per HTTP request, costing ~3ms each.
rule: "Stateless objects with non-trivial construction cost should be
cached at handler init, not constructed per request."
severity: warn
These don't go straight into the constitution — they queue in known-failures.md with a count. Same ID seen 3+ times across runs → promotion happens automatically.
SKILLS_TO_EXTRACT — reusable patterns:
SKILLS_TO_EXTRACT:
- name: sqlite-fts5-pattern
rationale: Canonical FTS5 setup with content/content_rowid + the three
triggers, used twice now (articles_fts, ae_artifacts_fts).
template_path: skills/extracted/sqlite-fts5-pattern.md
/speckit.agentic.skills materialises each entry into a file under skills/extracted/, with trigger, applies-when preconditions, template, and references. Candidates lacking a concrete trigger or concrete steps are skipped and reported rather than written — vague advice is not a skill. Under Claude Code these become loadable skills, so future runs pick up the relevant ones without anyone pasting them back in.
Without this phase, every run is independent. With it, marginal cost trends down and marginal quality trends up. Detail in The Compounding Layer.
Done when: both YAML blocks parse cleanly and known-failures.md is updated.
Fast-track: /speckit.agentic.quickfix
Hard preconditions, enforced before any edit:
- ≤ 2 files
- ≤ 20 lines
- No new files / deps / public-API changes
- Commit type in
{fix, docs, style, chore, typo, refactor} - No protected paths (configurable)
Trip any of these → escalate to full pipeline. Ceremony for a typo is worse than no process.
Why waterfall
Each phase produces an artefact that gates the next. Unfashionable, intentional. The whole structure exists to defeat The One-Shot Problem: small steps to keep each agent invocation one-shottable, written checkpoints between them.
Trade made explicit: heavy planning and review up front, almost mechanical implementation after. The thing being avoided is the ten-minute agent run that returns the wrong feature. Every hard decision is made and recorded before the implementer starts writing.
The alternative — "give the agent the repo, tell it to do the next thing" — is the one-shot trap with extra steps.