This walkthrough is written from the perspective of the person who owns the requirements: the Business Analyst, Product Owner or Product Manager. The team, adoption and presentation sections are also useful to Project Managers planning how SDD fits delivery.
1. What SDD is, in one picture
In spec-driven development the spec, not the code, is the source of truth: you change the spec first, and the AI agent regenerates plan, tasks and code from it. For a BA this is good news, because the artifact you already own (requirements) becomes the thing that actually drives the build.
This guide uses GitHub Spec Kit, the most widely used open-source SDD toolkit, with Claude Code as the agent. The concepts transfer to other tools (Amazon Kiro, Tessl, or plain Markdown specs in Claude Code).
The artifacts you will meet (all plain Markdown files in the repo):
| File | Answers | Primary owner | Changes how often |
| .specify/memory/constitution.md | Rules every feature must obey (quality, tech limits, testing) | Tech lead + team | Rarely |
| specs/<feature>/spec.md | WHAT and WHY: user stories, requirements, acceptance criteria | BA / Product | Per change request |
| specs/<feature>/plan.md | HOW: tech approach, data model, file structure | Developer | When spec or tech changes |
| specs/<feature>/tasks.md | Ordered, checkable work items | Generated, reviewed by dev | Regenerated from plan |
| Source code | The implementation | Agent + developer | Derived from tasks |
The lifecycle you will run:
- Constitution, once per project.
- Specify: write the feature spec.
- Clarify: the agent interrogates your spec for gaps.
- Plan: the technical design.
- Tasks: the work breakdown.
- Analyze: consistency check across spec, plan and tasks.
- Implement: the agent writes code task by task.
- Converge: the agent checks code against the spec and appends tasks for gaps; repeat 7 and 8 until it reports Converged.
- Change request: edit the spec, then rerun steps 3 to 8.
The key mindset shift: in classic agile a requirement is consumed once and then lives in a ticket graveyard. In SDD the spec is versioned in git next to the code and stays true, so QA tests against it and the next change starts from it.
Why SDD instead of vibe coding
Vibe coding (prompt, accept, re-prompt until it looks right) is fast for a prototype but leaves no record of intent. SDD keeps the AI speed and adds the governance a team needs to ship and maintain the result.
| Dimension | Vibe coding | Spec-driven development |
| Source of truth | The chat history, then the code | A versioned spec in git |
| Governance | None; whoever prompts decides | Constitution sets rules every feature must pass |
| Structure | Ad hoc prompts | Fixed lifecycle: specify → clarify → plan → tasks → implement |
| Ambiguity | Agent guesses silently | Clarify step surfaces open questions before code |
| Review | Read the diff and hope | Review the spec and plan before any code exists |
| Traceability | Hard to say why code exists | Each task traces to a requirement in spec.md |
| QA | Tests what the code does | Tests against acceptance criteria in the spec |
| Change | Re-prompt; drift accumulates | Update the spec first, then regenerate |
| Handover | Knowledge lives in one person’s prompts | Any person or agent can pick up from the artifacts |
Pros and cons of SDD
| Pros | Cons |
| Intent is written down and reviewed before code, so fewer rework cycles | Slower start: specify, clarify and plan add time before any code |
| Specs and plans give auditable, git-tracked governance | Overkill for throwaway prototypes and one-line fixes |
| BAs, QA and devs work from one shared artifact | Specs go stale if people skip the update-spec-first rule |
| AI output is more consistent because the agent has clear constraints | Output quality is capped by spec quality; vague specs still give vague code |
| Easier onboarding and handover: the why is in the repo | Learning curve for the commands, artifacts and review habits |
| Change requests start from a known baseline | Tooling (Spec Kit and similar) is young and changes quickly |
Rule of thumb: vibe code to explore an idea, switch to SDD once it needs to be shared, reviewed or maintained.
2. Phase 0: Project setup on your Mac (about 15 minutes)
You need three things: uv (a Python tool manager that installs Spec Kit), the Spec Kit CLI, and Claude Code running inside VS Code. Run everything below in Terminal (or VS Code’s built-in terminal).
- Install uv. curl -LsSf https://astral.sh/uv/install.sh | sh, then open a new terminal tab. (Homebrew users: brew install uv.) Spec Kit needs Python 3.11+; if you lack it, uv python install 3.12 handles it.
- Install the Spec Kit CLI. v. If that package name is not found, use the git form: uv tool install specify-cli –from git+https://github.com/github/spec-kit.git. Verify with specify check.
- Make sure Claude Code is reachable from VS Code. Either install the Claude Code extension from the VS Code Extensions panel, or install the CLI (curl -fsSL https://claude.ai/install.sh | bash) and run claude in VS Code’s terminal. Log in with your Claude subscription. The Claude Desktop app’s Code tab also works, but VS Code lets you see specs and code side by side, which is the point of this exercise.
- Create the project.
mkdir -p ~/dev/todo-sdd && cd ~/dev/todo-sdd
specify init --here --integration claude
code .
About specify: it is the Spec Kit command-line tool you installed above. specify init –here turns the current folder (~/dev/todo-sdd) into a Spec Kit project: it adds the /speckit-* skills for Claude Code under .claude/skills/, plus the templates, helper scripts and an empty constitution under .specify/. It writes no app code; that comes later from the specs.
Side note: VS Code. Instead of running everything here in VS Code you can run in Claude CLI – in terminal instead of code . type claude.
- Start Claude Code in the project folder (VS Code terminal: claude) and type /speckit. You should see the skills: constitution, specify, clarify, plan, tasks, analyze, implement, converge, checklist.
What init created:
todo-sdd/
├── .claude/skills/ ← the /speckit-* skills Claude Code runs
├── .specify/
│ ├── memory/constitution.md ← project rules (empty template)
│ ├── templates/ ← spec, plan and tasks templates
│ └── scripts/ ← helpers that create feature folders
└── (specs/ appears after your first /speckit-specify)
Open .specify/templates/spec-template.md now and skim it. That template is the shape every spec will take, and later your team can customise it to match your company’s requirement format.
Syntax note: Spec Kit docs show both /speckit-specify (current skills mode in Claude Code) and /speckit.specify (older command mode). Use whichever your / menu lists; this guide uses the hyphen form.
Side note: why this walkthrough skips git. To keep the example simple we use no commits, branches or merges. In real team work git is not optional, it is what makes SDD trustworthy: traceability (who changed which requirement, when and why), one reviewable diff where spec, plan and code change together, pull requests for BA/dev/QA sign-off, and a safe undo when a regeneration goes wrong. Section 9 describes that team flow. Without git, if Spec Kit ever picks the wrong feature, just tell Claude which folder to use (e.g. “use specs/001-todo-list”).
3. Phase 1: Constitution (the project’s non-negotiables)
The constitution holds rules that apply to every feature, so you never repeat them in individual specs. In a company it is where architecture standards, security rules, accessibility and testing policy live. You write it once; the team amends it rarely and deliberately.
Paste into Claude Code:
/speckit-constitution Create principles for a small browser-only web app:
1. Plain HTML, CSS and vanilla JavaScript only. No frameworks, no build step, no backend. The app must run by opening index.html directly in a browser.
2. All data persists in the browser's localStorage. No network calls.
3. Every functional requirement in a spec must map to at least one test. Tests are plain JavaScript in a tests.html page that runs in the browser and shows pass/fail.
4. Accessible by default: keyboard operable, labelled inputs, visible focus.
5. Keep it simple: no feature that the spec does not ask for.
What to do with the output: open .specify/memory/constitution.md and read it like a contract. Edit anything you disagree with directly in the file; it is just Markdown.
Why rule 3 matters to you: it forces traceability from requirement to test. That is the hook for QA later: every acceptance criterion you write will have a visible, runnable check.
Why rule 5 matters: AI agents tend to over-build (add filters, dark mode, animations nobody asked for). Telling it explicitly to build only what the spec says is what makes the spec authoritative.
4. Phase 2: Writing the spec (your core job)
The spec describes WHAT the user needs and WHY, never HOW it is built. No mention of localStorage, functions or HTML here; that belongs in the plan. A good test: a stakeholder with no technical background should be able to read and sign off the spec.
You can give the agent a one-liner, but as a BA you will get far better results feeding it structured input. Paste:
/speckit-specify A personal to-do list for a single user in the browser.
Goal: help a user capture tasks quickly and see what is left to do.
User stories:
- As a user, I can add a task by typing a title and pressing Enter or clicking Add, so I can capture it quickly.
- As a user, I can mark a task complete or not complete, so I can track progress.
- As a user, I can delete a task I no longer need.
- As a user, my tasks are still there when I close and reopen the browser.
- As a user, I can see how many tasks are still open.
Acceptance notes:
- A task title cannot be empty or only spaces.
- Completed tasks stay visible but look visually distinct (e.g. struck through).
- Newest tasks appear at the top.
Out of scope: user accounts, sharing, due dates, categories, sync between devices.
What happens: Spec Kit creates a folder specs/001-todo-list/ containing spec.md. Open it in VS Code (Cmd+Shift+V for Markdown preview). You will typically see:
- User scenarios with priorities (P1, P2…) and acceptance scenarios in Given / When / Then form. This is the part QA will turn into test cases.
- Functional requirements numbered FR-001, FR-002… Each must be testable. These IDs become your traceability keys across plan, tasks, tests and bug reports.
- Key entities (here, a Task with title and completed state), described in business terms.
- Success criteria, measurable outcomes (e.g. “a user can add a task in under 5 seconds”).
- Edge cases, and [NEEDS CLARIFICATION: …] markers wherever the agent had to guess.
Your review, as the BA (do this every time):
- Every FR is testable: could QA write a pass/fail check for it?
- No implementation detail leaked in (frameworks, storage, function names).
- Nothing was added that you did not ask for. Delete it if so; rule 5 of the constitution backs you.
- Out-of-scope items are listed explicitly. This prevents scope creep from the agent and from people.
- Edge cases make sense: very long titles? 500 tasks? duplicate titles?
Edit spec.md directly wherever you disagree. It is your document; the agent only drafted it.
Staying sharp: the risk of AI-drafted specs
When AI drafts the spec, it is tempting to skim and approve. A spec that looks complete is not the same as one that is complete, and over time this can dull a BA’s attention to detail. Three things keep this in proportion:
- The prompt is your thinking. The agent can only elaborate on what you give it. Structured input (goal, user stories, acceptance notes, out of scope) carries your analysis into the spec; a one-line prompt hands that analysis to the AI.
- The draft is a starting point, not a deliverable. The spec is yours to edit, cut and extend. Clarify questions and the review checklist above exist to challenge it.
- A miss costs less here, but not nothing. Because the spec drives regeneration, a missed detail is added later by editing the spec and rerunning the chain, much faster than reworking hand-written code. Still, a detail missed in the spec can reach production, so prevention remains the goal.
Habits that keep BAs thinking:
- Write your own list of key rules and edge cases before running specify, then compare it with the draft: what did the AI add that you hadn’t thought of, and what did it miss that you had?
- Treat clarify questions as a signal: if the agent had to guess, your input was incomplete.
- Use /speckit-checklist for a requirements-quality check (completeness, clarity, testability).
- Have dev and QA review the spec at sign-off; fresh eyes catch what the author assumed.
- When a miss is found later, note its category (limits, errors, permissions…) and add it to your team’s spec template.
5. Phase 3: Clarify (let the agent interrogate your spec)
Clarify is a structured requirements review: the agent scans the spec for ambiguity and asks you up to about five targeted questions, usually multiple choice with a recommended answer. Your answers are written back into spec.md under a Clarifications section, so the decision is recorded, not lost in chat.
Run:
/speckit-clarify
For this app expect questions such as: Can a task title be edited after creation? Is there a maximum title length? Should deleting ask for confirmation? What should the list show when empty?
How to answer: answer as the product owner would. If you genuinely do not know, say so; in a real team that question goes to the stakeholder, and the spec stays on hold. For the exercise, pick: no editing in v1 (we will add it later as a change request), max 200 characters, no delete confirmation, show “No tasks yet” when empty.
Why this step is gold for a BA: it is the conversation you normally have with developers in refinement, done before any developer time is spent. Treat the questions it raises as a checklist of what your future specs should already answer. Over a few features you will notice the same categories coming up (limits, empty states, errors, permissions), and you can bake them into your team’s spec template.
Tip: to see exactly what clarify changed, copy spec.md to spec-before.md before running it, then compare the two in VS Code (right-click one file, Select for Compare, then right-click the other, Compare with Selected).
6. Phase 4: Plan, tasks, analyze (the developer’s half)
This is where HOW enters. In a team the developer leads these steps and you review them for one thing only: does the plan still deliver every requirement in the spec, and nothing extra?
Plan. Give the technical direction:
/speckit-plan Single index.html with a linked app.js and styles.css. Store tasks as a JSON array in localStorage under the key “todos”. Each task has id, title, completed, createdAt. Tests in tests.html using a tiny hand-written assert helper, no libraries.
This produces plan.md plus supporting files such as data-model.md (the Task structure), research.md (decisions and alternatives considered) and quickstart.md (how to run and validate). It also runs a “constitution check”: if the plan breaks a principle (say, it pulls in a framework), it must flag and justify it.
Tasks.
/speckit-tasks
tasks.md is an ordered checklist (T001, T002…), grouped by user story, with tests written before the code that makes them pass. Tasks marked [P] can run in parallel. Each task should reference the requirement or story it serves; that is your traceability chain: FR-003 → T007 → test → code.
Analyze (a read-only quality gate, highly recommended):
/speckit-analyze
It cross-checks spec, plan and tasks and reports gaps: a requirement with no task, a task with no requirement, terminology drift, constitution violations. Fix findings in the relevant file before you implement. This is the cheapest bug fix you will ever make.
Reread all three artifacts before moving on. In a team this is the spec review / sign-off point, where BA, dev and QA approve them before any code exists.
7. Phase 5: Implement, converge, and run it
Now the agent writes code, working through tasks.md and ticking items off as it goes. Claude Code will ask permission before creating files or running commands; read what it proposes and approve.
/speckit-implement
Watch tasks.md in VS Code while it runs: boxes change from [ ] to [X]. When it finishes, check against the spec, not just “does it work”:
/speckit-converge
Converge compares the implementation with spec, plan and tasks. If something is missing (say, the “No tasks yet” empty state from clarify), it appends new tasks to tasks.md. Run implement then converge again until it reports Converged.
Run the app. No server needed:
open index.html
open tests.html
Do a QA pass yourself, straight from the spec. Open spec.md beside the browser and walk every Given / When / Then:
- Add a task with Enter, and with the Add button. Newest appears on top.
- Try an empty title and one of only spaces. Nothing is added.
- Paste 250 characters. It is limited to 200 (from clarify).
- Toggle complete; the item is struck through and the open counter changes.
- Delete a task. Reload the page; state persists.
- Delete everything; “No tasks yet” shows.
- Use only the keyboard (Tab, Enter, Space).
- tests.html shows all tests passing.
If something is wrong, decide which kind of wrong it is. This is the most important habit in SDD:
| Symptom | Real cause | Fix where |
| Code does not do what the spec says | Implementation bug | Ask the agent to fix it against FR-xxx; spec unchanged |
| Code does what the spec says, but that is not what the user wants | Spec defect | Fix spec.md first, then regenerate downstream (section 8) |
| Spec is right but the plan chose badly (e.g. slow rendering) | Plan defect | Fix plan.md, rerun tasks and implement |
Never silently patch code for a spec defect: the spec then lies, and the next regeneration reintroduces the bug.
That is v1 done: spec, plan, tasks, code and tests all agree.
8. Phase 6: Change requests (where SDD earns its keep)
The rule: change the spec first, then let everything downstream regenerate from it. Spec Kit supports two ways to record a change, and you should try both, because your team will have to pick one (Spec Kit guide).
| Model | How it works | Best for |
| Living spec | Edit the existing spec.md; plan and tasks are re-derived. The spec always describes current behavior. | Tweaks and extensions of an existing feature |
| Flow-forward | Each change is a new numbered spec folder; old ones stay as history. | New capabilities, audit trails, regulated work |
Change request A (living spec): “Users can edit a task’s title”
We deferred this during clarify. Now the stakeholder wants it.
- Edit specs/001-todo-list/spec.md yourself. Add a user story (“As a user, I can double-click a task to edit its title; Enter saves, Escape cancels”), a new requirement (e.g. FR-012, same rules as creation: not empty, max 200 characters), acceptance scenarios, and update the Clarifications entry that said “no editing in v1”. Hand-editing is the point: this is you, the BA, owning the contract.
Side note: adding the story alone won’t work, because in v1 a click on the title toggles completion, so a double-click never registers. Also change the existing completion requirement to “only the checkbox toggles completion; clicking the title does nothing”, add a scenario for it, define edit mode (Enter saves, Escape cancels, an empty title reverts), and add an Edit button for keyboard users. You don’t have to edit by hand: ask Claude Code to “update specs/001-todo-list/spec.md only, no code” with these points, then review what it changed.
- /speckit-clarify to let the agent probe the new story (what if the edit makes the title empty: revert or delete?).
- /speckit-plan (with no extra text, or “keep the existing approach”) so the plan absorbs the change.
- /speckit-tasks, then /speckit-analyze. Check that FR-012 has tasks and tests.
- /speckit-implement, then /speckit-converge until Converged.
- Retest in the browser: the new scenarios plus a quick regression run of the old ones.
Now reread what changed: spec, plan, tasks and code moved together as one unit. In a team, these changes go through spec review / sign-off together; here you are the reviewer.
Change request B (flow-forward): “Filter: All / Active / Completed”
This is a new capability, so it gets its own spec:
/speckit-specify Add filtering to the existing to-do list. Users can switch between All, Active and Completed views. The selected filter is remembered after reload. The open-task counter always shows open tasks regardless of filter. Builds on specs/001-todo-list.
This creates specs/002-…. Run the same chain (clarify, plan, tasks, analyze, implement, converge) and retest. Spec 001 stays untouched as the record of v1.
What to take away: in both models, nobody edits code first. If a developer finds a better approach mid-build, Spec Kit’s “flow-back” model lets them update the plan or spec, but the artifacts must agree again before implementation continues.
9. Working as a team: who does what
In a team the spec lives next to the code, and every change to it goes through spec review / sign-off by BA, dev and QA. That changes the BA role: you stop handing requirements over and start co-owning a versioned artifact with developers and QA.
| Stage | BA | Developer | QA | UI designer |
| Constitution | Contributes business rules (compliance, accessibility) | Owns tech standards | Owns testing policy | Adds design-system rule (e.g. every screen has a design.md entry) |
| Specify + clarify | Owns: writes input, runs specify, answers clarify, edits spec.md | Reviews for feasibility | Reviews for testability; flags vague criteria | Pairs with BA: frames per user story, writes design.md, answers UI questions from clarify |
| Spec review / sign-off | Author | Approver | Approver | Approver; frames marked Ready for dev |
| Plan + tasks + analyze | Reviews: does the plan cover every FR, nothing extra? | Owns | Reviews test approach in tasks | Reviews token translation and that UI tasks reference the right frames |
| Implement + converge | Available for questions; any answer goes into spec.md, not chat | Owns; reviews every agent change | Starts test cases from Given/When/Then | Answers visual questions; answers go into Figma + design.md |
| Verification | Acceptance against success criteria | Fixes bugs against FR IDs | Owns: tests trace to FR IDs | Visual review of the build against the frames |
| Change request | Owns the spec edit | Reruns plan → implement | Updates affected test cases | Updates Figma + design.md in the same review when the UI is affected |
Working agreements to propose to your team:
- Two review points per feature: a spec review / sign-off (spec, plan, tasks, no code) by all three roles, then a code review of the implementation. Catching a misunderstanding at the first one costs minutes.
- Decisions made in Slack, Teams or meetings are not real until they are in spec.md. The agent only knows what is written.
- Requirement IDs (FR-xxx) are used everywhere: Jira tickets, test cases, bug reports, commit messages. A Jira story can simply link to the spec file on the branch instead of duplicating the text.
- Bug triage always asks “code bug, spec bug or plan bug?” (the table in section 7).
- Only one active editor of a given spec.md at a time, like any shared document; merge conflicts in Markdown are easy, but conflicting intent is not.
- Developers still review every line of generated code. SDD shifts effort toward specs and review; it does not remove engineering judgement.
What changes for you as BA: you will write in Markdown in VS Code (or have the agent draft and you edit), learn enough git to branch and commit, and write requirements precisely enough to be executable. The upside is that your spec is no longer a document that goes stale after sprint one.
Working with a UI designer
The designer’s work lives in two places: Figma holds the visual design, and a one-page design.md next to spec.md connects it to the requirements. Spec Kit has no built-in design step, so this is a team convention layered on top.
Who owns what: spec.md owns behaviour (validation, limits, what happens on reload); Figma and design.md own look and layout. Neither repeats the other. If they conflict, the spec wins on behaviour and the design wins on visuals. A frame cannot express “Escape cancels editing”, so that always stays in the spec.
What goes in design.md:
# Design: 001-todo-list
Figma file: <link> · Frames marked Ready for dev on 2026-10-10
Design system: <name / version>, colours, spacing and type as Figma variables
## Frames
| Frame | Figma link | Covers |
| Task list, default | <frame link> | US1, FR-001, FR-003 |
| Task list, empty | <frame link> | FR-009 |
| Task item, completed | <frame link> | US2, FR-004 |
| Title too long / invalid | <frame link> | FR-002 |
## States to cover
Empty, error, focus, long text (200 chars)
## Accessibility notes
## Open design questions
The Covers column is the design’s traceability: each frame maps to user stories and FR IDs, the same keys QA and developers use.
How Figma links become code. Figma’s MCP server lets Claude Code open a frame from its link and read its layout, styles and variables.
- The designer gives each screen and state its own named frame, uses Figma variables for colors, spacing and type, and copies frame-level links (right-click → Copy link to selection; the link contains a node-id).
- At plan, the developer points the agent at the design, e.g. /speckit-plan … Follow specs/001-todo-list/design.md. Read each Figma frame via the Figma MCP. Turn Figma variables into CSS custom properties in styles.css. Plain HTML/CSS, no React. The plan records how the design is translated.
- Tasks reference specific frames. During implement the agent pulls each frame’s design context, variables and screenshot, then writes the HTML/CSS. Figma’s generated code is React + Tailwind by default, so it is a reference, not something to paste in.
- At verification the designer compares the running app to the frames; the agent can also compare a frame screenshot with the page.
Practical catches:
- Figma changes silently. A frame edited after approval means the link now points at something new. Mark frames Ready for dev, note the date or version in design.md, and route design changes through the same review as spec changes.
- Undrawn states get guessed. If no frame shows a state (e.g. a 200-character title wrapping), the agent invents one. Clarify should flag it and the designer adds the frame.
- Access. Each developer’s Claude Code needs the Figma MCP added and authenticated with their own Figma account, which typically needs a paid plan with a Dev or Full seat. Check your company’s licenses.
- Code Connect (optional). With a coded component library, Code Connect maps Figma components to it so the agent reuses them instead of rebuilding. Overkill for the to-do app, relevant for a real product.
- Visual defects get triaged too: code bug (build doesn’t match frame), design bug (frame is wrong or missing), or spec bug (behavior wrong).
Working with Jira
The spec is the source of truth; Jira tracks the work and points at the spec, never copies it. Requirements copied into two places drift apart within a sprint.
| Jira item | Maps to | What it contains |
| Epic | One spec folder (e.g. specs/001-todo-list) | One-line summary + links to spec.md and design.md |
| Story | One user story in the spec (US1, US2…) | Link to that spec section + the FR IDs it covers; acceptance criteria say “see spec” |
| Sub-task | Not used for tasks.md | Tasks are regenerated whenever the plan changes, so mirroring them creates churn; devs work from tasks.md, progress is tracked per story |
| Bug | A broken FR | FR ID it breaks + triage label: code, spec, plan or design |
| Change request | A spec edit | Links to the updated spec section; living-spec changes stay in the same epic, flow-forward changes get a new epic |
Working rules:
- Keys both ways, entered manually. spec.md records the epic key in its header and each story key next to its user story, e.g. “US1 (TODO-12)”. Jira links back to the spec file in the repo.
- Definition of Ready: a story moves to Ready for dev only when the spec review is approved and its design frames are marked Ready for dev.
- Bug routing: spec and design bugs go back to the BA or designer, not straight to a developer.
- QA test cases (in Jira or a test tool) reference FR IDs, so a failing test points to both a requirement and a story.
Side note: Jira as input. A BA can also start from a story already written in Jira and paste it into /speckit-specify as input. Once the spec exists it takes over, and the Jira description is replaced with a link so there is only one source.
Side note: automating with the Atlassian MCP. With the Atlassian MCP added to Claude Code, the agent can read a Jira story to seed /speckit-specify, create the epic and stories from an approved spec.md, write the keys back into the spec, and post links, comments and status changes. It acts with your Jira permissions, so have it propose issues for a human to confirm rather than bulk-creating them.
Bugs found in testing
Every bug is fixed in the layer it came from: code, spec, plan or design. That is what keeps the spec trustworthy after release.
- Log. QA records the bug with steps to reproduce, actual result, a screenshot, and the expected result quoted from the spec (FR ID plus the Given/When/Then it breaks). QA proposes a triage label. A bug that cannot point to any FR is either correct behavior or a spec gap, which makes it a spec bug.
- Triage. BA and a developer confirm the label (designer joins for visual issues) by asking: does the code match the spec, and does the spec match what the user needs?
- Fix, by type:
| Label | Who fixes | Where | Steps |
| Code | Developer | Code only | Spec untouched; fix against the FR with the bug extension (below) |
| Spec | BA | spec.md first | Edit spec → clarify → plan → tasks → analyze → implement → converge |
| Plan | Developer | plan.md | Edit plan → tasks → implement → converge |
| Design | Designer | Figma + design.md | Update frames → developer reruns affected UI tasks |
- Retest and close. QA retests the original scenario plus a short regression on related FRs. For spec and design bugs QA first updates the affected test cases, because the expected result changed. Close only when spec (or design), code and tests agree again.
Code bugs with Spec Kit’s bug extension (docs): install with specify extension add bug. The developer runs /speckit-bug-assess (the bug report as input), /speckit-bug-fix and /speckit-bug-test. Reports land in .specify/bugs/<slug>/ with a verdict of verified, partial or failed; only verified counts as fixed. Add a regression test tied to the FR ID and link the report from the bug.
Never patch code to fix a spec bug: the spec would then describe the wrong behavior, and the next regeneration brings the bug back.
With Jira: the bug is a Jira Bug linked to its story, carrying the FR ID and triage label as described in Working with Jira; status moves through the Jira workflow.
Without Jira: bugs live in the repo as Markdown files next to the spec they break, e.g. specs/001-todo-list/bugs/BUG-003-empty-title.md:
# BUG-003: Whitespace-only title is added
Status: Open (Open → Triaged → Fixed → Verified / Reopened)
Triage: code (code | spec | plan | design)
Severity: Medium
Breaks: FR-002, US1 scenario 2
Found by / date: QA, 2026-10-14
## Steps to reproduce
## Expected (from spec)
## Actual
## Screenshot / notes
## Fix reference
The flow is the same; people update the Status line instead of a Jira workflow, and the developer pastes the bug-extension report path into Fix reference. For an overview, ask Claude Code: “List all bugs under specs/ that aren’t Verified, grouped by triage type and FR.” Trade-offs: no assignment, notifications or dashboards, so this suits a small team or pilot; git becomes essential as the only history of status changes. If the repo is on GitHub or Bitbucket, their built-in Issues are a middle ground.
Multi-team / microservices (backend, frontend, API)
SDD works with split development roles; the API contract becomes the hand-off point. The BA still writes one business spec per feature, the API role turns it into a contract, and frontend and backend build against that contract in parallel.
- Specify + clarify (BA, unchanged). One spec.md per feature, in business terms, with no service boundaries. Frontend, backend and API developers all review it.
- Contract first (API role). Plan typically generates a contracts/ folder beside plan.md. The API engineer owns it and turns it into OpenAPI definitions: endpoints and which service owns each, request and response shapes, error codes, and the FR IDs each endpoint serves. The contract is reviewed and approved before anyone builds.
- Parallel plans and builds. Frontend and backend each run plan → tasks → implement → converge against the approved contract. Frontend works against mocks generated from the OpenAPI file; backend implements the endpoints with contract tests.
- Integration. QA runs end-to-end tests from the spec’s Given/When/Then against the real services; contract tests catch drift between services.
Where specs and contracts live depends on the repo setup (ask a dev lead whether services share one repository or each has its own):
| Setup | How to recognize it | Where specs live |
| Monorepo | One repository with a services/ (or similar) folder inside | One specs/ folder; tasks tagged [API], [BE], [FE]; closest to this walkthrough |
| Separate repos | Many repositories named after services (e.g. orders-api, web-frontend) | Feature spec and contracts in a shared specs repo; each service repo has its own Spec Kit setup that references the feature spec and FR IDs |
With separate repos, Claude Code only sees the repo it runs in, so the feature spec and contract must be reachable from each service: a shared repo, a published package, or a copied file. Spec Kit does not solve this; it is a team convention.
Other adjustments:
- Two-layer constitution: org-wide rules (API standards, security, versioning) plus per-service additions.
- Change requests get an impact check: if a spec change alters the contract, the API role decides whether it is backward-compatible or needs a new version before frontend and backend start.
- A fifth triage label, contract: services disagree about the API even though each matches its own plan.
- Jira: one epic per feature as before; split stories per layer (e.g. “US1 – API”, “US1 – FE”) or keep them whole with a Component field for the layer.
Compliance and audit
SDD is not audit-proof out of the box, but it is audit-friendly: it produces most of the evidence auditors ask for as a by-product of the work. Making it audit-proof means adding controls around it, several of which this walkthrough skipped for simplicity (git, PRs, recorded approvals).
What SDD already provides:
- Documented requirements and changes: spec.md versions, Clarifications and research.md record what was decided and why.
- Traceability: FR ID → task → test → code, plus Jira links and bug reports citing FR IDs; most of a traceability matrix, built as you go.
- Documented standards: the constitution is a written policy every plan is checked against.
- Verification evidence: analyze and converge reports, bug-extension verdicts, test results.
What you must add:
| Audit need | Why SDD alone falls short | Control to add |
| Tamper-evident history | Files can be edited silently | Git with a protected main branch; optionally signed commits |
| Recorded approvals | A sign-off in a meeting leaves no evidence | Approvals captured with name and date: PR approvals, a Jira workflow transition, or e-signature |
| Segregation of duties | One person could write, generate and approve | Author ≠ approver; the agent never merges; humans review all generated code |
| AI governance | Agent behavior and data handling are new risks | Approved tool list; no sensitive or regulated data in prompts; locked-down Claude Code permissions; agent changes always reviewed |
| Test evidence | Manual browser checks leave no record | Automated tests in CI with results stored per release |
| Spec = reality at release | Specs can drift after go-live | Converge before every release; tag the release with the spec versions it shipped |
Cautions:
- Regulated validation (e.g. GxP / 21 CFR Part 11): AI output is not deterministic, so validate the process and its outputs (reviews, tests), not the AI tool itself. The quality team decides how.
- Do not claim compliance. Say “SDD produces the evidence our controls need”. Whether it satisfies SOC 2, HIPAA or other frameworks is for the compliance team to confirm; mapping SDD artifacts to their control list is a good pilot step.
PO sign-off and audit trail
Spec Kit itself keeps no audit log, and neither does Claude Code in any useful sense. The real audit trail comes from the repo host (GitHub, Bitbucket or GitLab) and from Jira; for PO sign-off you normally use both.
| Source | What it records | Good enough for sign-off? |
| Spec Kit files | Dates inside spec.md (e.g. clarify sessions), editable by anyone | No: no identity, not tamper-evident |
| Claude Code | Session transcripts stored locally per user | No: local and personal; enterprise plans may have admin-level logs |
| Git history | Who changed which line and when, tied to an exact version | Partly: shows changes, not approval; identity can be faked unless commits are signed or verified |
| Pull request approvals | Who approved which exact version, and when | Yes: the strongest technical record |
| Jira issue history | Every field change and status transition, with user and timestamp | Yes, if the transition points to the exact spec version |
The key rule: an approval must point to an exact version of the spec. “PO approved the spec” means little when the spec keeps changing; “PO approved commit a1b2c3d” or “PO approved PR #42” is auditable.
How PO sign-off works:
- In the repo: a rule makes the PO a required approver for any change under specs/ (a CODEOWNERS file on GitHub; Bitbucket and GitLab have equivalents). A spec change cannot merge without the PO’s recorded approval.
- In Jira: a “Spec approved” status only the PO can move a story into, with a link to the approved PR or version.
Running both in parallel with Jira:
- The BA writes or changes the spec on a branch named with the Jira key (e.g. TODO-12-todo-list); Jira’s development panel then shows the branch and PR on the story automatically.
- Dev, QA, designer and PO review; the PO’s approval is required by the repo rule.
- When the spec PR merges, Jira automation moves the stories to Ready for dev and comments with the approved version: the Definition of Ready, with evidence.
- The implementation PR uses the same Jira keys; merging it moves stories to In QA.
- A change request repeats the cycle with a new spec PR and a new PO approval, linked from the CR ticket.
The result is two independent records pointing at each other: the repo proves what exactly was approved, Jira proves the process was followed.
Without git and PRs (as in the simplified walkthrough), the fallback is a PO-only Jira transition with a snapshot of the spec attached (e.g. a PDF export). It works, but it is weaker, because the attachment can drift from the live file.
Scaling to large projects
Large projects use SDD as many small feature specs on top of a shared layer of project-wide context, never as one big spec. The challenge shifts from writing specs to organizing them.
1. Break the work into feature-sized specs.
- Product level: a short overview (vision, capability map, feature list). Not a Spec Kit spec, but every spec links to it.
- Feature level: one Spec Kit spec per vertical slice, deliverable in about 1–2 weeks, typically 2–6 user stories.
- Rule of thumb: more than about 10 clarify questions or about 40 tasks means the spec is too big; split it.
- Jira: initiative → epic per feature spec → stories.
2. Build a shared context layer so specs don’t repeat it:
- Constitution: standards, security and NFRs (performance, accessibility), written once.
- Glossary and domain model: so terms mean the same thing in every spec.
- Architecture overview and API contract registry: so plans reuse existing services instead of reinventing them.
The agent cannot hold a large system in mind, so each feature spec points to just the shared documents it needs. That keeps the agent focused and consistent.
3. Make dependencies and ownership explicit.
- Each spec declares what it depends on (e.g. “Depends on 003-user-roles”) and what it must not break.
- Each product area has an owning BA and team, so two specs never silently redefine the same behavior.
- Analyze only checks within one feature, so add a periodic cross-feature review (BA lead plus architect) to catch conflicts between specs.
4. Keep the current picture readable. After dozens of flow-forward specs, nobody can tell what the system does today by reading them all. Mix the two models from section 8: a living spec per capability (e.g. “Reporting”) describes current behavior, and change specs feed into it. Spec Kit’s community extensions include archive and reconcile tools worth evaluating at that size.
5. Existing (brownfield) systems. Do not spec the whole legacy system up front. Spec only the area you are about to change; the agent can draft a current behavior spec from the code, which the BA corrects. Coverage grows with each change. Spec Kit has a guide for existing projects.
6. Big features at implementation time. Implement in phases (e.g. /speckit-implement Phase 1 only) and converge after each phase. Long, unattended agent runs over large task lists drift more.
Where it strains: spec sprawl, cross-feature consistency, and the upfront effort of the shared layer. Pilot on one bounded feature and grow the shared layer gradually.
When to adopt SDD
Consider SDD when a project is long-lived, involves several roles, and uses (or will use) AI coding agents; skip it for throwaway or tiny work. In a company, adopt it gradually through a pilot, not a mandate.
Project fit:
| Good fit | Poor fit |
| Product maintained for years | Throwaway prototypes, spikes, exploratory R&D |
| BA, dev, QA (and designer) need one shared truth | Tiny fixes, config changes, one-liners |
| Rework caused by misunderstood requirements | Requirements genuinely unknown, stakeholders unavailable |
| Devs already use AI agents and output is inconsistent | Mid-crunch on a critical release |
| Traceability or audit needs | No AI tool approval in place |
| Handovers, turnover, external vendors |
Rule of thumb: vibe code to explore; switch to SDD once the work needs to be shared, reviewed or maintained.
Company readiness checklist (before a pilot):
- Security and compliance have approved the AI coding tool and its data use.
- Basic repo practices exist (git, code review).
- A tech lead will own the constitution.
- BAs are willing to write precise, testable requirements in Markdown.
- Devs will review all agent output, not rubber-stamp it.
- Management accepts a slower start for less rework later.
Choosing the pilot project: a bounded 2–4 week feature, off the critical path, with real users; one small team (BA, dev, QA, plus a designer if there is UI); a comparable past feature as baseline for cycle time, rework and defects; low regulatory risk the first time.
Phased rollout:
- Pilot: one feature, success criteria agreed upfront.
- Team default: if the pilot beats the baseline, the team uses SDD for all new features; templates and conventions settle.
- Organization standard: more teams adopt it with a shared constitution, glossary and spec template; the scaling practices above apply.
Signals to pause or adjust: specs skipped or edited after the code, sign-offs becoming rubber stamps, or overhead outweighing rework savings on small items. Narrow the scope rather than abandon it.
SDD vs classic approaches
| Dimension | Classic (Jira stories, manual dev) | Classic + AI helpers (no specs) | SDD |
| Source of truth | Stories go stale once dev starts | Decisions buried in lost chat prompts | One versioned spec that stays true |
| Requirements quality | Gaps surface in dev or QA | Gaps filled by AI guesses | Clarify and analyze catch gaps before any code |
| Delivery speed | Slowest | Fast at first, slows with rework | Fast and repeatable once the spec is approved |
| Change requests | Manual rework, story rarely updated | Quick code edits nobody can trace | Edit the spec once; plan, tasks and code follow |
| Consistency | Varies by developer | Varies by prompt | Enforced by constitution and spec |
| Figma → UI | Rebuilt by eye, pixel drift | Partial help from screenshots | Agent reads frames and tokens directly |
| QA | Test cases written late, from vague stories | Same, with more untested AI code | Tests derived from acceptance criteria from day one |
| Rework | High: misunderstandings found late | High: AI builds the wrong thing fast | Low: misunderstandings caught at spec review |
| Traceability / audit | Manual, partial | Effectively none for AI changes | Requirement → task → test → code by design |
| Knowledge retention | Leaves with people | Leaves with people and prompts | Lives in the repo for every new joiner |
| BA impact | Requirements consumed once, then forgotten | Same | BA’s spec drives the build directly |
| Developer impact | Time spent on boilerplate | Time spent fixing AI drift | Time spent on design and review |
| Main cost | Slow delivery, rework | Hidden quality and audit debt | Upfront spec effort and new habits; offset by less rework |
AI already makes code cheap; SDD makes it correct, consistent and traceable.
About the upfront cost. SDD asks for an investment beyond spec writing. Everyone has a learning curve: BAs writing precise, testable specs in Markdown, developers moving from writing code to designing plans and reviewing agent output, QA tracing tests to requirement IDs. There is also a human side. People are used to how they work today, and a change this big naturally meets hesitation, so expect slower delivery during the first features. The good news is that IT teams are driven by curiosity and by making things work better. Given time to learn, a safe pilot and visible results, people usually move from skepticism to ownership, and start improving the process themselves.
10. Cheat sheet and troubleshooting
| Command (in Claude Code) | When | Input (text after the command) | Writes |
| /speckit-constitution | Once per project | Optional: your principles, e.g. /speckit-constitution Plain HTML/JS, no frameworks; every feature needs acceptance criteria | .specify/memory/constitution.md |
| /speckit-specify | New feature | Required: the what and why, no tech, e.g. /speckit-specify A to-do list where users add, complete and delete tasks | Folder specs/NNN-*/spec.md |
| /speckit-clarify | After specify, or after editing a spec | Optional: an area to focus on, e.g. /speckit-clarify focus on empty states and validation | Clarifications in spec.md |
| /speckit-checklist | Optional quality checklist for the spec | Optional: the domain to check, e.g. /speckit-checklist UX | checklists/ |
| /speckit-plan | After the spec is approved | Recommended: tech stack and constraints, e.g. /speckit-plan Single index.html, vanilla JS, localStorage | plan.md, data-model.md, research.md, quickstart.md |
| /speckit-tasks | After plan | None needed | tasks.md |
| /speckit-analyze | Before implementing | None needed | Report only |
| /speckit-implement | Build | Optional: scope, e.g. /speckit-implement Phase 1 only | Code, ticks in tasks.md |
| /speckit-converge | After implement, repeat until Converged | None needed | Appends gap tasks |
Troubleshooting:
- /speckit shows nothing: make sure Claude Code was started inside the project folder (where .claude/ lives), then restart the session. Try the dotted form /speckit.specify in case your version installed commands instead of skills.
- specify: command not found: open a new terminal after installing uv, or run uv tool update-shell.
- The agent works on the wrong feature: tell Claude which folder to use, e.g. “use specs/001-todo-list”.
- The agent added things you didn’t ask for: delete them from spec or plan, cite constitution rule 5, rerun analyze.
- Updating Spec Kit later: specify init –here –force –integration claude refreshes templates but can overwrite constitution.md and customized templates; copy those somewhere safe first.
Further reading: Spec Kit README · Evolving specs guide · spec-driven.md in the Spec Kit repo for the full methodology.

Leave a Reply