Key takeaways
- Spec-driven development (SDD) consists of writing a spec before generating code with an AI agent; the spec becomes the source of truth for both humans and machines.
- Three levels coexist: spec-first (spec written before the task), spec-anchored (spec kept for maintenance) and spec-as-source (only the spec is edited).
- The main risk is falling back into the waterfall cycle: too much Markdown, doubled reviews and specs that drift as soon as the code grows.
- A useful spec stays short, behavior-oriented, and links directly to user stories and their acceptance criteria in a living backlog.
- Acceptance criteria in GIVEN/WHEN/THEN format serve as a verifiable contract for both the AI agent and the team.
Why the spec is back on the table
A code assistant is a text box and not much else. No menus, no imposed structure. How do you make sure the code produced meets the need? The answer that has become dominant since 2025: write a spec before letting the agent code.
This movement has a name: spec-driven development (SDD). Birgitta Böckeler, Distinguished Engineer at Thoughtworks, gives an operational definition: write a spec before writing code with AI, with the spec becoming the source of truth for both the human and the agent (martinfowler.com).
GitHub, AWS with Kiro, Tessl and the BMad method all offer tools for this. The principle is simple: an initial prompt, a few instructions, and the LLM generates product specs, an implementation plan and a task list. Each document depends on the previous one. The human edits, the agent codes.
Three levels of SDD, not to be confused
Not all tools that claim to be SDD aim for the same thing. Böckeler distinguishes three levels (martinfowler.com):
- Spec-first: a carefully written spec is written before the task, then used in the AI-assisted development flow.
- Spec-anchored: the spec is kept after the task, to evolve and maintain the feature.
- Spec-as-source: the spec is the main file over time; the human never touches the code.
All tools are at least spec-first. Few embrace spec-anchored, even fewer spec-as-source. This is a point to check before adopting a tool: what happens to the spec in six months, once the code has moved on?
Another useful distinction: the spec is not the memory bank. The memory bank is the set of context files valid for all sessions (rules, product description, architecture). The spec, on the other hand, concerns only the feature currently being created or modified.
The trap: Markdown that buries agility
François Zaninotto, at Marmelab, tested SDD and his verdict is sharp: "Spec-Driven Development (SDD) revives the old idea of heavy documentation before coding — an echo of the Waterfall era" (marmelab.com).
His example is telling. With GitHub spec-kit, a simple feature — displaying today's date in a time-tracking app — produced 8 files and 1,300 lines of text. With Kiro, adding a "referred by" field to contacts generated three documents: requirements, design, tasks.
The problems he lists:
- Context blindness: the agent discovers context through text search and misses existing functions that need updating.
- Markdown madness: too much text, especially in the design phase; you read instead of thinking.
- Systematic bureaucracy: repetitions, imaginary edge cases, useless refinements.
- Fake agile: the generated "user stories" are not user stories. "As a system administrator, I want the referred by relationship to be stored in the database" is not a user story.
- Double code review: the technical spec already contains code, which must be reviewed before reviewing the final implementation.
- Diminishing returns: SDD shines on a new project, but stalls as soon as the codebase grows.
Zaninotto sums it up: "spending 80% of your time reading instead of thinking". SDD, in its heavy version, repeats the mistake of Big Design Up Front.
What a spec that holds up looks like
A useful spec is not a 40-page document. It is a structured, behavior-oriented artifact, written in natural language, that expresses a feature and guides the agent (martinfowler.com). Three elements are enough.
The behavior contract
What the module, function or endpoint must do. Preconditions, postconditions, invariants. Input, output and error types. No ambiguity.
The catalog of edge cases
What happens if the input is null? Empty? At maximum size? Negative? Unicode? Concurrent? The VSDD method suggests asking these questions explicitly to the agent so that it is exhaustive (gist.github.com).
Acceptance criteria
In GIVEN… WHEN… THEN… format. This is what Kiro produces in its requirements document, each requirement being a user story with its criteria (martinfowler.com). These criteria serve twice: for the agent to generate the code, and for the team to verify that the code does what was asked.
The key point: the spec must be verifiable. If you cannot write a test or an acceptance criterion that validates it, it is too vague.
Linking spec, user stories and backlog
A spec that lives in an isolated file is useless. It must be anchored in the backlog, in the same place as the user stories and their acceptance criteria.
The AI SDLC Scaffold framework proposes a folder structure that materializes this link: a 1-spec/ folder with goals/, user-stories/, requirements/, assumptions/, constraints/ subfolders, each with its own template (github.com). Each artifact has an identifier (US-, REQ-, ASM-, CON-) and lives in the repository, versioned with the code.
Concretely, in an agile project management tool:
- A user story in the backlog with its priority and points.
- Its acceptance criteria in GIVEN/WHEN/THEN format, in the description or as a checklist.
- The short spec attached to the story: contract, edge cases, non-functional constraints.
- The link to the generated code, to keep traceability.
This way of working fits naturally into a Kanban board with customizable columns, checklists and an activity log. For example, in Ever Earlier, a user story can carry its acceptance criteria as a checklist and its spec as an attachment or pinned comment, which avoids maintaining a separate Markdown folder that gets out of sync with the backlog.
The important thing is not the tool. It is that the spec stays alive: updated when the story evolves, archived when it is delivered, never left to rot in a corner.
A five-step method
Here is a concrete method for short specs that generate correct code.
- Start from intent, not code. Describe the feature in three sentences. If you cannot, the need is not clear.
- Write the user story and its acceptance criteria. Use the "As a… I want… so that…" format for the story, GIVEN/WHEN/THEN for the criteria. Three to five criteria maximum.
- Add the behavior contract and edge cases. Half a page is enough. Explicitly list degenerate inputs.
- Have the agent generate the code, then review it. The spec guides, it does not guarantee anything. Böckeler quotes GitHub: "Crucially, your role isn't just to steer. It's to verify." (martinfowler.com)
- Update the spec after delivery. If the code has diverged, the spec must reflect it. Otherwise, it becomes a documented lie.
A point of vigilance: agents do not always follow the spec. Zaninotto recounts that an agent marked the "verify implementation" task as done without writing a single unit test, instead writing manual test instructions (marmelab.com). Human verification remains non-negotiable.
What SDD does not solve
SDD is not a magic wand. Three limits to keep in mind.
The quality of training data. A study by Penn State and Oregon State University, published in Media Psychology, shows that most users do not detect a systematic bias in an AI system's training data, even when it is visible (psu.edu). Out of 769 participants across three experiments, the majority did not notice that happy faces were mostly white and sad faces mostly black. In other words: a well-written spec does not protect against a biased model.
Existing context. SDD shines on a new project. On a mature codebase, specs miss the context and slow down development. Zaninotto says it bluntly: "For large existing codebases, SDD is mostly unusable."
Spec maintenance. Who updates the spec when the code evolves? If the answer is "nobody", SDD becomes a dead documentation layer. Böckeler notes that the maintenance strategy is often left vague by the tools (martinfowler.com).
The reasonable conclusion: keep the best of SDD — the spec as a verifiable contract — without taking on its documentation heaviness. Short specs, anchored in the backlog, linked to user stories and their acceptance criteria. The rest is Markdown that serves no one.