Repository navigation
[Feature]: Stop grep-based spec alignment, structured IDs + context packs #4164
Description
Activity
- addedfeature-assessRun the Spec Kit idea-assessment pipeline on this feature requestRun the Spec Kit idea-assessment pipeline on this feature request
on Aug 19, 2026 github-actions commented
on Aug 19, 2026 on Aug 19, 2026 – with GitHub ActionsContributorMore actionsFeature assessment — structured-ids-context-packs · Stage 1/5: Intake
Idea Intake: Stop grep-based spec alignment, structured IDs + context packs
- Slug: structured-ids-context-packs
- Created: 2026-08-19
- Source: GitHub issue [Feature]: Stop grep-based spec alignment, structured IDs + context packs #4164 (pasted text)
- Type: improvement
Idea (as captured)
Problem Statement
I'm frustrated when a product change (especially a UI change) has to be folded into an existing feature. Agents find related work by grepping similar wording in spec.md, plan.md, and tasks.md. After a reword "charge on Pay" vs "show a confirm modal" the live AC/FR is often not retrieved, so another /speckit.specify pass or a manual edit adds a new requirement instead of updating the existing one.
The result is two live lines for the same behavior, or two live lines that contradict each other. /speckit.analyze is supposed to catch duplicates and conflicts, but it only reports what the model loaded. If the related line was never in context, the report is clean.
Proposed Solution
- Structured sidecar (tasks.yaml alongside tasks.md) — each task record: id, status, story, files, covers: [US1/AC2, FR-007]. Python script is the only mutator.
- Complete live inventory — script emits every live FR-/AC/SC/T- for the feature; specify consults before adding; analyze builds from it; implement loads one task plus covered records.
- Context pack (
specify context --task T014) — small JSON pack: that task + live ACs/FRs + matching plan bullets. Commands drive from pack/inventory, not whole Markdown dumps. - Local per-feature recall (opt-in phase 2) — local embedding index under
specs/<feature>/.index/. - Classify-on-write alignment — before adding a requirement: already-true → skip; same behavior new words → edit existing; contradicts live → conflict; genuinely new → one new ID.
Restated
The current Spec Kit spec-alignment workflow relies on agents grepping prose in three Markdown files, which silently misses renamed or paraphrased requirements and creates duplicate or contradictory live lines. The idea is to replace that grep-based discovery with a script-owned structured inventory of requirement and task IDs, plus a focused "context pack" command, so every SDD command operates on the full live set rather than a partial token-budget slice.
Origin & Context
- Raised by: harsha09
- Trigger: Practical frustration with spec alignment breaking down when UI copy changes rephrase existing requirements; /speckit.analyze not catching duplicates/conflicts reliably
First-Glance Unknowns
- [NEEDS CLARIFICATION: Does the proposal intend tasks.yaml to fully replace tasks.md as the writeable record, or exist alongside it as a derived artifact?]
- [NEEDS CLARIFICATION: Which commands are in scope for Phase 1 — specify, analyze, implement only, or also tasks, plan, converge, checklist?]
- [NEEDS CLARIFICATION: What is the backward-compatibility story for existing projects with no tasks.yaml?]
- [NEEDS CLARIFICATION: How does the live inventory handle multi-feature projects where IDs might collide across feature directories?]
- [NEEDS CLARIFICATION: Is the local embedding index (Phase 2) expected to be stored in-repo or excluded via .gitignore?]
- [NEEDS CLARIFICATION: What is the target scope for "context pack" — single task, feature, or cross-feature?]
Generated by 💡 Assess a Feature Request by Installing and Running Spec Kit for issue #4164 · 471.8 AIC · ⌖ 28.7 AIC · ⊞ 38.1K · ◷
github-actions commented
on Aug 19, 2026 on Aug 19, 2026 – with GitHub ActionsContributorMore actionsFeature assessment — structured-ids-context-packs · Stage 2/5: Research
Idea Research: Stop grep-based spec alignment, structured IDs + context packs
- Slug: structured-ids-context-packs
- Created: 2026-08-19
- Evidence confidence (overall): medium
Users & Demand
- Spec Kit users who iterate on UI/product features experience alignment breakage when requirement phrasing evolves — an agent adding "confirm modal" on top of an existing "charge on Pay" requirement creates silent duplicates. — [source: issue [Feature]: Stop grep-based spec alignment, structured IDs + context packs #4164, author harsha09] (confidence: high, cited)
- The issue is self-describing and specific — the reporter articulates a concrete failure mode (grep misses renamed requirements) with a clear symptom (two live lines for the same behavior), suggesting at least one active user hitting it in practice. — [source: issue [Feature]: Stop grep-based spec alignment, structured IDs + context packs #4164] (confidence: medium, cited)
- No usage data or support ticket volume available to quantify prevalence across the broader user base. — [NEEDS CLARIFICATION: How many users report this? Any prior community discussion?] (confidence: low, assumption)
Prior Art
- Internal — analyze command's requirements inventory (
templates/commands/analyze.md): The current/speckit.analyzealready extracts a "requirements inventory" (FR-/SC- IDs) and builds a coverage mapping per run. However, this inventory is ephemeral — reconstructed from markdown each time — rather than a persistent shared artifact. This confirms the ID-based approach is already partially acknowledged, but no cross-command persistence exists. — [source: codebase read] (confidence: high, cited) - Internal — tasks format already has T-IDs and US-labels (
templates/commands/tasks.md): Tasks use T001/T002 sequential IDs and[US1]story labels. There is nocovers:field or tasks.yaml sidecar in the current schema. — [source: codebase read] (confidence: high, cited) - Internal — spec template uses FR-001 identifiers (
templates/spec-template.md): FRs already carry structured IDs in specs; they are just not shared across commands via a script-owned persistent store. — [source: codebase read] (confidence: high, cited) - Internal — check-prerequisites script loads entire markdown files and outputs FEATURE_DIR; no structured sidecar or JSON inventory is produced. — [source:
scripts/bash/check-prerequisites.sh, codebase read] (confidence: high, cited) - External — Linear/JIRA/GitHub Issues all maintain persistent structured issue IDs that tools can query. The proposal applies this well-understood pattern to per-feature spec artifacts. — [source: general knowledge, ASSUMPTION] (confidence: medium, assumption)
- No prior closed Spec Kit issues or merged PRs found in this codebase addressing structured ID persistence. — [source: codebase exploration] (confidence: medium, cited)
Market & Context
- Alternative users rely on today: manually re-running
/speckit.analyzerepeatedly and hoping the model loads the right context window; manually editing spec.md to remove duplicates; staying disciplined about not rewording requirements (fragile). — [source: issue [Feature]: Stop grep-based spec alignment, structured IDs + context packs #4164] (confidence: high, cited) - Cost of doing nothing: duplicate or contradictory live requirements accumulate silently;
/speckit.implementmay build against the wrong or stale requirement; trust in the SDD pipeline erodes. — [source: issue [Feature]: Stop grep-based spec alignment, structured IDs + context packs #4164, codebase analysis] (confidence: high, cited) - The proposal is additive — tasks.yaml alongside tasks.md, context pack as a new command — so backward compatibility is achievable but requires explicit design. — [source: codebase structure analysis] (confidence: medium, cited)
Data & Constraints
- Token budget pressure is real: each command currently runs check-prerequisites which emits FEATURE_DIR and outputs whole markdown files. A context pack reducing that to a focused JSON directly addresses a known LLM context window constraint. — [source:
templates/commands/implement.md,templates/commands/analyze.md] (confidence: high, cited) - Scope is broad: the proposal spans at minimum 5 commands (specify, analyze, implement, tasks, converge), a new Python script (inventory mutator), a new file format (tasks.yaml), and an optional embedding subsystem. Phase 1 alone is medium-to-large appetite. — [source: codebase analysis] (confidence: high, cited)
- Python script consistency: the proposal requires Python for the inventory mutator. Spec Kit already ships Python scripts (
scripts/python/) alongside bash/PowerShell; this is consistent with existing architecture. — [source:scripts/python/, codebase read] (confidence: high, cited)
Evidence Against the Idea
- Complexity vs. benefit ratio is unclear for smaller projects: teams with few features and rarely-renamed requirements would carry the overhead of a YAML sidecar for little gain.
- Two-format maintenance burden: keeping tasks.md and tasks.yaml in sync introduces a new failure mode — if the Python mutator is bypassed, they drift. The enforcement mechanism is agent convention, not a hard constraint.
- Phase 2 (embeddings) adds significant complexity and storage cost with uncertain payoff — the proposal itself makes it opt-in, acknowledging this.
- Existing analyze command already does ID-based coverage mapping per run — the gap is persistence and sharing, which might be solvable with a lighter-weight approach than a full YAML sidecar (e.g., a generated inventory section in tasks.md, or a thin JSON file just for IDs).
Gaps & Open Questions
- [NEEDS CLARIFICATION: What is the migration path for existing projects that have tasks.md but no tasks.yaml?]
- [NEEDS CLARIFICATION: Who owns tasks.yaml on conflict — the human (tasks.md edit) or the script? How is drift detected and resolved?]
- [NEEDS CLARIFICATION: Is "context pack" a new CLI subcommand, a new agent skill, or both?]
- [NEEDS CLARIFICATION: How does multi-feature ID uniqueness work? Are IDs feature-scoped or global?]
- [NEEDS CLARIFICATION: Phase 2 embedding index — which embedding model, and is there a privacy concern for proprietary codebases?]
Sources
- GitHub issue [Feature]: Stop grep-based spec alignment, structured IDs + context packs #4164, [Feature]: Stop grep-based spec alignment, structured IDs + context packs #4164 (host: github.com, policy: allowlisted)
- Codebase reads:
templates/commands/analyze.md,templates/commands/tasks.md,templates/commands/implement.md,templates/spec-template.md,scripts/bash/check-prerequisites.sh,src/specify_cli/(local repo, no URL fetch)
Generated by 💡 Assess a Feature Request by Installing and Running Spec Kit for issue #4164 · 471.8 AIC · ⌖ 28.7 AIC · ⊞ 38.1K · ◷
github-actions commented
on Aug 19, 2026 on Aug 19, 2026 – with GitHub ActionsContributorMore actionsFeature assessment — structured-ids-context-packs · Stage 3/5: Problem
Problem Definition: Stop grep-based spec alignment, structured IDs + context packs
- Slug: structured-ids-context-packs
- Created: 2026-08-19
- Inputs used: intake.md | research.md
Problem Statement
When a Spec Kit user refines or renames a requirement mid-feature (a common occurrence in UI-heavy product work), the SDD pipeline's grep-based context loading silently misses the renamed line, causing agents to add a second live requirement instead of updating the existing one — producing duplicate or contradictory specifications that the current
/speckit.analyzecommand cannot reliably detect because it only reports on what it happened to load.Affected Users & Stakeholders
- Users: Spec Kit users iterating on product/UI features — they experience accumulating duplicate or contradictory FRs/ACs and are forced to rerun
/speckit.analyzerepeatedly without a reliable result. - Stakeholders: Spec Kit maintainers (github/spec-kit) — the reliability of the SDD pipeline is a core product promise; silent spec drift undermines that promise and trust in the toolchain.
Goals
- Every SDD command (at minimum: specify, analyze, implement) can consult a complete, authoritative list of live requirement and task IDs for the current feature before adding or modifying a requirement.
- Adding a requirement is a classified write operation: duplicate → skip, paraphrase → update existing, conflict → block, genuinely new → create new ID.
- A focused "context pack" reduces token waste (no whole-file dumps) and eliminates the missed-line failure mode from partial context windows.
- The approach is backward-compatible: projects without the new sidecar continue to work with existing behavior.
Non-Goals
- Replacing tasks.md as the human-readable artifact — it remains the primary format.
- Global or cross-feature ID namespacing beyond per-feature scoping.
- A hosted or cloud vector store for embeddings (Phase 2 is local and opt-in).
- Changing the Spec Kit spec/plan template formats (FR-001 naming stays).
- Real-time sync or conflict resolution between concurrent agent sessions.
Success Metrics
- Zero silent duplicate FRs/ACs after a requirement rename when using Phase 1 tooling (measurable by test: rename a requirement, run specify, verify no duplicate is added).
- Single-pass
/speckit.analyzereliably detects all coverage gaps and conflicts (no need to rerun; deterministic output on unchanged artifacts). - Context pack token reduction vs. whole-file dump (qualitative target; specific baseline [NEEDS CLARIFICATION: what is the current typical token count for a whole-file dump?]).
- Backward compatibility: existing projects without tasks.yaml continue to work without modification.
Cost of Inaction
Without this, spec quality degrades silently whenever requirements are reworded — a routine event in UI/product development. Agents build against stale or contradictory requirements; the output of
/speckit.analyzebecomes unreliable; developers lose trust in the SDD pipeline and fall back to manual spec hygiene, which is the exact workflow Spec Kit is meant to replace.Open Questions
- [NEEDS CLARIFICATION: What is the right granularity for the live inventory — per-feature only, or cross-feature?]
- [NEEDS CLARIFICATION: How is tasks.yaml kept in sync with tasks.md when a human edits tasks.md directly (the most common path today)?]
- [NEEDS CLARIFICATION: What commands are in scope for Phase 1 changes — specify and analyze only, or all SDD commands?]
- [NEEDS CLARIFICATION: Is there a preferred token budget target for the context pack?]
Generated by 💡 Assess a Feature Request by Installing and Running Spec Kit for issue #4164 · 471.8 AIC · ⌖ 28.7 AIC · ⊞ 38.1K · ◷
github-actions commented
on Aug 19, 2026 on Aug 19, 2026 – with GitHub ActionsContributorMore actionsFeature assessment — structured-ids-context-packs · Stage 4/5: Concept
Concept: Stop grep-based spec alignment, structured IDs + context packs
- Slug: structured-ids-context-packs
- Created: 2026-08-19
- Recommended option: Option B — Persistent ID index + classify-on-write
Options
Option A — Lightweight: Shared ID list in tasks.md
- Sketch: Extend the existing tasks.md format with a machine-readable header block (YAML frontmatter or a dedicated
## Live IDssection) listing all live FR-/AC-/T- IDs. The check-prerequisites script extracts this list and passes it to commands via JSON output. Commands consult the extracted list before adding requirements. No new file format; no new command. - Appetite: small (days)
- Trade-offs: Wins: zero new files, zero format change for humans, backward-compatible. Sacrifices: IDs still embedded in a markdown file; extraction is fragile if humans reformat; the
covers:field (linking tasks to FRs) still has no home. - Rabbit holes: Markdown frontmatter parsing is finicky across editors. Does not solve the context pack token-waste problem.
Option B — Persistent ID index + classify-on-write (recommended)
- Sketch: Alongside tasks.md, a Python script writes and reads
tasks.json— a machine-readable record of all live T-IDs with their status, associated story IDs, andcoverslinks to FR-/AC- IDs. The check-prerequisites script is extended to emit the live ID inventory. A newspecify context --task T014command returns a focused JSON pack: the task record, the FRs/ACs it covers, and matching plan bullets. Before any command adds a requirement, it calls the inventory query to classify the new requirement as duplicate/paraphrase/conflict/new. tasks.md remains the primary human-readable artifact; tasks.json is a derived index rebuilt when tasks.md changes (regenerate-on-read). - Appetite: medium (weeks)
- Trade-offs: Wins: eliminates the grep-based missed-line failure; context pack directly addresses token waste; consistent with existing Python script infrastructure. Sacrifices: adds a new derived file (tasks.json) and a generation step; agents must be updated to call the inventory before writing; introduces a sync dependency between tasks.md and tasks.json.
- Rabbit holes: Keeping tasks.json in sync with tasks.md when humans edit directly is the hardest part — the sync story needs explicit design. The classify-on-write logic ("same behavior, new words") is inherently heuristic without embeddings and could produce false positives/negatives.
Option C — Full structured inventory + local embeddings (Phase 1 + Phase 2 together)
- Sketch: Option B plus a local per-feature embedding index under
specs/<feature>/.index/for paraphrase-pair recall. The embedding index bridges the "confirm modal" ↔ "charge on Pay" gap that string matching misses. Embeddings are generated locally (no cloud); obsolete lines are excluded. - Appetite: large (months)
- Trade-offs: Wins: near-complete paraphrase detection even for radical rewording. Sacrifices: adds a Python embedding dependency (model download, disk storage, inference latency); significantly more complex; embedding quality depends on model choice.
- Rabbit holes: Model selection, in-repo .gitignore hygiene for the index, inference latency in agent loops, cross-platform support.
Recommendation
Option B — Persistent ID index + classify-on-write — addresses the core failure mode within a medium appetite, consistent with Spec Kit's existing Python script architecture. Phase 2 embeddings (Option C) should remain opt-in and are appropriately deferred until Phase 1 proves out the inventory approach.
The key design decision that must be resolved before specification: how tasks.json is kept in sync with tasks.md when a human edits tasks.md directly. Recommend a regenerate-on-read strategy (tasks.json is rebuilt from tasks.md whenever check-prerequisites runs) to eliminate the sync gap without requiring humans to use a script as the only editor.
Out of Scope (for the recommended option)
- Local embedding index / paraphrase recall (Phase 2, Option C)
- Cross-feature or global ID namespacing
- Changes to spec.md or plan.md formats
- Real-time or concurrent-session conflict resolution
- Hosted or cloud-based vector stores
Assumptions to Validate
- The regenerate-on-read approach (rebuild tasks.json from tasks.md on each command invocation) is fast enough to not add noticeable latency to the agent loop.
- The classify-on-write heuristic (string/ID matching without embeddings) is sufficient to catch the majority of duplicate/paraphrase cases in practice.
- Adding a
covers:field to the tasks.md task format is acceptable to existing Spec Kit users and backward-compatible with projects that omit it. - The new
specify contextcommand can be delivered as an additional Spec Kit skill/command without requiring a CLI version bump.
Generated by 💡 Assess a Feature Request by Installing and Running Spec Kit for issue #4164 · 471.8 AIC · ⌖ 28.7 AIC · ⊞ 38.1K · ◷
github-actions commented
on Aug 19, 2026 on Aug 19, 2026 – with GitHub ActionsContributorMore actionsFeature assessment — structured-ids-context-packs · Stage 5/5: Decision — verdict needs-clarification
Decision: Stop grep-based spec alignment, structured IDs + context packs
- Slug: structured-ids-context-packs
- Decided: 2026-08-19
- Verdict: needs-clarification
- Artifacts reviewed: intake.md | research.md | problem.md | concept.md
Scorecard
Criterion Rating Justification Problem validity strong The failure mode (grep-based context loading misses renamed requirements, producing silent duplicates/contradictions) is real, specific, and directly demonstrated. It undermines a core SDD pipeline promise. Evidence strength weak One user report with a clear description, but no usage data, no support ticket volume, no community upvotes or corroborating reports. Codebase analysis confirms the architectural gap is real, but demand signal is single-source. Value vs. inaction adequate Spec drift and unreliable analyze passes erode trust in the SDD pipeline. The cost of inaction compounds as more users iterate on UI features. Impact quantification is absent. Feasibility / appetite adequate Option B is medium appetite, architecturally consistent with existing Python scripts, and does not require new external dependencies. Key unknowns (sync strategy, command scope) are resolvable during specification. Strategic fit adequate Structured IDs and context packs align with Spec Kit's goal of reliable, deterministic SDD. The proposal extends existing patterns (FR-IDs, T-IDs, Python scripts) rather than replacing them. Risk posture adequate Main risks (tasks.md/tasks.json sync drift, classify-on-write false positives) are identified and the recommended regenerate-on-read mitigation is credible. Phase 2 risk is deferred by making it opt-in. Verdict & Rationale
needs-clarification — The problem is real and the concept is credible, but the evidence strength is weak (single user report; no usage data, support ticket volume, or broader demand signal). Per the pipeline's downgrade rule, weak evidence strength blocks a
goverdict regardless of other scores.Additionally, several design decisions that
/speckit.specifywould need to act on remain unresolved: the tasks.md ↔ tasks.json sync strategy (regenerate-on-read is assumed but not validated), the scope of command changes for Phase 1, and the API shape of thespecify contextcommand. These are resolvable questions, not blockers in principle — but they should be answered before specification can produce a reliable spec.This is a promising, internally-consistent idea that deserves a
goonce the evidence base is strengthened and the open design questions are answered.If needs-clarification
- Blocking questions:
- [NEEDS CLARIFICATION: Is there broader demand signal beyond this single report? Are there other issues, discussions, or community requests corroborating this failure mode?]
- [NEEDS CLARIFICATION: What is the exact sync strategy for tasks.json when a human edits tasks.md directly — regenerate-on-read (every command invocation), on-commit hook, or manual?]
- [NEEDS CLARIFICATION: Which SDD commands are in scope for Phase 1 ID inventory integration — specify and analyze only, or also implement, converge, tasks?]
- [NEEDS CLARIFICATION: What is the API shape of the
specify contextcommand — new CLI subcommand, new agent skill, or both?] - [NEEDS CLARIFICATION: Is
covers:intended as a new required field in tasks.md task lines, or optional/additive?]
- Revisit stage: research (to gather broader demand signal); define (to resolve scope and design questions once demand is confirmed)
Generated by 💡 Assess a Feature Request by Installing and Running Spec Kit for issue #4164 · 471.8 AIC · ⌖ 28.7 AIC · ⊞ 38.1K · ◷
- addedfeature-needs-clarificationFeature assessment verdict: needs clarificationFeature assessment verdict: needs clarification
on Aug 19, 2026 I would like to explore a focused contribution for this issue, but the proposed scope is fairly broad. I suggest starting with the smallest Phase 1 slice: a backward-compatible, script-generated requirements/task inventory and a narrow context-pack interface, leaving embeddings and broader command rewrites out of scope.
Before implementing, could a maintainer confirm the preferred Phase 1 boundary, migration behavior for existing projects without the sidecar, and whether this direction is wanted? If approved, I would be happy to prepare the design and implementation and be assigned to the issue.
Disclosure: I used ChatGPT to review the repository guidance and draft this comment.
After a careful review by myself, this doesn't need to go into core as proposed — the whole of Phase 1 can be delivered as an installable, opt-in bundle with zero core changes, which is also the right way to build the adoption signal the assessment flagged as missing.
Concretely, I'd suggest a paired extension + preset:
Extension (
speckit-inventory) — the read-only primitive:- A
scripts/inventory extractor that regenerates every liveFR-/NFR-/SC-/T-ID from the existingspec.md/tasks.mdon each run and prints JSON. Purely additive, no sidecar, no mutation — so thetasks.yaml↔tasks.mdsync problem simply doesn't arise (nothing is a second source of truth). - A namespaced
speckit.inv.contextcommand that returns a focused context pack (a task plus the requirement IDs it covers) instead of whole-file dumps. - Optional
before_specify/before_analyzehooks.
Preset (
inventory-alignment) — wires the primitive into the core flow via template overrides (presets sit above extensions in the resolution stack):- A
wrapoverride ofspeckit.specify/speckit.analyze(using the{CORE_TEMPLATE}placeholder) that prepends a "load inventory, then classify each requirement as already-true / edit-existing / conflict / genuinely-new" pre-pass, then falls through to the original command body. - An
appendoverride oftasks-templateadding the optionalcovers:field.
Please keep embeddings out (that was the large/opt-in Phase 2). This combination gives essentially all of Phase 1 without a core PR, and if it gets real usage we can revisit promoting the inventory primitive into core.
One genuine gap worth noting:
preset.yml'srequires:only declaresspeckit_version, so there's no formal "requires extension X" dependency to guarantee the extension is installed alongside the preset. That's a small, separable core enhancement I'd track on its own issue rather than block this on.We would love to see a PR that will list the community extension and community preset that would support this.
Disclosure: this comment was prepared with GitHub Copilot (model: Claude Opus 4.8) acting on my behalf.
- A
- added a commit that references this issue
on Aug 20, 2026
Problem Statement
I'm frustrated when a product change (especially a UI change) has to be folded into an existing feature. Agents find related work by grepping similar wording in spec.md, plan.md, and tasks.md. After a reword “charge on Pay” vs “show a confirm modal” the live AC/FR is often not retrieved, so another /speckit.specify pass or a manual edit adds a new requirement instead of updating the existing one.
The result is two live lines for the same behavior, or two live lines that contradict each other. /speckit.analyze is supposed to catch duplicates and conflicts, but it only reports what the model loaded. If the related line was never in context, the report is clean. I end up rerunning analyze over and over; it is not a reliable alignment step.
There is no script-owned inventory of live IDs (FR-, US/AC, T-xxx) that specify, analyze, and implement all share. Commands dump whole Markdown files (wasted tokens) and still miss the one line that mattered. I need a way to keep the spec set unambiguous no silent duplicates, no contradictory live requirements without another round of hope-and-grep.
Proposed Solution
Keep spec.md / plan.md / tasks.md as the human artifacts. Add a script-owned live inventory so specify, analyze, and implement stop grepping prose to decide what exists.
Structured sidecar for tasks (and later spec/plan records)
Alongside tasks.md, write tasks.yaml (YAML on disk, JSON on stdout). Each task is a record: id, status, story, files, covers: [US1/AC2, FR-007]. A Python script is the only mutator: parse Markdown once, then list / get / mark-done / coverage. Agents do not regex checkboxes.
Complete live inventory, every command
The script emits every live FR- / AC / SC / T- for the current feature. That list is the source of “what exists.” Specify consults it before adding a requirement. Analyze builds its report from it. Implement loads one task plus the records it covers. No command should discover the spec by grepping similar text.
Context pack instead of whole files
specify context --task T014 (or equivalent) prints a small JSON pack: that task, the live ACs/FRs it cites, matching plan bullets, and optional research hits. /speckit.implement and /speckit.analyze drive from the pack/inventory, not from dumping three Markdown files. That removes the missed-line failure and cuts tokens.
Local, per-feature recall only (optional second phase)
If an ID link is missing, a local embedding index under specs//.index/ can pair paraphrases (“confirm modal” <-> live “charge on Pay”). Not a global or hosted vector store. Default retrieval excludes obsolete lines so old UI copy cannot be treated as live. Embeddings improve recall; they are not the source of truth.
Alignment is classify-on-write, not another analyze loop
Before adding a requirement, resolve against the full live set (IDs, then local similarity):
/speckit.analyze then reports coverage holes, unpaired IDs, and remaining high-score pairs over that structure. It should be a single deterministic pass, not a hunt that has to be rerun until the model gets lucky.
Phase 1 (core-shaped): parser + tasks.yaml + covers + context pack.
Phase 2 (opt-in): local per-feature index for paraphrase pairs.
Alternatives Considered
No response
Component
Other
AI Agent (if applicable)
None
Use Cases
No response
Acceptance Criteria
No response
Additional Context
No response