research(docs): reinforce documentation, enforce its maintenance under agentic development, and shape a knowledge base for agents #208

Open
opened 2026-10-02 15:49:30 +00:00 by gabogg · 1 comment
Owner

Goal

Research how to keep this repo's documentation accurate as the code changes, especially under agentic development, and how to shape it into a knowledge base that development agents and a future in-product AI integration can consult.

Three threads:

  1. Reinforcing documentation: making docs harder to get wrong and easier to trust (structure, single sources, generated sections, cross-links that are checked).
  2. Enforcing maintenance on the go: catching doc drift when it happens, in the same commit, PR or review pass, rather than in a later audit. Agents are the main risk: they change code fluently and leave docs, plans and PR descriptions behind.
  3. Knowledge bases for agents and AI features: what a development agent, or an early AI integration in the product, should be able to look up, in what form, and how freshness is guaranteed.

Evidence from recent work (2026-10-02 reviews)

Drift that review passes caught by hand, all from agent-written changes:

  • Stale PR descriptions saying "placeholder, implementation has not started" after implementation landed (#177, #178, #168), and stale checklists and doc-check counts (#169).
  • Plan and mapping docs that misdescribe the change: a priority contract described as a comment after it became code (#169); a citation of a test that doesn't exist (#168); items already done on master credited to the PR (#167).
  • The glossary disagreeing with the code: CONTEXT.md says "cycle D+1's boundary", while the spec and code use the civil date + 1 (#177).
  • Avoided glossary terms reintroduced in new schema descriptions: "portal group" and "business Zone" (#168).
  • ADRs that contradict their own PR's behaviour (#177, "history is never re-cut").

What exists today

  • CONTEXT.md (glossary), docs/adr/, docs/agents/* and AGENTS.md, plus vendored skills under .claude/skills/.
  • scripts/check_docs.py in CI and pre-commit, which checks local links and HTTP routes.
  • The two-axis review passes (Standards and Spec) posted through scripts/wt review.
  • Memory files for agent sessions (outside the repo).

Questions to research

  • Mechanical checks: which kinds of drift can a script catch?
    • doc references to symbols, files and tests that no longer exist;
    • glossary Avoid terms appearing in code, schemas or docs;
    • ADR status consistency;
    • PR-description sections that are still templated;
    • "code in area X changed, but its owning doc didn't" (a doc ownership map).
  • Review-time enforcement: should the review passes gain an explicit documentation axis or checklist, and what does it check?
  • Agent guidance: where should instructions live so agents actually follow them? This covers the size and layering of AGENTS.md, skills, hooks and per-directory notes. How do we stop agents committing process noise into docs?
  • Single sources and generated docs: which facts should be generated from code rather than written: API catalog, config keys, enum and label lists, glossary term usage?
  • Knowledge base shape:
    • plain Markdown in the repo, an index or search layer, or a server agents query (e.g. MCP);
    • what content to include (glossary, ADRs, API, runbooks, domain rules) and what never to expose, such as production data or credentials;
    • how freshness is guaranteed.
  • Consumers: development agents vs a future in-product assistant. Do they need the same knowledge base or different ones?
  • Prior art: documentation-as-code practices, docs linting tools, how other agent-heavy repos keep docs honest. Use primary sources.

Outcome

  • A findings document in the repo, under docs/research/ or wherever the repo convention settles. It should give a recommendation per thread and say what fits this repo's size.
  • Follow-up issues for the chosen mechanisms. This issue changes no code by itself.

Not tied to a milestone.

🤖 Generated with Claude Code

## Goal Research how to keep this repo's documentation accurate as the code changes, especially under **agentic development**, and how to shape it into a **knowledge base** that development agents and a future in-product AI integration can consult. Three threads: 1. **Reinforcing documentation:** making docs harder to get wrong and easier to trust (structure, single sources, generated sections, cross-links that are checked). 2. **Enforcing maintenance on the go:** catching doc drift when it happens, in the same commit, PR or review pass, rather than in a later audit. Agents are the main risk: they change code fluently and leave docs, plans and PR descriptions behind. 3. **Knowledge bases for agents and AI features:** what a development agent, or an early AI integration in the product, should be able to look up, in what form, and how freshness is guaranteed. ## Evidence from recent work (2026-10-02 reviews) Drift that review passes caught by hand, all from agent-written changes: - **Stale PR descriptions** saying "placeholder, implementation has not started" after implementation landed (#177, #178, #168), and stale checklists and doc-check counts (#169). - **Plan and mapping docs that misdescribe the change:** a priority contract described as a comment after it became code (#169); a citation of a test that doesn't exist (#168); items already done on master credited to the PR (#167). - **The glossary disagreeing with the code**: CONTEXT.md says "cycle D+1's boundary", while the spec and code use the civil date + 1 (#177). - **Avoided glossary terms reintroduced** in new schema descriptions: "portal group" and "business Zone" (#168). - **ADRs that contradict their own PR's behaviour** (#177, "history is never re-cut"). ## What exists today - `CONTEXT.md` (glossary), `docs/adr/`, `docs/agents/*` and `AGENTS.md`, plus vendored skills under `.claude/skills/`. - `scripts/check_docs.py` in CI and pre-commit, which checks local links and HTTP routes. - The two-axis review passes (Standards and Spec) posted through `scripts/wt review`. - Memory files for agent sessions (outside the repo). ## Questions to research - **Mechanical checks:** which kinds of drift can a script catch? - doc references to symbols, files and tests that no longer exist; - glossary _Avoid_ terms appearing in code, schemas or docs; - ADR status consistency; - PR-description sections that are still templated; - "code in area X changed, but its owning doc didn't" (a doc ownership map). - **Review-time enforcement:** should the review passes gain an explicit documentation axis or checklist, and what does it check? - **Agent guidance:** where should instructions live so agents actually follow them? This covers the size and layering of `AGENTS.md`, skills, hooks and per-directory notes. How do we stop agents committing process noise into docs? - **Single sources and generated docs:** which facts should be generated from code rather than written: API catalog, config keys, enum and label lists, glossary term usage? - **Knowledge base shape:** - plain Markdown in the repo, an index or search layer, or a server agents query (e.g. MCP); - what content to include (glossary, ADRs, API, runbooks, domain rules) and what never to expose, such as production data or credentials; - how freshness is guaranteed. - **Consumers:** development agents vs a future in-product assistant. Do they need the same knowledge base or different ones? - **Prior art:** documentation-as-code practices, docs linting tools, how other agent-heavy repos keep docs honest. Use primary sources. ## Outcome - A findings document in the repo, under `docs/research/` or wherever the repo convention settles. It should give a recommendation per thread and say what fits this repo's size. - Follow-up issues for the chosen mechanisms. This issue changes no code by itself. Not tied to a milestone. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Author
Owner

This was generated by AI during triage.

Triage decision: needs-triage → ready-for-agent

Why it is ready. The goal has three clear threads. The evidence is concrete: drift found by hand in PRs #167, #168, #169, #177 and #178. The current tooling is listed, and the outcome is defined as a findings document plus follow-up issues, with no code change. An agent can do all of it unattended from the repo, the PR history and primary sources. Deciding which mechanisms to adopt stays with the maintainer, through the follow-up issues.

What was settled in triage:

  1. Document location. The issue said "under docs/research/ or wherever the repo convention settles", and no convention existed. It is now docs/research/<issue number>-<slug>.md, shared with #217, #199 and #235.
  2. Thread 3 is kept small. The knowledge base for a future in-product assistant is speculative next to threads 1 and 2, which have evidence. Give it a short "options and what fits our size" section, not a design.
  3. No milestone. This is a single issue with a single PR.

Agent Brief

Category: enhancement, with research (needs investigation before it can be specified further)
Summary: Find out how to keep this repo's documentation accurate under agentic development, and how to shape it into a knowledge base for agents. Deliver a findings document with a recommendation per thread, and follow-up issues.

Current behavior:

  • Documentation today: CONTEXT.md (glossary with Avoid terms), ADRs in docs/adr/, agent docs in docs/agents/, AGENTS.md, and vendored skills under .claude/skills/.
  • Mechanical enforcement: only scripts/check_docs.py, which runs in CI and pre-commit and checks local links and HTTP routes.
  • Review-time enforcement: the two-axis review passes (Standards and Spec), posted through scripts/wt review, with no explicit documentation axis.
  • The gap: drift is caught only by hand in review: stale PR descriptions, plan docs that misdescribe the change, a glossary contradicting code, Avoid terms reintroduced, ADRs contradicting their own PR. The issue body lists the cases.

Desired behavior:

  1. A findings document at docs/research/208-documentation-and-agent-knowledge.md, organised by the three threads (reinforcing, enforcing on the go, knowledge base), with a recommendation per thread sized for this repo.

  2. A drift catalogue. For each drift case in the issue's evidence list, read the PR, classify the drift, and say which proposed mechanism would have caught it, if any: a script, a review checklist item, agent guidance, or a generated doc. Mechanisms that catch nothing from the real evidence should be ranked down.

  3. For each proposed mechanical check, the issue lists these candidates:

    • stale references to symbols, files or tests;
    • glossary Avoid terms appearing in code, schemas or docs;
    • ADR status consistency;
    • PR descriptions still templated;
    • a doc-ownership map ("area X changed, its doc didn't").

    For each: estimate its signal on current master with a throwaway script run from the scratchpad (not committed), and report hits, true positives, false positives and the expected maintenance cost.

  4. Agent guidance: recommend where instructions should live and how big they should be: AGENTS.md vs docs/agents/* vs skills vs hooks vs per-directory notes. Explain how to stop agents committing process noise (review rounds, scratch notes) into durable docs.

  5. Single sources: list which facts should be generated from code rather than hand-written: API catalog, config keys, enum and label lists, glossary term usage. Say how each would be generated and checked.

  6. Knowledge base: compare plain Markdown, an index or search layer, and a queryable server (e.g. MCP). Cover what to include, what must never be exposed (production data, credentials, personal data of cardholders or staff), and how freshness is guaranteed. Then say whether development agents and a future in-product assistant need the same base.

  7. Prior art from primary sources (tool docs, the projects' own docs), cited inline: docs-as-code practice, docs linters (e.g. Vale, markdownlint, lychee), and how agent-heavy repos keep docs honest.

  8. Follow-up issues for each recommended mechanism, labelled needs-triage (adoption is the maintainer's choice), each linked from the doc.

Key interfaces / places to look (by concept):

  • The docs check script and the CI and pre-commit hooks that run it: the existing enforcement seam to extend.
  • The review pass template and scripts/wt review: where a documentation axis or checklist would go.
  • The CONTEXT.md Avoid entries: the input for an Avoid-term linter.
  • PR descriptions and review threads of #167, #168, #169, #177 and #178 (tea pr <N> --comments, scripts/wt review list/show <N>): the evidence.

Acceptance criteria:

  • docs/research/208-documentation-and-agent-knowledge.md exists, with one section per thread and a recommendation for each.
  • The drift catalogue covers every case in the issue's evidence list, maps each one to the mechanism that would have caught it, and links the PR.
  • Each proposed mechanical check reports a hit / false-positive estimate on current master, with the master SHA used.
  • Every non-repo claim cites a primary source.
  • The knowledge-base section states the never-expose list and a freshness mechanism.
  • Follow-up issues exist for each recommended mechanism and are linked from the doc.
  • If docs/research/ doesn't exist yet when this lands: add docs/research/README.md (a one-line index per document) and a link to it from docs/README.md. If #217's or another research PR got there first, rebase and add a line to its index.
  • scripts/check_docs.py passes, and the change goes through a PR (scripts/wt new docs/<slug>), never straight to master.

Out of scope:

  • Implementing any check, hook, generator or knowledge-base server. Those are follow-ups.
  • Rewriting AGENTS.md, CONTEXT.md, ADRs or skills, beyond the index link to docs/research/.
  • Designing the in-product AI assistant itself.
  • Fixing the individual drift cases found. Note them in the catalogue; file an issue only if one is still wrong on master.
> *This was generated by AI during triage.* ## Triage decision: `needs-triage` → `ready-for-agent` **Why it is ready.** The goal has three clear threads. The evidence is concrete: drift found by hand in PRs #167, #168, #169, #177 and #178. The current tooling is listed, and the outcome is defined as a findings document plus follow-up issues, with no code change. An agent can do all of it unattended from the repo, the PR history and primary sources. Deciding which mechanisms to adopt stays with the maintainer, through the follow-up issues. **What was settled in triage:** 1. **Document location.** The issue said "under `docs/research/` or wherever the repo convention settles", and no convention existed. It is now `docs/research/<issue number>-<slug>.md`, shared with #217, #199 and #235. 2. **Thread 3 is kept small.** The knowledge base for a *future* in-product assistant is speculative next to threads 1 and 2, which have evidence. Give it a short "options and what fits our size" section, not a design. 3. **No milestone.** This is a single issue with a single PR. --- ## Agent Brief **Category:** enhancement, with `research` (needs investigation before it can be specified further) **Summary:** Find out how to keep this repo's documentation accurate under agentic development, and how to shape it into a knowledge base for agents. Deliver a findings document with a recommendation per thread, and follow-up issues. **Current behavior:** - **Documentation today:** `CONTEXT.md` (glossary with *Avoid* terms), ADRs in `docs/adr/`, agent docs in `docs/agents/`, `AGENTS.md`, and vendored skills under `.claude/skills/`. - **Mechanical enforcement:** only `scripts/check_docs.py`, which runs in CI and pre-commit and checks local links and HTTP routes. - **Review-time enforcement:** the two-axis review passes (Standards and Spec), posted through `scripts/wt review`, with no explicit documentation axis. - **The gap:** drift is caught only by hand in review: stale PR descriptions, plan docs that misdescribe the change, a glossary contradicting code, *Avoid* terms reintroduced, ADRs contradicting their own PR. The issue body lists the cases. **Desired behavior:** 1. A findings document at **`docs/research/208-documentation-and-agent-knowledge.md`**, organised by the three threads (reinforcing, enforcing on the go, knowledge base), with a recommendation per thread sized for this repo. 2. **A drift catalogue.** For each drift case in the issue's evidence list, read the PR, classify the drift, and say which proposed mechanism would have caught it, if any: a script, a review checklist item, agent guidance, or a generated doc. Mechanisms that catch nothing from the real evidence should be ranked down. 3. **For each proposed mechanical check**, the issue lists these candidates: - stale references to symbols, files or tests; - glossary *Avoid* terms appearing in code, schemas or docs; - ADR status consistency; - PR descriptions still templated; - a doc-ownership map ("area X changed, its doc didn't"). For each: estimate its signal on **current master** with a throwaway script run from the scratchpad (not committed), and report hits, true positives, false positives and the expected maintenance cost. 4. **Agent guidance:** recommend where instructions should live and how big they should be: AGENTS.md vs `docs/agents/*` vs skills vs hooks vs per-directory notes. Explain how to stop agents committing process noise (review rounds, scratch notes) into durable docs. 5. **Single sources:** list which facts should be generated from code rather than hand-written: API catalog, config keys, enum and label lists, glossary term usage. Say how each would be generated and checked. 6. **Knowledge base:** compare plain Markdown, an index or search layer, and a queryable server (e.g. MCP). Cover what to include, what must never be exposed (production data, credentials, personal data of cardholders or staff), and how freshness is guaranteed. Then say whether development agents and a future in-product assistant need the same base. 7. **Prior art from primary sources** (tool docs, the projects' own docs), cited inline: docs-as-code practice, docs linters (e.g. Vale, markdownlint, lychee), and how agent-heavy repos keep docs honest. 8. Follow-up issues for each recommended mechanism, labelled `needs-triage` (adoption is the maintainer's choice), each linked from the doc. **Key interfaces / places to look (by concept):** - The docs check script and the CI and pre-commit hooks that run it: the existing enforcement seam to extend. - The review pass template and `scripts/wt review`: where a documentation axis or checklist would go. - The `CONTEXT.md` *Avoid* entries: the input for an *Avoid*-term linter. - PR descriptions and review threads of #167, #168, #169, #177 and #178 (`tea pr <N> --comments`, `scripts/wt review list/show <N>`): the evidence. **Acceptance criteria:** - [ ] `docs/research/208-documentation-and-agent-knowledge.md` exists, with one section per thread and a recommendation for each. - [ ] The drift catalogue covers every case in the issue's evidence list, maps each one to the mechanism that would have caught it, and links the PR. - [ ] Each proposed mechanical check reports a hit / false-positive estimate on current master, with the master SHA used. - [ ] Every non-repo claim cites a primary source. - [ ] The knowledge-base section states the never-expose list and a freshness mechanism. - [ ] Follow-up issues exist for each recommended mechanism and are linked from the doc. - [ ] If `docs/research/` doesn't exist yet when this lands: add `docs/research/README.md` (a one-line index per document) and a link to it from `docs/README.md`. If #217's or another research PR got there first, rebase and add a line to its index. - [ ] `scripts/check_docs.py` passes, and the change goes through a PR (`scripts/wt new docs/<slug>`), never straight to master. **Out of scope:** - Implementing any check, hook, generator or knowledge-base server. Those are follow-ups. - Rewriting `AGENTS.md`, `CONTEXT.md`, ADRs or skills, beyond the index link to `docs/research/`. - Designing the in-product AI assistant itself. - Fixing the individual drift cases found. Note them in the catalogue; file an issue only if one is still wrong on master.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
gabogg/hikcentral#208
No description provided.