research(tests): faster test suite and right-sized testing for agentic workflows #217
Labels
No labels
blocked
bug
enhancement
high-priority
low-priority
needs-info
needs-triage
ready-for-agent
ready-for-human
referenced
research
wontfix
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
gabogg/hikcentral#217
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Goal
Make testing fast for agentic development. This project is developed by several agents working in parallel: a coordinating agent reviews and up to about 5 coding agents implement. A maintainer supervises them and handles real-world interactions.
Agents run the full test suite many times per task, and that now takes a significant share of every run. Research two things:
Facts (measured 2026-10-02 on master
a7e89c6)The suite:
node --testfiles.test_counting_integrity::test_gap_across_cycle_rollover_is_published_exactly_once: 4.4 s;test_wt_merge::test_no_actions_run_fails_after_the_start_grace: 2.7 s, a realsleep;test_frontend_modulessyntax checks: 2–3 s;pytest-xdistis not installed.tests/conftest.py, as AGENTS.md requires: "never mock production state when in-memory sqlite or ASGI test clients can test real flows".Where the suite runs:
always_run: true,pass_filenames: false), which is about 2.2 min per commit, including docs-only commits.Questions to research
sleeps in tests (for example thewt mergegrace test'sWT_CI_START_GRACE).pytest-xdist(or similar) safe with the SQLite isolation and the async fixtures? What speed-up does it give on this machine and in CI?pytest-testmon, coverage-based selection, or a hand-kept map from paths to tests), with the full suite still enforced in CI. How safe is each approach in this codebase?wtsubcommand (e.g.scripts/wt test [--changed|--full]), that picks the right scope and prints compact output (failures only, like RTK).node --testsuite isn't in pre-commit or CI as far as current docs say. Check that, and decide where it belongs.Outcome
researchconvention), with measured before and after numbers for each recommended change.AGENTS.mdanddocs/standards/git-and-workflow.md.Related: #216 (network-timeout tests), #208 (documentation and agent-knowledge research), #157 (slow statistics routes). Not tied to a milestone.
🤖 Generated with Claude Code
Tracked in draft PR #218 (fix/test-isolation), which delivers the findings document and the safe quick wins; policy changes wait for maintainer sign-off here.
Maintainer note for the research: remote test execution as a primary option
This is a point to evaluate when the research is done, not a finding.
The maintainer has a VPS that is much stronger than the development laptop. Running the test suite remotely, as the default for agents, could beat any amount of local optimisation, so the research should treat "engineer a way to run tests on the remote server" as a first-class option rather than an afterthought.
What already exists:
vps-proxy, reachable over WireGuard) already hosts Forgejo and the Forgejo Actions runner (pythonlabel,hikcentral-ciimage) that runs CI for every PR push.Questions to answer:
The mechanism:
teaon a pushed WIP ref;rsyncorgit pushof the worktree state to a scratch area on the VPS, then running pytest overssh;Compare latency, from command to first failure, against local runs.
Uncommitted work: agents often test before committing. How does the working tree reach the server: a diff, an rsync, or a temporary commit pushed to a scratch ref?
Concurrency: about five agents testing at once. Do separate checkouts and DBs per job isolate them, and how many parallel jobs can the VPS run (CPU and RAM)? Does it still help when CI runs on the same box?
The interface: one entry point (e.g.
scripts/wt test --remote [--changed|--full]) with compact output (failures only), the same result format as a local run, and a clear local fallback when the VPS or VPN is unreachable.Safety:
What it means for policy:
Measure:
pytest-xdist;🤖 Generated with Claude Code
Triage decision:
needs-triage→ready-for-agentWhy it is ready. The questions are concrete. The baseline is measured (master
a7e89c6: 520 passed, 1 skipped, 134 s). The outcome is defined: a findings document plus follow-up issues. The one thing that needs judgement, changing when tests run or the "all tests pass" invariant, is already fenced off as maintainer sign-off. An agent can do everything else unattended on the dev machine.What changed in triage:
ready-for-humanfor that reason. Here, the findings document only names remote execution as an option and links #235. Don't benchmark the VPS in this issue.docs/research/<issue number>-<slug>.md. The same rule applies to #208, #199 and #235.fix/test-isolation), which also closes #216 and #210. This brief is the contract for #218's #217 part.Agent Brief
Category: enhancement, with
research(needs investigation before it can be specified further)Summary: Measure and cut the test suite's wall time, and recommend how much testing agents should run and when, in a findings document with before/after numbers.
Current behavior:
node --testfiles exists. Whether pre-commit or CI runs them is unverified.Desired behavior:
docs/research/217-test-suite-speed.md. It answers each of the issue's questions 1–4, in order, with a recommendation per question that fits a repo this size.needs-triagein this milestone. These cover test selection by change, fast/slow tiers, the pre-commit hook's scope, docs-only commits skipping pytest, ascripts/wt testentry point, tree-hash result caching, and recording the tests run in review passes. Mechanical follow-ups with no open decision may beready-for-agent.node --testfiles today, and where should they run?Key interfaces / places to look (by concept):
always_run,pass_filenames) and the CI workflow's test job. Read them to describe the current state, but don't change them.scripts/wt merge's CI start-grace setting (WT_CI_START_GRACE), the source of the real sleep in its test.How to work:
scripts/wt pr 218). Land #216's network guard and #210's fixes first, so every "after" number is measured on a hermetic suite.-n autoat least 5 times in a row, all green. Then give the speed-up on this machine, and estimate it for CI from the runner's core count.Acceptance criteria:
docs/research/217-test-suite-speed.mdexists and answers questions 1–4, each with a recommendation.--durations=25report are in the doc and the PR.-n autoruns are recorded, and the dependency is added to the test requirements. If not, the doc says why.node --testquestion is answered with evidence (the config lines that do or don't run them).docs/research/doesn't exist yet when this lands: adddocs/research/README.md(a one-line index per document) and a link to it fromdocs/README.md. If #208's or another research PR got there first, rebase and add a line to its index.scripts/check_docs.pyand the full pytest suite pass.Out of scope:
gabogg referenced this issue2026-10-03 12:06:23 +00:00
Decision record: #217 maintainer design interview
Posted at the maintainer's explicit request to make the #217 interview checkable from the tracker and supply the decision source requested by PR #265 review pass 1, r50, P2-1. This is an agent-prepared transcript summary: quoted answers are the maintainer's words from the conversation, not statements independently authored by the agent. Questions are condensed; accepted recommendations are reproduced in the resulting rules.
The interview concerned the policy following research PR #218. The agreed policy is recorded in documentation PR #265; this comment records its authority, not implementation of the runner or hook. PR #265 is not edited by this action.
D1 — Local versus remote full-suite enforcement (Q1)
Question: Must an agent run the full suite locally before submitting a revision, or can remote CI supply the full-suite gate?
Maintainer answer:
The maintainer also required preventing an agent from forcing a merge when CI fails after the final review.
Resulting rule — confirmed: Local testing may be scoped to the change. Remote CI independently runs the full suite. Completing the final review never bypasses CI: a failed, pending or missing CI result cannot authorize a merge. The successful CI result must apply to the revision being merged.
D2 — Evidence-based reuse and honest reporting (Q2)
Question: Can agents reuse a previous successful execution when the tested inputs have not changed?
Maintainer answer:
Resulting rule — confirmed: Reuse qualifying completed passing evidence, and explicitly report that a previous result was reused, with its evidence, rather than claiming another execution. Failed or incomplete executions are not successful evidence. The concrete matching boundaries were settled in D5, D9 and D12–D14.
D3 — Measure concurrent demand rather than assume a scheduling optimum (Q3)
Question: Should optimization prioritize one fast full-suite run or concurrent agent throughput?
Maintainer answer:
The maintainer explicitly expressed uncertainty about whether sequential execution would be faster.
Resulting rule — confirmed: Reduce unnecessary scope and measure CPU use and concurrent request performance before choosing scheduling limits. No claim that sequential execution, unrestricted parallel execution, or a particular worker count is optimal was confirmed. D15 settles the initial release behavior while that study remains outstanding.
D4 — Evidence-aware commit hook; repeat lint and formatting (Q4)
Question: Should the local hook select/run relevant tests itself, or leave execution to the agent?
Maintainer answer:
Resulting rule — confirmed: The hook checks for applicable completed passing evidence and runs the selected relevant tests when that evidence is absent. An earlier agent execution counts only if it satisfies the matching and coverage rules. Lint and formatting continue to run on every hook invocation.
D5 — Conservative complete-snapshot matching first (Q5)
Question: Initially require identical complete source snapshots, or reuse results across changes declared unrelated to the selected tests?
Maintainer answer:
The recommendation was identical complete source snapshots, test scope, command and environment, with dependency-aware reuse deferred until its invalidation rules are validated.
Resulting rule — confirmed: Start with conservative complete-snapshot matching plus equivalent test command/settings and environment. Any source or test-content change invalidates earlier evidence under this initial policy. Dependency-aware reuse across different snapshots is deferred, not approved as an initial implementation shortcut. D9 permits coverage-equivalent superset evidence without relaxing snapshot matching.
D6 — Shared local evidence outside isolated worktrees (Q6)
Question: Begin with coordination on one machine, or immediately share results across local and remote machines?
Maintainer answer:
The recommendation was a shared local directory outside individual worktrees, atomic records, environment identity, and an interface that can accommodate remote storage later.
Resulting rule — confirmed: Begin with shared local evidence storage outside individual worktrees; publish completed records atomically and retain environment identity. Cross-machine storage/service implementation is future work. Storage coordination does not override D7's independent-execution decision.
D7 — Independent concurrent executions (Q7)
Question: If matching tests are already running for another agent, should a second request join that execution?
Maintainer answer:
Resulting rule — confirmed: Concurrent requests execute independently; they do not join or wait for an in-progress result. Preserve separate execution records. Completed matching passing results remain eligible for reuse. Changed test contents already yield different snapshot identities; differing commands, parameters or environments also require distinct matching conditions.
D8 — Explicit test groups and conservative full fallback (Q8)
Question: How should “relevant tests” be delimited, particularly for shared or unknown changes?
Maintainer answer: The maintainer first asked:
After an explicit boundary proposal was presented, the maintainer answered:
That round covered the clarified Q8 mapping, Q10 investigation and Q12 staged-content rules.
Resulting rule — confirmed: Use a versioned explicit mapping from changed paths to test groups, including indirect consumers; combine groups for changes spanning multiple areas. Shared initialization, configuration, schemas, dependencies and test infrastructure select the full local suite. Unmapped paths select the full suite until their scope is documented. Validate the mapping against recent changes before enabling selection.
The agreed starting groups were door/access-cycle/reconciliation/API; combined occupancy/statistics for occupancy, calibration, analytics or schedules; and frontend/static-asset/i18n/relevant-security tests for frontend changes. These are boundaries to validate, not a complete proven path map or a claim of full-suite-equivalent coverage. Naive filename matching was not accepted as sufficient.
D9 — A full pass covers a subset only with unchanged inputs (Q9)
Question: Can an earlier full-suite pass satisfy a later subset request?
Maintainer answer: The maintainer challenged the original example:
The clarification was that any intervening source/test change invalidates the pass; reuse applies only to repeated requests without changed inputs, such as a manual execution followed by the hook. The maintainer then answered:
Resulting rule — confirmed: A full pass predating a modification cannot validate the modified subset. A matching full-suite result may satisfy a subset request only if the unchanged snapshot/environment/settings match and every requested test demonstrably executed. A request alone does not change the snapshot. A subset pass cannot certify the full suite.
D10 — Investigate conflicting results before choosing permanent handling (Q10)
Question: What should happen when independent executions apparently sharing an identity produce a pass and a failure?
Maintainer answer:
When the question was refined to preserving separate records and comparing input differences, edits during execution, interference and nondeterminism, the maintainer answered:
Resulting rule — confirmed: Preserve separate records and investigate the actual causes of disagreement before choosing a permanent conflict policy. Candidate conditions include incorrectly matched inputs/environment, changing files, shared state/resources, test order, clocks and scheduling. #221 supplies a concrete order-dependence case; other hypothesized causes still need evidence.
Proposed, not confirmed: The earlier recommendation to automatically suspend all reuse for an identity until every disagreement is resolved was not accepted as a settled permanent policy. Automatic retry limits, blacklist duration, and pass-versus-failure precedence were not decided.
D11 — A fix after the final review requires a new review and investigation (Q11)
Question: If CI fails after the final review and an agent fixes it, does the new change require another review?
Maintainer answer:
Resulting rule — confirmed: A fix that lands after the final review requires a new review of the fix and why validation missed the failure. Distinguish deliberate omission of local testing from a miss caused by the optimized testing system. If that system caused the miss, file a separate issue for the testing-system defect. Successful CI on the new head and the required new review remain prerequisites for merge.
Proposed, not confirmed: A new review's exact template, pass numbering and whether it must repeat every prior review axis were not separately specified. “Focused” must not mean omitting the required investigation.
D12 — Match the actual staged/post-format tested contents (Q12)
Question: With partial staging and formatter changes, must hook reuse match the contents actually visible after formatting and pre-commit isolation?
Maintainer answer:
This was the round explicitly including Q12.
Resulting rule — confirmed: Yes. Reuse requires the actual tested contents visible after formatting and staging isolation to match. Otherwise run the selected tests again. HEAD or the initial index hash alone is insufficient. A successful execution against unstaged changes does not certify different staged contents.
D13 — Skipped tests and reusable success (Q13)
Question: Can a successful process exit with skipped requested tests qualify for reuse?
Maintainer answer:
This answer covered Q13–Q15. Q13 recommended every requested test pass except explicitly approved skips with recorded conditions.
Resulting rule — confirmed: Every requested test must pass, except explicitly approved skips whose conditions are recorded. Missing executables must not silently count as validation. A zero process exit alone is insufficient evidence of requested coverage.
Proposed, not confirmed: The exact approved-skip list and skip-approval mechanism remain to be defined.
D14 — Isolated immutable execution inputs (Q14)
Question: Execute tests directly in a changing worktree, or against an isolated immutable snapshot?
Maintainer answer:
This answer covered Q13–Q15. Q14 recommended an isolated snapshot containing applicable source, configuration and Git context.
Resulting rule — confirmed: Test an isolated immutable snapshot of the observable inputs, including applicable configuration and Git context. Evidence belongs to that snapshot. Later worktree edits create a different identity and invalidate reuse for the changed contents.
Proposed, not confirmed: The precise snapshot-construction mechanism, fingerprint schema, exhaustive input list, evidence expiry/retention and exclusion rules were not individually approved. They must be designed and validated against the agreed identity rule.
D15 — CPU study does not block the first scoped runner (Q15)
Question: Must CPU/scheduling measurements precede the initial runner release?
Maintainer answer:
This answer covered Q13–Q15. Q15 recommended a separate study, independent requests and existing serial execution within each request initially.
Resulting rule — confirmed: The initial scoped runner need not wait for the CPU study. Keep requests independent and preserve serial execution within each request initially. Measure CPU time/utilization and concurrent completion performance before choosing workers or admission limits. This does not endorse a particular future limit or parallelism default.
D16 — Research closure requires an accurate linked issue handoff
Question: Should research review/merge produce implementation issues rather than treating the findings as delivered implementation?
Maintainer answer:
The maintainer also directed:
In the explicit request for this decision record, the maintainer required inclusion of:
Resulting rule — confirmed: Research does not close until accepted actionable findings have an accurate linked follow-up issue handoff. Existing issues can provide that handoff after their contradictory briefs are corrected; create additional issues for uncovered work. Related open test issues should be linked as examples, distinguishing them from work actually delivered by the research PR. Research closure is not implementation completion.
Proposed, not confirmed as an independent interview decision: The exact handoff checklist and blanket requirement that every follow-up join the research milestone were agent elaborations. Milestone membership follows the current repository rule: an issue/PR joins only if the milestone cannot be completed without it; unrelated review follow-ups normally remain outside. This record does not approve the older blanket membership wording in PR #265.
Consolidated confirmation and scope
After the complete design had been documented, the agent asked whether the consolidated policy captured shared understanding and whether the documentation branch could be pushed. The maintainer answered:
That confirmation establishes the agreed design above; it does not convert unasked implementation details or the explicitly deferred conflict policy into confirmed decisions. The maintainer subsequently requested a separate linked draft PR, examples on #218, and the
ready-for-agentlabel for #217. PR #265 is the resulting documentation draft; #218 remains the research/isolation vehicle.Direct follow-up candidates already exist: #241 (runner), #242 (hook/guidance), #238 (parallel/CI evaluation), #240 (fixture profiling), and #235 (remote execution). Their older requirements are not automatically endorsed by this interview. Reconcile them with the decisions and measured findings before treating them as implementation contracts. Related examples on #218 include #221, #250, #260, #184, #213, #76, #261, #225 and #60.
Maintainer confirmation of D16 (2026-10-03). The decision record above attributes the sentence "research doesn't close until it hands off to accurate, linked follow-up issues" to the maintainer. That wording came from a coordinator prompt, not from the maintainer. Asked directly, the maintainer confirmed the rule in their own words:
D16 is therefore confirmed: performing the research a research PR requests must produce linked follow-up issues for its actionable findings, and the research isn't finished until that handoff exists. Milestone membership of those issues follows the current rule from #266: an issue joins only if the milestone cannot be completed without it. The handoff checklist itself stays proposed, not confirmed.