research(tests): run the test suite remotely on the VPS as the agents' default #235

Open
opened 2026-10-03 07:43:50 +00:00 by gabogg · 0 comments
Owner

This was generated by AI during triage.

Blocked by: #218

Split from #217 during triage (2026-10-03). It carries the maintainer's note of 2026-10-03 on #217, "remote test execution as a primary option", so #217 and its PR #218 can ship the local findings and quick wins without waiting on VPS work.

Question

Can agents run the test suite on the VPS (vps-proxy, over WireGuard) by default, and does that beat local runs on the laptop once #217's local optimisations have landed?

The VPS already hosts Forgejo and the Forgejo Actions runner (python label, hikcentral-ci image), so a remote Python environment for this suite exists. What's open is how an agent triggers it during work, without a push and a full CI run.

Questions to answer

  1. Mechanism. Compare these on latency from command to first failure, against a local run:
    • a dedicated "test now" workflow dispatched with tea on a pushed WIP ref;
    • rsync or git push of the worktree state to a scratch area on the VPS, then pytest over ssh;
    • a long-lived remote worker that agents queue test jobs to.
  2. Uncommitted work. Agents test before committing. How does the working tree reach the server: a diff, an rsync, or a temporary commit pushed to a scratch ref? What gets missed (untracked files, ignored files)?
  3. Concurrency. About five agents testing at once. Do separate checkouts and DBs per job isolate them? How many parallel jobs fit in the VPS's CPU and RAM? Does it still help when CI runs on the same box?
  4. Interface. One entry point (e.g. scripts/wt test --remote [--changed|--full]) with compact output (failures only), the same result format as a local run, and a clear local fallback when the VPS or VPN is unreachable.
  5. Safety.
    • Remote runs stay hermetic: no production data and no secrets beyond what CI already has.
    • No reach into HikCentral. This depends on #216's session-level network guard.
    • Scratch checkouts are cleaned up afterwards.
  6. Policy implications. Recommendations only; each needs the maintainer's sign-off.
    • Could the pre-commit hook delegate to the remote run?
    • Does a remote pass on the exact tree hash count toward the AGENTS.md rule that all tests must pass?
    • Can CI reuse that result instead of re-running?

Measure

  • The same suite on the laptop vs the VPS: wall time, cold vs warm environment, with and without pytest-xdist.
  • The round-trip overhead of each mechanism.

Why ready-for-human, not ready-for-agent

The questions are settled, but an unattended (AFK) agent can't do this work:

  • It writes to a shared production host. The VPS runs Forgejo and the CI runner. Creating scratch areas, installing environments, starting workers and load-testing five parallel jobs are all remote writes. Claude Code's auto mode has denied VPS writes before, and they should happen with the maintainer present.
  • It needs the WireGuard tunnel up and a real benchmark of the box under load. Both need a supervised session.
  • The results feed policy decisions (pre-commit, the "all tests pass" invariant, CI reuse) that are the maintainer's to make.

An agent can do the work in a supervised session: design, local measurements, a prototype of the entry point, and remote commands run with the maintainer's approval.

Blocked by

  • #218: the hermetic suite (#216's network guard, #210's clean worktree) is a safety precondition for running anywhere but the laptop. Its post-fix timings are the local baseline to beat.

Outcome

  • A findings document at docs/research/<this issue number>-remote-test-execution.md (convention set in the #217 / #208 triage briefs), with measured numbers per mechanism and a recommendation.
  • Follow-up issues for the chosen mechanism and any policy change, each needing the maintainer's sign-off.

Related: #217 (parent research), #218, #216.

> *This was generated by AI during triage.* Blocked by: #218 Split from #217 during triage (2026-10-03). It carries the maintainer's note of 2026-10-03 on #217, "remote test execution as a primary option", so #217 and its PR #218 can ship the local findings and quick wins without waiting on VPS work. ## Question Can agents run the test suite on the VPS (vps-proxy, over WireGuard) **by default**, and does that beat local runs on the laptop once #217's local optimisations have landed? The VPS already hosts Forgejo and the Forgejo Actions runner (`python` label, `hikcentral-ci` image), so a remote Python environment for this suite exists. What's open is how an agent triggers it *during* work, without a push and a full CI run. ## Questions to answer 1. **Mechanism.** Compare these on latency from command to first failure, against a local run: - a dedicated "test now" workflow dispatched with `tea` on a pushed WIP ref; - rsync or `git push` of the worktree state to a scratch area on the VPS, then pytest over ssh; - a long-lived remote worker that agents queue test jobs to. 2. **Uncommitted work.** Agents test before committing. How does the working tree reach the server: a diff, an rsync, or a temporary commit pushed to a scratch ref? What gets missed (untracked files, ignored files)? 3. **Concurrency.** About five agents testing at once. Do separate checkouts and DBs per job isolate them? How many parallel jobs fit in the VPS's CPU and RAM? Does it still help when CI runs on the same box? 4. **Interface.** One entry point (e.g. `scripts/wt test --remote [--changed|--full]`) with compact output (failures only), the same result format as a local run, and a clear local fallback when the VPS or VPN is unreachable. 5. **Safety.** - Remote runs stay hermetic: no production data and no secrets beyond what CI already has. - No reach into HikCentral. This depends on #216's session-level network guard. - Scratch checkouts are cleaned up afterwards. 6. **Policy implications.** Recommendations only; each needs the maintainer's sign-off. - Could the pre-commit hook delegate to the remote run? - Does a remote pass on the exact tree hash count toward the AGENTS.md rule that all tests must pass? - Can CI reuse that result instead of re-running? ## Measure - The same suite on the laptop vs the VPS: wall time, cold vs warm environment, with and without pytest-xdist. - The round-trip overhead of each mechanism. ## Why `ready-for-human`, not `ready-for-agent` The questions are settled, but an unattended (AFK) agent can't do this work: - **It writes to a shared production host.** The VPS runs Forgejo and the CI runner. Creating scratch areas, installing environments, starting workers and load-testing five parallel jobs are all remote writes. Claude Code's auto mode has denied VPS writes before, and they should happen with the maintainer present. - **It needs the WireGuard tunnel up** and a real benchmark of the box under load. Both need a supervised session. - **The results feed policy decisions** (pre-commit, the "all tests pass" invariant, CI reuse) that are the maintainer's to make. An agent can do the work *in a supervised session*: design, local measurements, a prototype of the entry point, and remote commands run with the maintainer's approval. ## Blocked by - #218: the hermetic suite (#216's network guard, #210's clean worktree) is a safety precondition for running anywhere but the laptop. Its post-fix timings are the local baseline to beat. ## Outcome - A findings document at `docs/research/<this issue number>-remote-test-execution.md` (convention set in the #217 / #208 triage briefs), with measured numbers per mechanism and a recommendation. - Follow-up issues for the chosen mechanism and any policy change, each needing the maintainer's sign-off. Related: #217 (parent research), #218, #216.
Sign in to join this conversation.
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
gabogg/hikcentral#235
No description provided.