research(occupancy): decide the policy for a single silent counting camera #236

Open
opened 2026-10-03 07:44:07 +00:00 by gabogg · 1 comment
Owner

This was generated by AI during triage.

Blocked by: #199

Split from #199 during triage (2026-10-03). #199 now covers phase A only: documenting how the app behaves today when one counting camera goes silent. This issue is phase B: deciding what the app should do. It takes over #199's questions 4 and 5, and any part of question 1 that phase A can't answer from stored data.

Why this is a separate issue, labelled needs-triage

These are policy and product decisions. They belong to the maintainer and shouldn't be guessed by an implementing agent:

  • how severe a partial camera outage is for a published day;
  • whether published days get revised when late data arrives;
  • what supervisors see.

They also shouldn't be decided before phase A's findings exist. Without a traced failure case and evidence of what HikCentral reports, any policy would rest on assumptions. So this issue stays needs-triage until #199 closes, then gets a grilling session (/triage → grilling + domain-modeling). The decisions land in CONTEXT.md and an ADR before any implementation issue is cut.

Questions to settle (once #199 has reported)

  1. Detection. Using phase A's evidence, what tells a silent camera apart from a legitimately quiet entrance? Examples: an expected schedule, a freshness window, the share of traffic seen on comparable days, explicit fields from HikCentral. Do not treat a zero count as failure on its own.
  2. Severity. Should a day with a partial camera outage be Estimate, Unreliable (excluded), or still measured with an explicit coverage warning? Does that depend on camera role (ENTRANCE / EXIT / BIDIRECTIONAL), the entrance's share of traffic, outage duration, or whether another camera covers the same path? Don't invent a weighting without defensible evidence; phase A reports whether any exists.
  3. Recovery and revision. If the camera's data arrives later, is a published day recomputed? How is that revision recorded, and what happens to statistics comparisons and calibration learning that already used the day?
  4. Operator/supervisor diagnostic. What should identify the camera, the interval, the affected metrics and the confidence before the day is used in statistics or calibration? Where does it show up?
  5. Remaining HikCentral behaviour. Any part of #199's question 1 that phase A couldn't answer offline needs a check against the live system, ideally a controlled outage of one camera. Only the maintainer can arrange that.

Outcome

  • Decisions recorded in CONTEXT.md (terms) and an ADR (policy).
  • Implementation issues cut from those decisions, in this milestone.

Related: #199 (phase A), #129 (statistics data-quality marker), ADR 0005 (one camera per Camera Group), ADR 0006 (published occupancy and calibration honesty).

> *This was generated by AI during triage.* Blocked by: #199 Split from #199 during triage (2026-10-03). #199 now covers **phase A only**: documenting how the app behaves today when one counting camera goes silent. This issue is **phase B**: deciding what the app *should* do. It takes over #199's questions 4 and 5, and any part of question 1 that phase A can't answer from stored data. ## Why this is a separate issue, labelled `needs-triage` These are policy and product decisions. They belong to the maintainer and shouldn't be guessed by an implementing agent: - how severe a partial camera outage is for a published day; - whether published days get revised when late data arrives; - what supervisors see. They also shouldn't be decided before phase A's findings exist. Without a traced failure case and evidence of what HikCentral reports, any policy would rest on assumptions. So this issue stays `needs-triage` until #199 closes, then gets a grilling session (`/triage` → grilling + domain-modeling). The decisions land in `CONTEXT.md` and an ADR before any implementation issue is cut. ## Questions to settle (once #199 has reported) 1. **Detection.** Using phase A's evidence, what tells a silent camera apart from a legitimately quiet entrance? Examples: an expected schedule, a freshness window, the share of traffic seen on comparable days, explicit fields from HikCentral. Do not treat a zero count as failure on its own. 2. **Severity.** Should a day with a partial camera outage be **Estimate**, **Unreliable** (excluded), or still measured with an explicit coverage warning? Does that depend on camera role (`ENTRANCE` / `EXIT` / `BIDIRECTIONAL`), the entrance's share of traffic, outage duration, or whether another camera covers the same path? Don't invent a weighting without defensible evidence; phase A reports whether any exists. 3. **Recovery and revision.** If the camera's data arrives later, is a published day recomputed? How is that revision recorded, and what happens to statistics comparisons and calibration learning that already used the day? 4. **Operator/supervisor diagnostic.** What should identify the camera, the interval, the affected metrics and the confidence before the day is used in statistics or calibration? Where does it show up? 5. **Remaining HikCentral behaviour.** Any part of #199's question 1 that phase A couldn't answer offline needs a check against the live system, ideally a controlled outage of one camera. Only the maintainer can arrange that. ## Outcome - Decisions recorded in `CONTEXT.md` (terms) and an ADR (policy). - Implementation issues cut from those decisions, in this milestone. Related: #199 (phase A), #129 (statistics data-quality marker), ADR 0005 (one camera per Camera Group), ADR 0006 (published occupancy and calibration honesty).
Author
Owner

Phase A Findings: Open Questions Requiring Live HikCentral Outage Observation (Ref #199 / PR #248)

As documented in docs/research/199-single-camera-outage-behaviour.md (PR #248), offline analysis of repo OpenAPI specifications and client models confirms that Artemis passenger-flow payloads (resourceGroupRealTimeCount) contain only cumulative integers (enterNum, exitNum) with no health flags, sensor codes, or per-camera status fields.

The following questions cannot be resolved from repo documentation or client models alone and require controlled empirical testing on physical hardware by the maintainer:

  1. Behavior of resourceGroupRealTimeCount during camera outage:
    When a physical counting camera loses power or network link, does Artemis:
    • Keep returning the camera's resource group with frozen cumulative numbers?
    • Omit the group from data.list entirely?
    • Return an error code at the response or group level?
  2. Asynchronous event notifications:
    Does HikCentral emit an asynchronous alarm or event over the OpenAPI event subscription (port 7016/7017) when a camera drops offline (comparable to door hardware alarms 131585–131588)?
  3. Behavior of historical queries (statisticsTotalNumByTime):
    During camera outages, how are historical hourly intervals reported in Artemis: omitted rows, zero counts, or a degraded completeness score?
  4. Auxiliary camera health endpoints:
    Does HikCentral expose an authoritative camera online/offline status API (e.g., via video/resource endpoints) that could be correlated with people-counting groups to detect silent cameras?
## Phase A Findings: Open Questions Requiring Live HikCentral Outage Observation (Ref #199 / PR #248) As documented in `docs/research/199-single-camera-outage-behaviour.md` (PR #248), offline analysis of repo OpenAPI specifications and client models confirms that Artemis passenger-flow payloads (`resourceGroupRealTimeCount`) contain only cumulative integers (`enterNum`, `exitNum`) with no health flags, sensor codes, or per-camera status fields. The following questions cannot be resolved from repo documentation or client models alone and require controlled empirical testing on physical hardware by the maintainer: 1. **Behavior of `resourceGroupRealTimeCount` during camera outage**: When a physical counting camera loses power or network link, does Artemis: - Keep returning the camera's resource group with frozen cumulative numbers? - Omit the group from `data.list` entirely? - Return an error code at the response or group level? 2. **Asynchronous event notifications**: Does HikCentral emit an asynchronous alarm or event over the OpenAPI event subscription (port 7016/7017) when a camera drops offline (comparable to door hardware alarms 131585–131588)? 3. **Behavior of historical queries (`statisticsTotalNumByTime`)**: During camera outages, how are historical hourly intervals reported in Artemis: omitted rows, zero counts, or a degraded `completeness` score? 4. **Auxiliary camera health endpoints**: Does HikCentral expose an authoritative camera online/offline status API (e.g., via video/resource endpoints) that could be correlated with people-counting groups to detect silent cameras?
Sign in to join this conversation.
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
gabogg/hikcentral#236
No description provided.