investigate single-camera data loss and statistics quality #199

Open
opened 2026-10-02 14:52:36 +00:00 by gabogg · 1 comment
Owner

Question

What happens when one active counting camera stops providing data while the other cameras continue? Today we know a facility-wide ingestion gap can affect a business day's quality and eligibility for statistics, but we have not established whether a single silent camera is detected, how much it affects published totals, or whether any weighting/coverage policy exists. Do not assume a zero count means a failed camera: a quiet entrance can legitimately have no passages.

Current context to inspect

  • Published event totals include registered, active, non-excluded cameras (app/db/occupancy_repository.py); this identifies which cameras count, but does not by itself establish per-camera completeness.
  • CONTEXT.md defines an ingestion gap as failure to obtain a successful passenger-flow reading and distinguishes missing data from low traffic. Statistics quality markers and calibration eligibility use day-level evidence.
  • Historical flow staging records camera-hour coverage for imported data; determine whether equivalent coverage exists for live data and whether it reaches daily verdicts and the statistics deck.

Triage / investigation questions

  1. What does HikCentral return when one camera is offline or missing while a camera group still responds? Can group totals mask the missing camera? Which source fields distinguish zero passages from absent/unreliable readings?
  2. What per-camera freshness, expected schedule, and coverage does the app currently persist? Trace one silent-camera interval through ingestion, daily audit/verdict, statistics quality marker, comparison eligibility, and calibration learning.
  3. Is there any existing weight, expected contribution, or threshold for losing one camera? If not, should impact depend on camera role, direction, entrance share, outage duration, and whether another camera covers the same path? Avoid inventing an estimate without defensible evidence.
  4. Should a partial camera outage mark a day Estimated, Unreliable/excluded, or still measured with an explicit coverage warning? What are the rules for recovering later data and revising a published day?
  5. What operator/supervisor diagnostic should identify the camera, interval, affected metrics, and confidence before the day is used in statistics or calibration?

Ready-to-spec outcome

Document the current behavior with a reproducible one-camera failure case, then decide the policy for detection, severity, publication, recovery, and tests. Cover both a truly silent camera and a legitimately quiet camera so the policy does not mistake zero traffic for failure.

Related: statistics data-quality marker work #129 and historical flow coverage/review. This issue is exploratory and does not presume a weighted system already exists.

## Question What happens when **one active counting camera stops providing data** while the other cameras continue? Today we know a facility-wide ingestion gap can affect a business day's quality and eligibility for statistics, but we have not established whether a single silent camera is detected, how much it affects published totals, or whether any weighting/coverage policy exists. Do not assume a zero count means a failed camera: a quiet entrance can legitimately have no passages. ## Current context to inspect - Published event totals include registered, active, non-excluded cameras (`app/db/occupancy_repository.py`); this identifies which cameras count, but does not by itself establish per-camera completeness. - `CONTEXT.md` defines an ingestion gap as failure to obtain a successful passenger-flow reading and distinguishes missing data from low traffic. Statistics quality markers and calibration eligibility use day-level evidence. - Historical flow staging records camera-hour coverage for imported data; determine whether equivalent coverage exists for live data and whether it reaches daily verdicts and the statistics deck. ## Triage / investigation questions 1. What does HikCentral return when one camera is offline or missing while a camera group still responds? Can group totals mask the missing camera? Which source fields distinguish zero passages from absent/unreliable readings? 2. What per-camera freshness, expected schedule, and coverage does the app currently persist? Trace one silent-camera interval through ingestion, daily audit/verdict, statistics quality marker, comparison eligibility, and calibration learning. 3. Is there any existing weight, expected contribution, or threshold for losing one camera? If not, should impact depend on camera role, direction, entrance share, outage duration, and whether another camera covers the same path? Avoid inventing an estimate without defensible evidence. 4. Should a partial camera outage mark a day Estimated, Unreliable/excluded, or still measured with an explicit coverage warning? What are the rules for recovering later data and revising a published day? 5. What operator/supervisor diagnostic should identify the camera, interval, affected metrics, and confidence before the day is used in statistics or calibration? ## Ready-to-spec outcome Document the current behavior with a reproducible one-camera failure case, then decide the policy for detection, severity, publication, recovery, and tests. Cover both a truly silent camera and a legitimately quiet camera so the policy does not mistake zero traffic for failure. Related: statistics data-quality marker work #129 and historical flow coverage/review. This issue is exploratory and does not presume a weighted system already exists.
Author
Owner

This was generated by AI during triage.

Triage decision: needs-triage → split, then ready-for-agent (phase A)

Why split. As written, #199 mixes three kinds of work that need different people:

Part Questions Who can do it Where it lives now
Trace current behaviour in code and stored data Q2, Q3, and the parts of Q1 answerable offline an AFK agent #199 (this issue, phase A)
Observe live HikCentral with a camera actually offline the rest of Q1 maintainer only. HikCentral is real hardware; tests and local runs must never reach it. #236
Decide policy: severity, publication, revision, diagnostic Q4, Q5 maintainer decision, after the evidence #236, needs-triage, blocked by #199

Leaving it all in one issue would either stall an agent on decisions it may not make, or tempt it to make them. Policy settled by an implementing agent without maintainer sign-off has caused trouble before.

What triage established (redundancy check):

  • No per-camera freshness or coverage exists for live ingestion. Camera-hour coverage exists only for Historical Flow Staging (imported data held for admin review). I searched the application code, CONTEXT.md and docs/ for camera-hour, per-camera, coverage and silent/stale-camera concepts.
  • Day-level quality is facility-wide. The Data-Quality Marker (Estimate / Unreliable) comes from the Cycle Verdict, ingestion-gap estimates and missing data. An Ingestion Gap is a facility-wide failure to get a successful reading, not a per-camera one.
  • Group totals masking a missing camera is mostly moot here. Per CONTEXT.md and ADR 0005, HikCentral reports flow per Camera Group, and each group holds exactly one camera. The exception is the Multi-camera group state, which the trace should cover.
  • No .out-of-scope/ entry relates to this.

Milestone: "Single silent counting camera" groups #199 and #236, plus the implementation issues #236 will produce.


Agent Brief

Category: enhancement, with research (needs investigation before it can be specified further)
Summary: Document, with a reproducible case, exactly what the app does today when one active counting camera goes silent while the others keep reporting, and contrast it with a legitimately quiet camera. Change no behaviour.

Current behavior (as known before the investigation):

  • Published event totals include every registered, active, non-excluded counting camera.
  • Quality and eligibility verdicts are made per business day, facility-wide: Ingestion Gap, Cycle Completeness, Data Trust, Cycle Verdict, Data-Quality Marker, comparison coverage, Weekday Baseline eligibility, and calibration learning.
  • Nobody has traced what a single silent camera does to each of those.

Desired behavior (the deliverables):

  1. Findings document at docs/research/199-single-camera-outage-behaviour.md. It answers, with evidence:

    • Q1 (offline part). From the Artemis/ISAPI passenger-flow API documentation in the repo and the response shapes the client parses: what does a group's reading contain, and which fields, if any, tell "zero passages" apart from "no or unreliable reading" (status, timestamps, missing records vs zero records)? List what can't be answered without a live outage as open questions for #236. Don't guess.
    • Q2. What per-camera state does the app persist for live data (last reading time, per-camera hourly counts, expected schedule)? Then trace one silent-camera interval stage by stage: passenger-flow polling and ingestion → ingestion-gap detection → Cycle Completeness and Data Trust → Cycle Verdict → Data-Quality Marker → statistics comparison eligibility and Weekday Baseline → calibration (exit multiplier) learning. At each stage, say whether the silent camera is noticed, ignored, or silently lowers totals.
    • Q3. Does any weight, expected contribution, or threshold for losing one camera exist anywhere: config, settings, code constants? State plainly if none exists.
    • Contrast. Do the same trace for a legitimately quiet camera, and show whether today's system treats the two cases identically.
    • Multi-camera group. One paragraph on how the trace differs when the silent camera is inside a multi-camera group.
  2. Characterization tests in the test suite, deterministic and offline, with the HikCentral clients mocked as AGENTS.md requires. They feed a facility where one camera returns nothing (or zero) for a business-hours interval while the others report traffic, and a second scenario where one camera is legitimately quiet. They assert today's observable outputs at the stages above: published totals, the day's quality state and marker, comparison eligibility, and calibration inclusion. Each test's docstring says it pins current behaviour pending the #236 policy, so a later behaviour change updates them deliberately.

  3. Optional production evidence, read-only. If the prod DB is reachable over WireGuard, look for real past intervals where one camera reported nothing during open hours while others had traffic, and report how those days were classified. Rules:

    • query in place, read-only (mode=ro SQLite URI), piping a script over SSH to the app's venv Python;
    • never copy the DB, never write, and never put cardholder or personal data in the doc (camera names and counts are fine).

    If it isn't reachable, say so and skip this step.

Key interfaces / places to look (by concept):

  • The occupancy repository's query for counted cameras (registered, active, non-excluded): it defines which cameras contribute.
  • The passenger-flow polling path in the monitor/occupancy services, and how a failed or empty group reading is recorded.
  • The day quality functions: the period quality summary that feeds the statistics deck, the day coverage and quality-state helpers, gap spreading, and the minimum comparison coverage setting.
  • The calibration learning inputs: which cycles count as trusted.
  • Historical Flow Staging's camera-hour coverage, as the existing model of per-camera coverage, for comparison.

Acceptance criteria:

  • docs/research/199-single-camera-outage-behaviour.md answers Q1 (the offline part), Q2 and Q3, plus the quiet-camera contrast and the multi-camera note, each with the code concept or test that shows it.
  • Every stage in the Q2 trace is marked detects, ignores or silently lowers totals, with evidence.
  • The characterization tests for the silent and quiet scenarios pass offline and say they pin pre-#236 behaviour.
  • The open questions that need a live outage are listed in the doc and copied as a comment on #236.
  • If prod was queried: the method is stated, the query is read-only, and the doc contains no personal data.
  • If docs/research/ doesn't exist yet when this lands: add docs/research/README.md (a one-line index per document) and a link to it from docs/README.md. If another research PR got there first, rebase and add a line to its index.
  • Ruff, scripts/check_docs.py and the full pytest suite pass. The PR joins the "Single silent counting camera" milestone.

Out of scope:

  • Any behaviour change: detection, warnings, weighting, marking days Estimate or Unreliable, revision of published days. That is #236's policy, then its implementation issues.
  • Contacting or testing against live HikCentral, or arranging a real camera outage. That's the maintainer's job, under #236.
  • Inventing a weighting or threshold. Report what exists; recommend nothing numeric without evidence.
  • Changing CONTEXT.md or ADRs. New terms are proposed in the doc for #236's grilling.
> *This was generated by AI during triage.* ## Triage decision: `needs-triage` → split, then `ready-for-agent` (phase A) **Why split.** As written, #199 mixes three kinds of work that need different people: | Part | Questions | Who can do it | Where it lives now | |---|---|---|---| | **Trace current behaviour** in code and stored data | Q2, Q3, and the parts of Q1 answerable offline | an AFK agent | **#199 (this issue, phase A)** | | **Observe live HikCentral** with a camera actually offline | the rest of Q1 | maintainer only. HikCentral is real hardware; tests and local runs must never reach it. | #236 | | **Decide policy**: severity, publication, revision, diagnostic | Q4, Q5 | maintainer decision, after the evidence | **#236**, `needs-triage`, blocked by #199 | Leaving it all in one issue would either stall an agent on decisions it may not make, or tempt it to make them. Policy settled by an implementing agent without maintainer sign-off has caused trouble before. **What triage established (redundancy check):** - **No per-camera freshness or coverage exists for live ingestion.** Camera-hour coverage exists only for **Historical Flow Staging** (imported data held for admin review). I searched the application code, `CONTEXT.md` and `docs/` for camera-hour, per-camera, coverage and silent/stale-camera concepts. - **Day-level quality is facility-wide.** The **Data-Quality Marker** (Estimate / Unreliable) comes from the **Cycle Verdict**, ingestion-gap estimates and missing data. An **Ingestion Gap** is a facility-wide failure to get a successful reading, not a per-camera one. - **Group totals masking a missing camera is mostly moot here.** Per `CONTEXT.md` and ADR 0005, HikCentral reports flow per **Camera Group**, and each group holds exactly one camera. The exception is the **Multi-camera group** state, which the trace should cover. - No `.out-of-scope/` entry relates to this. **Milestone:** "Single silent counting camera" groups #199 and #236, plus the implementation issues #236 will produce. --- ## Agent Brief **Category:** enhancement, with `research` (needs investigation before it can be specified further) **Summary:** Document, with a reproducible case, exactly what the app does today when one active counting camera goes silent while the others keep reporting, and contrast it with a legitimately quiet camera. Change no behaviour. **Current behavior (as known before the investigation):** - Published event totals include every registered, active, non-excluded counting camera. - Quality and eligibility verdicts are made per business day, facility-wide: Ingestion Gap, Cycle Completeness, Data Trust, Cycle Verdict, Data-Quality Marker, comparison coverage, Weekday Baseline eligibility, and calibration learning. - Nobody has traced what a single silent camera does to each of those. **Desired behavior (the deliverables):** 1. **Findings document** at **`docs/research/199-single-camera-outage-behaviour.md`**. It answers, with evidence: - **Q1 (offline part).** From the Artemis/ISAPI passenger-flow API documentation in the repo and the response shapes the client parses: what does a group's reading contain, and which fields, if any, tell "zero passages" apart from "no or unreliable reading" (status, timestamps, missing records vs zero records)? List what can't be answered without a live outage as **open questions for #236**. Don't guess. - **Q2.** What per-camera state does the app persist for live data (last reading time, per-camera hourly counts, expected schedule)? Then trace **one silent-camera interval** stage by stage: passenger-flow polling and ingestion → ingestion-gap detection → Cycle Completeness and Data Trust → Cycle Verdict → Data-Quality Marker → statistics comparison eligibility and Weekday Baseline → calibration (exit multiplier) learning. At each stage, say whether the silent camera is noticed, ignored, or silently lowers totals. - **Q3.** Does any weight, expected contribution, or threshold for losing one camera exist anywhere: config, settings, code constants? State plainly if none exists. - **Contrast.** Do the same trace for a **legitimately quiet** camera, and show whether today's system treats the two cases identically. - **Multi-camera group.** One paragraph on how the trace differs when the silent camera is inside a multi-camera group. 2. **Characterization tests** in the test suite, deterministic and offline, with the HikCentral clients mocked as AGENTS.md requires. They feed a facility where one camera returns nothing (or zero) for a business-hours interval while the others report traffic, and a second scenario where one camera is legitimately quiet. They assert **today's** observable outputs at the stages above: published totals, the day's quality state and marker, comparison eligibility, and calibration inclusion. Each test's docstring says it pins current behaviour pending the #236 policy, so a later behaviour change updates them deliberately. 3. **Optional production evidence, read-only.** If the prod DB is reachable over WireGuard, look for real past intervals where one camera reported nothing during open hours while others had traffic, and report how those days were classified. Rules: - query **in place, read-only** (`mode=ro` SQLite URI), piping a script over SSH to the app's venv Python; - never copy the DB, never write, and never put cardholder or personal data in the doc (camera names and counts are fine). If it isn't reachable, say so and skip this step. **Key interfaces / places to look (by concept):** - The occupancy repository's query for counted cameras (registered, active, non-excluded): it defines which cameras contribute. - The passenger-flow polling path in the monitor/occupancy services, and how a failed or empty group reading is recorded. - The day quality functions: the period quality summary that feeds the statistics deck, the day coverage and quality-state helpers, gap spreading, and the minimum comparison coverage setting. - The calibration learning inputs: which cycles count as trusted. - Historical Flow Staging's camera-hour coverage, as the existing model of per-camera coverage, for comparison. **Acceptance criteria:** - [ ] `docs/research/199-single-camera-outage-behaviour.md` answers Q1 (the offline part), Q2 and Q3, plus the quiet-camera contrast and the multi-camera note, each with the code concept or test that shows it. - [ ] Every stage in the Q2 trace is marked *detects*, *ignores* or *silently lowers totals*, with evidence. - [ ] The characterization tests for the silent and quiet scenarios pass offline and say they pin pre-#236 behaviour. - [ ] The open questions that need a live outage are listed in the doc and copied as a comment on #236. - [ ] If prod was queried: the method is stated, the query is read-only, and the doc contains no personal data. - [ ] If `docs/research/` doesn't exist yet when this lands: add `docs/research/README.md` (a one-line index per document) and a link to it from `docs/README.md`. If another research PR got there first, rebase and add a line to its index. - [ ] Ruff, `scripts/check_docs.py` and the full pytest suite pass. The PR joins the "Single silent counting camera" milestone. **Out of scope:** - Any behaviour change: detection, warnings, weighting, marking days Estimate or Unreliable, revision of published days. That is #236's policy, then its implementation issues. - Contacting or testing against live HikCentral, or arranging a real camera outage. That's the maintainer's job, under #236. - Inventing a weighting or threshold. Report what exists; recommend nothing numeric without evidence. - Changing `CONTEXT.md` or ADRs. New terms are proposed in the doc for #236's grilling.
Sign in to join this conversation.
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
gabogg/hikcentral#199
No description provided.