[data-veracity] Passenger flow sync silently stops ingesting after any upstream counter reset #26

Closed
opened 2026-09-21 13:43:26 +00:00 by gabogg · 2 comments
Owner

Filed from a data-veracity audit of the ingestion and aggregation pipeline on master, carried out against the KPI set that the Executive Statistics Deck (PR #20 / RFC-ARCH-2026-004) intends to publish. Each issue names the deck KPIs it corrupts.

Problem

OccupancyManager.sync_passenger_flow_from_artemis_async (app/services/occupancy_service.py:1764) reconciles against HikCentral's cumulative group counter:

# app/services/occupancy_service.py:1898-1899
delta_in  = max(0, current_artemis_in  - total_local_in)
delta_out = max(0, current_artemis_out - total_local_out)

current_artemis_* comes from /artemis/api/aiapplication/v1/people/resourceGroupRealTimeCount, which returns a running total. total_local_* is our own accumulated total for the current business cycle.

Whenever the upstream counter goes down — HikCentral's own daily rollover, an NVR or camera reboot, a resource-group reconfiguration, a manual counter clear — the subtraction goes negative and max(0, ...) clamps it to zero.

From that moment the pipeline records nothing until the upstream counter climbs back above our local total. On a busy day that can be several hours of completely lost traffic, and it fails silently: events_created is 0, no warning is logged, no anomaly flag is raised. The dashboard simply shows a flat line, which is indistinguishable from a genuinely quiet period.

The inverse failure

The clamp also hides the opposite case. Our cycle starts at daily_reset_time (04:00); HikCentral's counter resets on its own schedule. At our 04:00 boundary total_local_* drops to ~0 while the upstream counter is still mid-period at N, so the next poll computes delta = N - 0 = N and injects one synthetic event of N persons stamped at 04:00:00.

That phantom spike lands inside the nocturnal quiet window (calibration_window_start–calibration_window_end, default 03:30–04:30) — exactly the window used to derive the exit multiplier. A counter-rollover artifact is therefore fed straight into k.

The state to fix it already exists and is never read

# app/services/occupancy_service.py:41
self._last_group_readings: dict[str, dict[str, int]] = {}
# app/services/occupancy_service.py:1925  (written every poll)
self._last_group_readings[g_code] = {"in": current_artemis_in, "out": current_artemis_out}

_last_group_readings is initialised and written on every poll and read nowhere in the codebase. A poll-over-poll delta against this field is what would make reset detection possible; right now it is dead state.

KPIs corrupted

Gross Ingress · Gross Egress · Raw Net · Calibrated Residual · Peak Occupancy · Peak Velocity · Mean Dwell · Portal Attribution · every Week and Month roll-up built on them.

Suggested fix

  1. Compute the delta poll-over-poll from _last_group_readings, not cycle-total-over-cycle-total.
  2. Treat current < last as a counter reset: log it, re-seed _last_group_readings from the new value, emit no events for that tick, and write an anomaly marker for the cycle.
  3. Keep the cycle-total reconciliation only as a slow drift correction (e.g. once per minute), and cap any single correction so a rollover can never be published as one giant event.
  4. Fail loudly: if consecutive polls produce delta == 0 while the upstream counter is non-zero and moving, raise a FLAG_INGESTION_STALLED on the cycle.

Needs confirmation

The exact reset epoch and semantics of resourceGroupRealTimeCount on this HikCentral version — does it reset at local midnight, at a configurable boundary, or never? The fix above is correct either way, but the reconciliation window depends on the answer.

> Filed from a data-veracity audit of the ingestion and aggregation pipeline on `master`, carried out against the KPI set that the Executive Statistics Deck (PR #20 / `RFC-ARCH-2026-004`) intends to publish. Each issue names the deck KPIs it corrupts. ## Problem `OccupancyManager.sync_passenger_flow_from_artemis_async` (`app/services/occupancy_service.py:1764`) reconciles against HikCentral's **cumulative** group counter: ```python # app/services/occupancy_service.py:1898-1899 delta_in = max(0, current_artemis_in - total_local_in) delta_out = max(0, current_artemis_out - total_local_out) ``` `current_artemis_*` comes from `/artemis/api/aiapplication/v1/people/resourceGroupRealTimeCount`, which returns a running total. `total_local_*` is our own accumulated total for the current business cycle. Whenever the upstream counter goes **down** — HikCentral's own daily rollover, an NVR or camera reboot, a resource-group reconfiguration, a manual counter clear — the subtraction goes negative and `max(0, ...)` clamps it to zero. From that moment the pipeline records **nothing** until the upstream counter climbs back above our local total. On a busy day that can be several hours of completely lost traffic, and it fails **silently**: `events_created` is `0`, no warning is logged, no anomaly flag is raised. The dashboard simply shows a flat line, which is indistinguishable from a genuinely quiet period. ## The inverse failure The clamp also hides the opposite case. Our cycle starts at `daily_reset_time` (04:00); HikCentral's counter resets on its own schedule. At our 04:00 boundary `total_local_*` drops to ~0 while the upstream counter is still mid-period at `N`, so the next poll computes `delta = N - 0 = N` and injects **one synthetic event of N persons stamped at 04:00:00**. That phantom spike lands inside the nocturnal quiet window (`calibration_window_start`–`calibration_window_end`, default 03:30–04:30) — exactly the window used to derive the exit multiplier. A counter-rollover artifact is therefore fed straight into `k`. ## The state to fix it already exists and is never read ```python # app/services/occupancy_service.py:41 self._last_group_readings: dict[str, dict[str, int]] = {} # app/services/occupancy_service.py:1925 (written every poll) self._last_group_readings[g_code] = {"in": current_artemis_in, "out": current_artemis_out} ``` `_last_group_readings` is initialised and written on every poll and **read nowhere in the codebase**. A poll-over-poll delta against this field is what would make reset detection possible; right now it is dead state. ## KPIs corrupted Gross Ingress · Gross Egress · Raw Net · Calibrated Residual · Peak Occupancy · Peak Velocity · Mean Dwell · Portal Attribution · every Week and Month roll-up built on them. ## Suggested fix 1. Compute the delta **poll-over-poll** from `_last_group_readings`, not cycle-total-over-cycle-total. 2. Treat `current < last` as a **counter reset**: log it, re-seed `_last_group_readings` from the new value, emit **no** events for that tick, and write an anomaly marker for the cycle. 3. Keep the cycle-total reconciliation only as a slow drift correction (e.g. once per minute), and cap any single correction so a rollover can never be published as one giant event. 4. Fail loudly: if consecutive polls produce `delta == 0` while the upstream counter is non-zero and moving, raise a `FLAG_INGESTION_STALLED` on the cycle. ## Needs confirmation The exact reset epoch and semantics of `resourceGroupRealTimeCount` on this HikCentral version — does it reset at local midnight, at a configurable boundary, or never? The fix above is correct either way, but the reconciliation window depends on the answer.
Author
Owner

✅ Design settled (grilling session)

The needs-info (reset epoch of resourceGroupRealTimeCount) is not a blocker — the fix is correct regardless of the exact reset boundary; the epoch only tunes the drift-reconciliation cadence and can be confirmed against prod later.

Settled spec

  • Poll-over-poll delta from _last_group_readings (currently written every poll and read nowhere — occupancy_service.py:41, 1925), not cycle-total-over-cycle-total.
  • Reset detection: current < last ⇒ treat as a counter reset — log it, re-seed _last_group_readings from the new value, emit no events for that tick, write an anomaly marker on the cycle.
  • Drift correction: keep cycle-total reconciliation only as a slow correction (≈ once/minute), capped adaptively at ≈ 3× the group's trailing per-minute rate (e.g. last 15 min). Anything above the cap is treated as a reset/anomaly and re-seeded, never published — so a rollover can never surface as one giant event.
  • FLAG_INGESTION_STALLED: fires on ≥ 5 consecutive failed/empty polls (the silent early-returns at :1786 / :1845 / :1951), and is tied into #31's ingestion-gap ledger — a stall is the start of a gap, not a sensor fault. Plus a safety assertion: if the raw counter advances while emitted events for the group are 0, flag it (should never trip post-fix; if it does, something re-broke).

Acceptance criteria (supersede "Needs confirmation")

  • Deltas computed poll-over-poll from _last_group_readings.
  • current < last re-seeds and emits nothing; cycle marked.
  • Drift correction capped adaptively at ≈3× trailing per-minute rate.
  • FLAG_INGESTION_STALLED on ≥5 consecutive failed/empty polls, linked to the #31 gap ledger.
  • A simulated upstream reset (counter drops mid-cycle) loses no subsequent traffic and injects no 04:00 phantom spike into the calibration window.
  • Tests cover both the reset case and the stall case.

Re-tagged ready-for-agent. (Reset-epoch confirmation against prod remains an optional tuning refinement, not a blocker.)

## ✅ Design settled (grilling session) The `needs-info` (reset epoch of `resourceGroupRealTimeCount`) is **not a blocker** — the fix is correct regardless of the exact reset boundary; the epoch only tunes the drift-reconciliation cadence and can be confirmed against prod later. ### Settled spec - **Poll-over-poll delta** from `_last_group_readings` (currently written every poll and read nowhere — `occupancy_service.py:41, 1925`), not cycle-total-over-cycle-total. - **Reset detection**: `current < last` ⇒ treat as a counter reset — log it, re-seed `_last_group_readings` from the new value, emit **no** events for that tick, write an anomaly marker on the cycle. - **Drift correction**: keep cycle-total reconciliation only as a slow correction (≈ once/minute), **capped adaptively** at ≈ **3× the group's trailing per-minute rate** (e.g. last 15 min). Anything above the cap is treated as a reset/anomaly and re-seeded, never published — so a rollover can never surface as one giant event. - **`FLAG_INGESTION_STALLED`**: fires on **≥ 5 consecutive failed/empty polls** (the silent early-returns at `:1786 / :1845 / :1951`), and is **tied into #31's ingestion-gap ledger** — a stall is the start of a gap, not a sensor fault. Plus a safety assertion: if the raw counter advances while emitted events for the group are 0, flag it (should never trip post-fix; if it does, something re-broke). ### Acceptance criteria (supersede "Needs confirmation") - [ ] Deltas computed poll-over-poll from `_last_group_readings`. - [ ] `current < last` re-seeds and emits nothing; cycle marked. - [ ] Drift correction capped adaptively at ≈3× trailing per-minute rate. - [ ] `FLAG_INGESTION_STALLED` on ≥5 consecutive failed/empty polls, linked to the #31 gap ledger. - [ ] A simulated upstream reset (counter drops mid-cycle) loses no subsequent traffic and injects no 04:00 phantom spike into the calibration window. - [ ] Tests cover both the reset case and the stall case. Re-tagged `ready-for-agent`. (Reset-epoch confirmation against prod remains an optional tuning refinement, not a blocker.)
Author
Owner

Being addressed in draft PR #68 together with #26, #31 and #30 (passenger flow ingestion honesty). The PR description lists the settled spec plus the open design points that still need a decision (restart cold start, a floor for the adaptive cap, a shared anomaly/gap ledger table, and how trust rules see reconstructed hours).

Being addressed in draft **PR #68** together with #26, #31 and #30 (passenger flow ingestion honesty). The PR description lists the settled spec plus the open design points that still need a decision (restart cold start, a floor for the adaptive cap, a shared anomaly/gap ledger table, and how trust rules see reconstructed hours).
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
gabogg/hikcentral#26
No description provided.