[data-veracity] Passenger flow sync silently stops ingesting after any upstream counter reset #26
Labels
No labels
blocked
bug
enhancement
high-priority
low-priority
needs-info
needs-triage
ready-for-agent
ready-for-human
referenced
research
wontfix
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
gabogg/hikcentral#26
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Problem
OccupancyManager.sync_passenger_flow_from_artemis_async(app/services/occupancy_service.py:1764) reconciles against HikCentral's cumulative group counter:current_artemis_*comes from/artemis/api/aiapplication/v1/people/resourceGroupRealTimeCount, which returns a running total.total_local_*is our own accumulated total for the current business cycle.Whenever the upstream counter goes down — HikCentral's own daily rollover, an NVR or camera reboot, a resource-group reconfiguration, a manual counter clear — the subtraction goes negative and
max(0, ...)clamps it to zero.From that moment the pipeline records nothing until the upstream counter climbs back above our local total. On a busy day that can be several hours of completely lost traffic, and it fails silently:
events_createdis0, no warning is logged, no anomaly flag is raised. The dashboard simply shows a flat line, which is indistinguishable from a genuinely quiet period.The inverse failure
The clamp also hides the opposite case. Our cycle starts at
daily_reset_time(04:00); HikCentral's counter resets on its own schedule. At our 04:00 boundarytotal_local_*drops to ~0 while the upstream counter is still mid-period atN, so the next poll computesdelta = N - 0 = Nand injects one synthetic event of N persons stamped at 04:00:00.That phantom spike lands inside the nocturnal quiet window (
calibration_window_start–calibration_window_end, default 03:30–04:30) — exactly the window used to derive the exit multiplier. A counter-rollover artifact is therefore fed straight intok.The state to fix it already exists and is never read
_last_group_readingsis initialised and written on every poll and read nowhere in the codebase. A poll-over-poll delta against this field is what would make reset detection possible; right now it is dead state.KPIs corrupted
Gross Ingress · Gross Egress · Raw Net · Calibrated Residual · Peak Occupancy · Peak Velocity · Mean Dwell · Portal Attribution · every Week and Month roll-up built on them.
Suggested fix
_last_group_readings, not cycle-total-over-cycle-total.current < lastas a counter reset: log it, re-seed_last_group_readingsfrom the new value, emit no events for that tick, and write an anomaly marker for the cycle.delta == 0while the upstream counter is non-zero and moving, raise aFLAG_INGESTION_STALLEDon the cycle.Needs confirmation
The exact reset epoch and semantics of
resourceGroupRealTimeCounton this HikCentral version — does it reset at local midnight, at a configurable boundary, or never? The fix above is correct either way, but the reconciliation window depends on the answer.✅ Design settled (grilling session)
The
needs-info(reset epoch ofresourceGroupRealTimeCount) is not a blocker — the fix is correct regardless of the exact reset boundary; the epoch only tunes the drift-reconciliation cadence and can be confirmed against prod later.Settled spec
_last_group_readings(currently written every poll and read nowhere —occupancy_service.py:41, 1925), not cycle-total-over-cycle-total.current < last⇒ treat as a counter reset — log it, re-seed_last_group_readingsfrom the new value, emit no events for that tick, write an anomaly marker on the cycle.FLAG_INGESTION_STALLED: fires on ≥ 5 consecutive failed/empty polls (the silent early-returns at:1786 / :1845 / :1951), and is tied into #31's ingestion-gap ledger — a stall is the start of a gap, not a sensor fault. Plus a safety assertion: if the raw counter advances while emitted events for the group are 0, flag it (should never trip post-fix; if it does, something re-broke).Acceptance criteria (supersede "Needs confirmation")
_last_group_readings.current < lastre-seeds and emits nothing; cycle marked.FLAG_INGESTION_STALLEDon ≥5 consecutive failed/empty polls, linked to the #31 gap ledger.Re-tagged
ready-for-agent. (Reset-epoch confirmation against prod remains an optional tuning refinement, not a blocker.)Being addressed in draft PR #68 together with #26, #31 and #30 (passenger flow ingestion honesty). The PR description lists the settled spec plus the open design points that still need a decision (restart cold start, a floor for the adaptive cap, a shared anomaly/gap ledger table, and how trust rules see reconstructed hours).
gabogg referenced this issue2026-09-24 13:13:51 +00:00