[data-veracity] Event timestamps are poll time, not traversal time — outages collapse into a single bucket #31

Closed
opened 2026-09-21 13:43:31 +00:00 by gabogg · 2 comments
Owner

Filed from a data-veracity audit of the ingestion and aggregation pipeline on master, carried out against the KPI set that the Executive Statistics Deck (PR #20 / RFC-ARCH-2026-004) intends to publish. Each issue names the deck KPIs it corrupts.

Problem

Every ingested event is stamped with the moment the poll ran, not the moment anyone walked through a door:

# app/services/occupancy_service.py:1909, 1921
await self.record_counting_event_async(c_code, "IN", count=c_delta, timestamp=now, camera_name=c_name)

At the normal 3-second cadence (monitor_service.py, last_passenger_sync >= 3.0) this is harmless — 3 s of skew is invisible in an hourly bucket.

It stops being harmless the moment a poll does not run. sync_passenger_flow_from_artemis_async returns early on a failed resource-group fetch (:1786), a failed real-time-count fetch (:1845), or any exception (:1951), and the loop's outer except (monitor_service.py) swallows errors and sleeps. Deltas keep accumulating upstream and are then attributed in full to the single timestamp of the next successful poll.

A 20-minute Artemis outage therefore produces a 20-minute hole followed by one timestamp holding 20 minutes of traffic.

KPIs corrupted

  • 24H Diurnal Flux curve — a flat gap followed by a vertical spike that never happened.
  • Peak Flow Velocity (ν_max = max_h(I_h + E_h)) — the recovery hour wins, permanently.
  • Hourly Kinetics matrix and the Week 7×24 intensity heatmap — the spike lands in whichever cell the recovery fell in.
  • Peak Occupancy timestamp — pinned to recovery (see the peak-ordering issue).
  • Trust rules R1/R2 in evaluate_cycle_integrity_async — FLAG_TRUNCATED_HOURS (fewer than 10 active hours) and FLAG_BURST_COUNTER_FLUSH (any hour above 35% of volume) will both fire, so the cycle is auto-excluded from k learning and the month's Trust Index drops. The trust engine correctly detects the symptom but attributes it to the sensors rather than to our own polling gap.

Suggested fix

Three options, in order of preference:

  1. Ingest with real timestamps. If HikCentral exposes a historical/interval passenger flow query (.../people/advance/... over a time range), poll that on recovery and backfill events at their true buckets. This is the only option that actually restores the data.
  2. Spread the backlog. If only the cumulative counter is available, distribute a recovered delta uniformly across the gap [last_successful_poll, now] rather than piling it on one instant, and mark those events with a raw_payload flag so they are identifiable as reconstructed.
  3. At minimum, record the gap. Persist a per-cycle ingestion-gap ledger (start, end, delta recovered) and surface it. An hour of reconstructed data must not be presented with the same confidence as an hour of measured data — the deck's 95% CI band should widen across a gap.

Whichever is chosen, evaluate_cycle_integrity_async should distinguish FLAG_INGESTION_GAP (our fault) from FLAG_BURST_COUNTER_FLUSH (sensor fault). Right now they are indistinguishable in the calibration ledger.

Decision needed

Which of the three. Option 1 depends on an upstream endpoint we have not yet confirmed exists.

> Filed from a data-veracity audit of the ingestion and aggregation pipeline on `master`, carried out against the KPI set that the Executive Statistics Deck (PR #20 / `RFC-ARCH-2026-004`) intends to publish. Each issue names the deck KPIs it corrupts. ## Problem Every ingested event is stamped with the moment the poll ran, not the moment anyone walked through a door: ```python # app/services/occupancy_service.py:1909, 1921 await self.record_counting_event_async(c_code, "IN", count=c_delta, timestamp=now, camera_name=c_name) ``` At the normal 3-second cadence (`monitor_service.py`, `last_passenger_sync >= 3.0`) this is harmless — 3 s of skew is invisible in an hourly bucket. It stops being harmless the moment a poll does not run. `sync_passenger_flow_from_artemis_async` returns early on a failed resource-group fetch (`:1786`), a failed real-time-count fetch (`:1845`), or any exception (`:1951`), and the loop's outer `except` (`monitor_service.py`) swallows errors and sleeps. Deltas keep accumulating upstream and are then attributed **in full to the single timestamp of the next successful poll**. A 20-minute Artemis outage therefore produces a 20-minute hole followed by one timestamp holding 20 minutes of traffic. ## KPIs corrupted - **24H Diurnal Flux curve** — a flat gap followed by a vertical spike that never happened. - **Peak Flow Velocity** (`ν_max = max_h(I_h + E_h)`) — the recovery hour wins, permanently. - **Hourly Kinetics matrix** and the Week **7×24 intensity heatmap** — the spike lands in whichever cell the recovery fell in. - **Peak Occupancy timestamp** — pinned to recovery (see the peak-ordering issue). - **Trust rules R1/R2** in `evaluate_cycle_integrity_async` — `FLAG_TRUNCATED_HOURS` (fewer than 10 active hours) and `FLAG_BURST_COUNTER_FLUSH` (any hour above 35% of volume) will both fire, so the cycle is auto-excluded from `k` learning and the month's **Trust Index** drops. The trust engine correctly detects the symptom but attributes it to the sensors rather than to our own polling gap. ## Suggested fix Three options, in order of preference: 1. **Ingest with real timestamps.** If HikCentral exposes a historical/interval passenger flow query (`.../people/advance/...` over a time range), poll that on recovery and backfill events at their true buckets. This is the only option that actually restores the data. 2. **Spread the backlog.** If only the cumulative counter is available, distribute a recovered delta uniformly across the gap `[last_successful_poll, now]` rather than piling it on one instant, and mark those events with a `raw_payload` flag so they are identifiable as reconstructed. 3. **At minimum, record the gap.** Persist a per-cycle ingestion-gap ledger (start, end, delta recovered) and surface it. An hour of reconstructed data must not be presented with the same confidence as an hour of measured data — the deck's 95% CI band should widen across a gap. Whichever is chosen, `evaluate_cycle_integrity_async` should distinguish `FLAG_INGESTION_GAP` (our fault) from `FLAG_BURST_COUNTER_FLUSH` (sensor fault). Right now they are indistinguishable in the calibration ledger. ## Decision needed Which of the three. Option 1 depends on an upstream endpoint we have not yet confirmed exists.
Author
Owner

✅ Design settled (grilling session)

Chosen: Option 2 + Option 3 combined. Option 1 (backfill from a historical query) is dropped — the Artemis catalog exposes no per-interval passenger-count history: only people/resourceGroupRealTimeCount (realtime cumulative), people/advance/resourceGroupList, and people/statisticsHeatMapByTime (a spatial camera heatmap, not a time series). There is no endpoint to backfill from.

Settled spec

  • Gap trigger: reconstruct only when elapsed since the last successful poll ≥ 60 s (≈20 missed 3 s cadences). Below that, keep stamping now — harmless.
  • Spread: distribute the recovered delta proportionally across the hourly buckets the gap spans (aggregation is hourly), not piled on one instant. Tag reconstructed events in raw_payload as reconstructed.
  • Flag split: new FLAG_INGESTION_GAP (our fault) distinct from FLAG_BURST_COUNTER_FLUSH (sensor fault).
  • Trust treatment (agreed middle ground): a reconstructed span is excluded from k learning (it is not measured data) but is not counted as a sensor fault — it must not tank the monthly Trust Index the way a real burst-flush does.
  • Ledger + CI: persist a per-cycle ingestion-gap ledger (start, end, delta); the deck widens its 95% CI band across reconstructed spans so reconstructed hours never read with measured confidence.

Added to scope

  • CONTEXT.md: add glossary entries for ingestion gap (our-fault data hole) vs burst counter flush (sensor-fault), so the two are distinguishable in the domain, not just the code.

Acceptance criteria (supersede the "Decision needed" section)

  • Gap ≥ 60 s triggers reconstruction; shorter gaps stamp now.
  • Recovered delta spread proportionally across spanned hourly buckets; reconstructed events flagged in raw_payload.
  • FLAG_INGESTION_GAP exists and is distinct from FLAG_BURST_COUNTER_FLUSH.
  • Reconstructed spans excluded from k learning but not penalising the Trust Index as sensor faults.
  • Ingestion-gap ledger persisted; deck CI widens across reconstructed spans.
  • CONTEXT.md documents ingestion gap vs burst counter flush.
  • Backend tests: a simulated ≥60 s outage produces spread buckets (no spike), a ledger row, and FLAG_INGESTION_GAP — and does not drop the Trust Index as a sensor fault.

Re-tagged ready-for-agent.

## ✅ Design settled (grilling session) **Chosen: Option 2 + Option 3 combined.** Option 1 (backfill from a historical query) is **dropped** — the Artemis catalog exposes no per-interval passenger-count history: only `people/resourceGroupRealTimeCount` (realtime cumulative), `people/advance/resourceGroupList`, and `people/statisticsHeatMapByTime` (a *spatial* camera heatmap, not a time series). There is no endpoint to backfill from. ### Settled spec - **Gap trigger**: reconstruct only when elapsed since the last successful poll **≥ 60 s** (≈20 missed 3 s cadences). Below that, keep stamping `now` — harmless. - **Spread**: distribute the recovered delta **proportionally across the hourly buckets** the gap spans (aggregation is hourly), not piled on one instant. Tag reconstructed events in `raw_payload` as reconstructed. - **Flag split**: new `FLAG_INGESTION_GAP` (our fault) **distinct** from `FLAG_BURST_COUNTER_FLUSH` (sensor fault). - **Trust treatment** (agreed middle ground): a reconstructed span is **excluded from `k` learning** (it is not measured data) but is **not** counted as a *sensor* fault — it must not tank the monthly Trust Index the way a real burst-flush does. - **Ledger + CI**: persist a per-cycle ingestion-gap ledger `(start, end, delta)`; the deck **widens its 95% CI band** across reconstructed spans so reconstructed hours never read with measured confidence. ### Added to scope - **CONTEXT.md**: add glossary entries for **ingestion gap** (our-fault data hole) vs **burst counter flush** (sensor-fault), so the two are distinguishable in the domain, not just the code. ### Acceptance criteria (supersede the "Decision needed" section) - [ ] Gap ≥ 60 s triggers reconstruction; shorter gaps stamp `now`. - [ ] Recovered delta spread proportionally across spanned hourly buckets; reconstructed events flagged in `raw_payload`. - [ ] `FLAG_INGESTION_GAP` exists and is distinct from `FLAG_BURST_COUNTER_FLUSH`. - [ ] Reconstructed spans excluded from `k` learning but not penalising the Trust Index as sensor faults. - [ ] Ingestion-gap ledger persisted; deck CI widens across reconstructed spans. - [ ] `CONTEXT.md` documents ingestion gap vs burst counter flush. - [ ] Backend tests: a simulated ≥60 s outage produces spread buckets (no spike), a ledger row, and `FLAG_INGESTION_GAP` — and does **not** drop the Trust Index as a sensor fault. Re-tagged `ready-for-agent`.
Author
Owner

Being addressed in draft PR #68 together with #26, #31 and #30 (passenger flow ingestion honesty). The PR description lists the settled spec plus the open design points that still need a decision (restart cold start, a floor for the adaptive cap, a shared anomaly/gap ledger table, and how trust rules see reconstructed hours).

Being addressed in draft **PR #68** together with #26, #31 and #30 (passenger flow ingestion honesty). The PR description lists the settled spec plus the open design points that still need a decision (restart cold start, a floor for the adaptive cap, a shared anomaly/gap ledger table, and how trust rules see reconstructed hours).
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
gabogg/hikcentral#31
No description provided.