[data-veracity] Trust Index conflates sensor faults with quiet business days, on hardcoded thresholds #33

Closed
opened 2026-09-21 13:43:34 +00:00 by gabogg · 1 comment
Owner

Filed from a data-veracity audit of the ingestion and aggregation pipeline on master, carried out against the KPI set that the Executive Statistics Deck (PR #20 / RFC-ARCH-2026-004) intends to publish. Each issue names the deck KPIs it corrupts.

Problem

The deck publishes Trust Index as a headline Month KPI (trusted_cycles_ratio, calibration_trust_index) and renders converged/failed nights in the calibration ledger. That index comes from evaluate_cycle_integrity_async (app/services/occupancy_service.py:826), whose four rules mix two unrelated concepts:

# R1 Temporal dispersion
if active_hours < 10 and total_vol > 0:            flags.append("FLAG_TRUNCATED_HOURS")
# R2 Burst concentration
if (max_hourly_vol / total_vol) > 0.35:            flags.append("FLAG_BURST_COUNTER_FLUSH")
# R3 Ratio bounds
if ratio < 0.80 or ratio > 1.30:                   flags.append("FLAG_RATIO_OUT_OF_BOUNDS")
# R4 Minimum footfall
min_entries = 1000 if is_holiday else 5000
if total_in < min_entries and total_vol > 0:       flags.append("FLAG_INSUFFICIENT_VOLUME")

R2 and R3 are data-quality signals — they detect counter flushes and directional sensor bias. R1 and R4 are business-activity signals: a genuinely short trading day or a genuinely quiet Tuesday trips them even when every sensor behaved perfectly.

All four collapse into is_trusted = len(flags) == 0, which decides:

  • whether the cycle is AUTO_EXCLUDED from compute_ewma_multiplier (so k never learns from quiet days), and
  • the Trust Index the deck publishes as a measure of calibration health.

A month with several quiet days reports a degraded Trust Index and a red-marked calibration ledger, implying a sensor problem that does not exist.

Hardcoded site-specific thresholds

5000 / 1000 entries, 10 active hours, 0.35 burst share, [0.80, 1.30] ratio bounds — none are in occupancy_config. A second deployment, a smaller mall, or a renovation period silently mass-quarantines cycles.

The same pattern repeats in SQL:

-- app/db/occupancy_repository.py:1410  quarantine_and_deduplicate_calibration_logs_async
WHERE is_trusted = 1 AND ( raw_net_flow < -1000 OR total_in > 40000
                           OR computed_exit_multiplier < 0.5 OR computed_exit_multiplier > 2.0
                           OR completeness_score < 90 )

total_in > 40000 auto-quarantines any exceptionally busy day as a counter flush. A large mall on a holiday weekend can legitimately exceed that — and the busiest days are precisely the ones whose peak the deck most wants to report.

KPIs corrupted

Trust Index (Month headline) · trusted_cycles_ratio on every period · the nocturnal calibration ledger's status column · k̂ itself, because quiet and very busy cycles are excluded from EWMA learning, biasing the multiplier toward mid-volume days.

Suggested fix

  1. Split the flag vocabulary into DATA_QUALITY_* (R2, R3, counter-flush, ratio bounds) and LOW_ACTIVITY_* (R1, R4). Only DATA_QUALITY_* should clear is_trusted and gate EWMA learning.
  2. Publish two separate figures: Data Trust (sensor reliability) and Cycle Completeness (business coverage). The deck currently has one tile where it needs two, or one tile with an explicit definition.
  3. Move every threshold into occupancy_config with the current values as defaults, and surface them in the admin UI.
  4. Make total_in > 40000 a relative rule — e.g. more than 3σ above the trailing 30-cycle mean — so it scales with the site instead of assuming one.
> Filed from a data-veracity audit of the ingestion and aggregation pipeline on `master`, carried out against the KPI set that the Executive Statistics Deck (PR #20 / `RFC-ARCH-2026-004`) intends to publish. Each issue names the deck KPIs it corrupts. ## Problem The deck publishes **Trust Index** as a headline Month KPI (`trusted_cycles_ratio`, `calibration_trust_index`) and renders converged/failed nights in the calibration ledger. That index comes from `evaluate_cycle_integrity_async` (`app/services/occupancy_service.py:826`), whose four rules mix two unrelated concepts: ```python # R1 Temporal dispersion if active_hours < 10 and total_vol > 0: flags.append("FLAG_TRUNCATED_HOURS") # R2 Burst concentration if (max_hourly_vol / total_vol) > 0.35: flags.append("FLAG_BURST_COUNTER_FLUSH") # R3 Ratio bounds if ratio < 0.80 or ratio > 1.30: flags.append("FLAG_RATIO_OUT_OF_BOUNDS") # R4 Minimum footfall min_entries = 1000 if is_holiday else 5000 if total_in < min_entries and total_vol > 0: flags.append("FLAG_INSUFFICIENT_VOLUME") ``` R2 and R3 are **data-quality** signals — they detect counter flushes and directional sensor bias. R1 and R4 are **business-activity** signals: a genuinely short trading day or a genuinely quiet Tuesday trips them even when every sensor behaved perfectly. All four collapse into `is_trusted = len(flags) == 0`, which decides: - whether the cycle is `AUTO_EXCLUDED` from `compute_ewma_multiplier` (so `k` never learns from quiet days), and - the Trust Index the deck publishes as a measure of *calibration health*. A month with several quiet days reports a degraded Trust Index and a red-marked calibration ledger, implying a sensor problem that does not exist. ## Hardcoded site-specific thresholds `5000` / `1000` entries, `10` active hours, `0.35` burst share, `[0.80, 1.30]` ratio bounds — none are in `occupancy_config`. A second deployment, a smaller mall, or a renovation period silently mass-quarantines cycles. The same pattern repeats in SQL: ```sql -- app/db/occupancy_repository.py:1410 quarantine_and_deduplicate_calibration_logs_async WHERE is_trusted = 1 AND ( raw_net_flow < -1000 OR total_in > 40000 OR computed_exit_multiplier < 0.5 OR computed_exit_multiplier > 2.0 OR completeness_score < 90 ) ``` `total_in > 40000` auto-quarantines any exceptionally busy day as a counter flush. A large mall on a holiday weekend can legitimately exceed that — and the *busiest* days are precisely the ones whose peak the deck most wants to report. ## KPIs corrupted Trust Index (Month headline) · `trusted_cycles_ratio` on every period · the nocturnal calibration ledger's status column · `k̂` itself, because quiet and very busy cycles are excluded from EWMA learning, biasing the multiplier toward mid-volume days. ## Suggested fix 1. Split the flag vocabulary into `DATA_QUALITY_*` (R2, R3, counter-flush, ratio bounds) and `LOW_ACTIVITY_*` (R1, R4). Only `DATA_QUALITY_*` should clear `is_trusted` and gate EWMA learning. 2. Publish two separate figures: **Data Trust** (sensor reliability) and **Cycle Completeness** (business coverage). The deck currently has one tile where it needs two, or one tile with an explicit definition. 3. Move every threshold into `occupancy_config` with the current values as defaults, and surface them in the admin UI. 4. Make `total_in > 40000` a *relative* rule — e.g. more than 3σ above the trailing 30-cycle mean — so it scales with the site instead of assuming one.
Author
Owner

Being addressed in draft PR #71, one of four [data-veracity] drafts declared on 2026-09-23 (#68, #69, #70, #71). Each will be triaged, reviewed and implemented in order; the PR description lists the open design points to settle first.

Being addressed in draft **PR #71**, one of four [data-veracity] drafts declared on 2026-09-23 (#68, #69, #70, #71). Each will be triaged, reviewed and implemented in order; the PR description lists the open design points to settle first.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
gabogg/hikcentral#33
No description provided.