[data-veracity] Trust rules and KPI queries read different datasets (missing camera join) #35
Labels
No labels
blocked
bug
enhancement
high-priority
low-priority
needs-info
needs-triage
ready-for-agent
ready-for-human
referenced
research
wontfix
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
gabogg/hikcentral#35
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Problem
Every aggregate that feeds a published KPI joins
counting_camerasand filters excluded/inactive cameras:get_hourly_flow_distribution_async(app/db/occupancy_repository.py:1391) does not:No join, no exclusion filter. It therefore sums events from excluded cameras, inactive cameras, and orphan resource-group codes that no KPI counts.
Why it matters
This query is the sole input to trust rules R1 (temporal dispersion) and R2 (burst concentration) in
evaluate_cycle_integrity_async. So the engine that decides whether a cycle is trustworthy is looking at a different dataset from the one whose numbers get published. Excluding a malfunctioning camera removes it from the KPIs but leaves it driving the trust verdict — which is the opposite of the intent, since exclusion exists precisely to quarantine a bad sensor.It also means the hourly volumes used for trust evaluation will not reconcile against the Gross Ingress/Egress the deck shows for the same cycle, which will look like a bug to whoever first compares them.
Also: it cannot be reused for the deck
The query returns a single
SUM(count)with no direction split, and it skips hours with no events entirely (solen(rows)is "active hours", not 24). The deck's diurnal series needs 24 dense buckets split by direction with occupancy and CI bounds per bucket —DiurnalTimeseriesBucketinapp/schemas/occupancy_models.py.get_bucketed_cycle_flow_async(:1471) already has the correct join, filter, direction split and bucket parameterisation. Phase 2 should build the diurnal series on that, and this issue should not be worked around by extending the unfiltered query.Suggested fix
JOIN counting_cameras ... AND c.is_excluded = 0 AND c.is_active = 1toget_hourly_flow_distribution_async.get_bucketed_cycle_flow_async(bucket_seconds=3600), so there is exactly one query shape for "counts over time" and it cannot drift again.gabogg referenced this issue2026-09-22 15:22:42 +00:00