[data-veracity] Business cycle boundaries use server local time and ignore facility_utc_offset_minutes #27

Closed
opened 2026-09-21 13:43:27 +00:00 by gabogg · 1 comment
Owner

Filed from a data-veracity audit of the ingestion and aggregation pipeline on master, carried out against the KPI set that the Executive Statistics Deck (PR #20 / RFC-ARCH-2026-004) intends to publish. Each issue names the deck KPIs it corrupts.

Problem

Every business-cycle boundary is computed in the server's timezone:

# app/services/occupancy_service.py:117  get_business_day_epoch_bounds
st = time.localtime(t)
today_reset_st = time.struct_time((st.tm_year, st.tm_mon, st.tm_mday, reset_h, reset_m, 0, ...))
today_reset_epoch = time.mktime(today_reset_st)

Same pattern in get_completed_business_cycle_bounds (:160), get_active_schedule_info_async (:190), get_today_midnight_epoch (:110), and in SQL via DATE(..., 'unixepoch', 'localtime') in quarantine_and_deduplicate_calibration_logs_async.

Meanwhile settings.facility_utc_offset_minutes exists and is broadcast to every client (monitor_service.py:39,49,117,151, occupancy_service.py:550,995,1040,1335,1693). Its existence is an explicit acknowledgement that the server clock and the facility clock are not assumed to be the same — but it is used only for display, never for cycle math.

Why this destroys the data, not just the labels

If the gateway runs UTC (the default for most VPS and container images) and the facility is UTC−4, then:

  • The "04:00 reset" fires at 00:00 facility time. Every daily KPI slices the wrong 24 hours, and the mall's evening peak is split across two reported cycles.
  • The nocturnal quiet window (03:30–04:30) runs at 23:30–00:30 facility time — while the mall is still emptying, not when it is verifiably empty.

The second one is the serious one: the quiet-window residual Δ = I − k·E is the entire basis of the proportional calibration model. Evaluating it during closing traffic produces a large spurious residual, which drives k away from truth via compute_ewma_multiplier, which then mis-scales every calibrated occupancy figure the deck publishes.

Secondary: DST and the exact-86400 assumption

time.struct_time is built by copying tm_isdst from the current time and then passed to mktime; around a DST transition this yields an ambiguous or wrong epoch. The code then hardcodes cycle_start + 86400.0, and RFC-ARCH-2026-004 §2.2 asserts "Duration is exactly 86,400 seconds". On a DST-observing deployment a business cycle is genuinely 23 or 25 hours. Venezuela does not observe DST, so impact is currently nil — but the assumption is baked into both the code and the RFC and should be stated rather than assumed.

KPIs corrupted

All of them, plus cycle attribution itself (traffic assigned to the wrong day), plus k and therefore Trust Index and the calibrated occupancy envelope.

Suggested fix

  1. Introduce a single facility_now() / facility_bounds() seam that applies settings.facility_utc_offset_minutes (or, better, a real IANA facility_timezone setting) and route all cycle math through it.
  2. Replace time.mktime + copied tm_isdst with explicit datetime + ZoneInfo arithmetic.
  3. Replace SQL 'localtime' modifiers with epoch ranges computed by that seam.
  4. Add a startup assertion that logs server TZ, facility offset, and the resolved next reset boundary, so a misconfiguration is visible on boot rather than a month later in the numbers.
  5. Add a test that pins the server to UTC, the facility to UTC−4, and asserts the cycle boundary lands at 04:00 facility time.
> Filed from a data-veracity audit of the ingestion and aggregation pipeline on `master`, carried out against the KPI set that the Executive Statistics Deck (PR #20 / `RFC-ARCH-2026-004`) intends to publish. Each issue names the deck KPIs it corrupts. ## Problem Every business-cycle boundary is computed in the **server's** timezone: ```python # app/services/occupancy_service.py:117 get_business_day_epoch_bounds st = time.localtime(t) today_reset_st = time.struct_time((st.tm_year, st.tm_mon, st.tm_mday, reset_h, reset_m, 0, ...)) today_reset_epoch = time.mktime(today_reset_st) ``` Same pattern in `get_completed_business_cycle_bounds` (`:160`), `get_active_schedule_info_async` (`:190`), `get_today_midnight_epoch` (`:110`), and in SQL via `DATE(..., 'unixepoch', 'localtime')` in `quarantine_and_deduplicate_calibration_logs_async`. Meanwhile `settings.facility_utc_offset_minutes` exists and is broadcast to every client (`monitor_service.py:39,49,117,151`, `occupancy_service.py:550,995,1040,1335,1693`). Its existence is an explicit acknowledgement that **the server clock and the facility clock are not assumed to be the same** — but it is used only for display, never for cycle math. ## Why this destroys the data, not just the labels If the gateway runs UTC (the default for most VPS and container images) and the facility is UTC−4, then: - The "04:00 reset" fires at **00:00 facility time**. Every daily KPI slices the wrong 24 hours, and the mall's evening peak is split across two reported cycles. - The nocturnal quiet window (03:30–04:30) runs at **23:30–00:30 facility time** — while the mall is still emptying, not when it is verifiably empty. The second one is the serious one: the quiet-window residual `Δ = I − k·E` is the entire basis of the proportional calibration model. Evaluating it during closing traffic produces a large spurious residual, which drives `k` away from truth via `compute_ewma_multiplier`, which then mis-scales **every** calibrated occupancy figure the deck publishes. ## Secondary: DST and the exact-86400 assumption `time.struct_time` is built by copying `tm_isdst` from the current time and then passed to `mktime`; around a DST transition this yields an ambiguous or wrong epoch. The code then hardcodes `cycle_start + 86400.0`, and `RFC-ARCH-2026-004` §2.2 asserts *"Duration is exactly 86,400 seconds"*. On a DST-observing deployment a business cycle is genuinely 23 or 25 hours. Venezuela does not observe DST, so impact is currently nil — but the assumption is baked into both the code and the RFC and should be stated rather than assumed. ## KPIs corrupted All of them, plus cycle attribution itself (traffic assigned to the wrong day), plus `k` and therefore Trust Index and the calibrated occupancy envelope. ## Suggested fix 1. Introduce a single `facility_now()` / `facility_bounds()` seam that applies `settings.facility_utc_offset_minutes` (or, better, a real IANA `facility_timezone` setting) and route **all** cycle math through it. 2. Replace `time.mktime` + copied `tm_isdst` with explicit `datetime` + `ZoneInfo` arithmetic. 3. Replace SQL `'localtime'` modifiers with epoch ranges computed by that seam. 4. Add a startup assertion that logs server TZ, facility offset, and the resolved next reset boundary, so a misconfiguration is visible on boot rather than a month later in the numbers. 5. Add a test that pins the server to UTC, the facility to UTC−4, and asserts the cycle boundary lands at 04:00 facility time.
Author
Owner

Being addressed in draft PR #69, one of four [data-veracity] drafts declared on 2026-09-23 (#68, #69, #70, #71). Each will be triaged, reviewed and implemented in order; the PR description lists the open design points to settle first.

Being addressed in draft **PR #69**, one of four [data-veracity] drafts declared on 2026-09-23 (#68, #69, #70, #71). Each will be triaged, reviewed and implemented in order; the PR description lists the open design points to settle first.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
gabogg/hikcentral#27
No description provided.