[data-veracity] Uncalibrated deployments silently report a foreign site's exit multiplier #36

Closed
opened 2026-09-21 13:43:37 +00:00 by gabogg · 2 comments
Owner

Filed from a data-veracity audit of the ingestion and aggregation pipeline on master, carried out against the KPI set that the Executive Statistics Deck (PR #20 / RFC-ARCH-2026-004) intends to publish. Each issue names the deck KPIs it corrupts.

Problem

k = 1.1162 is hardcoded as the fallback in at least four places:

# app/services/occupancy_service.py:736
def compute_ewma_multiplier(self, history, initial_k: float = 1.1162, gamma: float = 0.25) -> float:
# app/services/occupancy_service.py:778
async def refresh_active_multiplier_async(self, initial_k: float = 1.1162, gamma: float = 0.25)
# app/db/occupancy_repository.py:1006, :1065
k_hat = float(exit_multiplier if exit_multiplier is not None else cfg.get("active_exit_multiplier", 1.1162))

compute_ewma_multiplier returns initial_k unchanged when there is no trusted history, so a fresh deployment — or any site that has not yet accumulated a trusted cycle — publishes calibrated occupancy figures scaled by a constant measured at a different facility.

1.1162 is a real empirical result for this mall's sensors. It is not a neutral default. As a fallback it means the system's least-informed state produces numbers that look calibrated and carry a ~11.6% systematic correction nobody measured.

Why it is hard to notice

  • get_sample_maturity_info correctly reports INITIAL (0/14 jornadas), but that badge lives in the admin UI, not on the deck.
  • The deck's headline tiles (Calibrated Residual, Peak Occupancy, Mean Dwell) will all be computed with k = 1.1162 and displayed with no indication that the multiplier is assumed rather than derived.
  • compute_multiplier_variance returns the default 0.0002 when N < 2, so the 95% CI band renders narrow — the system is at its least certain precisely when it draws its most confident-looking band.

KPIs corrupted

Calibrated Residual · Peak Occupancy · Riemann Ō · Mean Dwell · the 95% CI band width · k̂ stability tile — all of them, on any deployment before calibration matures.

Suggested fix

  1. Default k to 1.0 (identity) when there is no trusted history. An uncalibrated system should report raw counts, which are honest, rather than counts scaled by a borrowed constant.
  2. Move the site's empirical seed into occupancy_config (initial_exit_multiplier) rather than into function signatures, so it is a deployment decision with a visible value.
  3. Widen, don't narrow, the default variance when N < 2 — the CI should be at its widest when the sample is empty. Deriving the default from the guardrail range ([0.90, 1.35]) would be defensible; 0.0002 is not.
  4. Propagate maturity to the deck: when trusted_days_count < 7, the calibrated tiles should carry an UNCALIBRATED state chip. RFC-ARCH-2026-004 has no provision for this and should gain one.

Decision needed

Whether an uncalibrated deployment should show raw numbers (recommended) or refuse to show calibrated tiles at all.

> Filed from a data-veracity audit of the ingestion and aggregation pipeline on `master`, carried out against the KPI set that the Executive Statistics Deck (PR #20 / `RFC-ARCH-2026-004`) intends to publish. Each issue names the deck KPIs it corrupts. ## Problem `k = 1.1162` is hardcoded as the fallback in at least four places: ```python # app/services/occupancy_service.py:736 def compute_ewma_multiplier(self, history, initial_k: float = 1.1162, gamma: float = 0.25) -> float: # app/services/occupancy_service.py:778 async def refresh_active_multiplier_async(self, initial_k: float = 1.1162, gamma: float = 0.25) # app/db/occupancy_repository.py:1006, :1065 k_hat = float(exit_multiplier if exit_multiplier is not None else cfg.get("active_exit_multiplier", 1.1162)) ``` `compute_ewma_multiplier` returns `initial_k` unchanged when there is no trusted history, so a fresh deployment — or any site that has not yet accumulated a trusted cycle — publishes **calibrated** occupancy figures scaled by a constant measured at a different facility. `1.1162` is a real empirical result for *this* mall's sensors. It is not a neutral default. As a fallback it means the system's least-informed state produces numbers that look calibrated and carry a ~11.6% systematic correction nobody measured. ## Why it is hard to notice - `get_sample_maturity_info` correctly reports `INITIAL (0/14 jornadas)`, but that badge lives in the admin UI, not on the deck. - The deck's headline tiles (Calibrated Residual, Peak Occupancy, Mean Dwell) will all be computed with `k = 1.1162` and displayed with no indication that the multiplier is assumed rather than derived. - `compute_multiplier_variance` returns the default `0.0002` when `N < 2`, so the 95% CI band renders **narrow** — the system is at its least certain precisely when it draws its most confident-looking band. ## KPIs corrupted Calibrated Residual · Peak Occupancy · Riemann Ō · Mean Dwell · the 95% CI band width · k̂ stability tile — all of them, on any deployment before calibration matures. ## Suggested fix 1. Default `k` to **1.0** (identity) when there is no trusted history. An uncalibrated system should report raw counts, which are honest, rather than counts scaled by a borrowed constant. 2. Move the site's empirical seed into `occupancy_config` (`initial_exit_multiplier`) rather than into function signatures, so it is a deployment decision with a visible value. 3. Widen, don't narrow, the default variance when `N < 2` — the CI should be at its widest when the sample is empty. Deriving the default from the guardrail range (`[0.90, 1.35]`) would be defensible; `0.0002` is not. 4. Propagate maturity to the deck: when `trusted_days_count < 7`, the calibrated tiles should carry an `UNCALIBRATED` state chip. `RFC-ARCH-2026-004` has no provision for this and should gain one. ## Decision needed Whether an uncalibrated deployment should show raw numbers (recommended) or refuse to show calibrated tiles at all.
Author
Owner

✅ Design settled (grilling session)

Note — reversed from this issue's original suggestion. The issue proposed defaulting k to 1.0 (identity) for uncalibrated deployments. Decision: do not fall back to 1.0. Apply the seeded k from day one and mark the figures explicitly unreliable until calibration matures — one consistent formula, honesty via a state chip and a wide CI.

Settled spec

  • Seed applied from day one, sourced from occupancy_config.initial_exit_multiplier (a visible per-deployment value), not from hardcoded function defaults. Remove the scattered 1.1162 literals (occupancy_service.py:736,778, occupancy_repository.py:1006,1065). This mall's config carries 1.1162; a fresh deployment defaults to 1.0 and the operator sets a measured value when they have one. (Biased to this site for now — accepted.)
  • UNCALIBRATED / SIN CALIBRAR state chip on every calibrated tile (Calibrated Residual, Peak Occupancy, Mean Dwell, k̂ stability) until 14 trusted business cycles. (RFC-ARCH-2026-004 gains this state.)
  • Wide default CI when N < 2: derive from the guardrail — SD ≈ (1.35 − 0.90) / 4 ≈ 0.11 — replacing the overconfident 0.0002. Widest band precisely when the sample is empty.
  • Guardrail range: keep the running [0.90, 1.35] (occupancy_service.py:753,941); fix CONTEXT.md:88's stale [0.80, 1.30] to match. Resolves a live code/doc contradiction. Changing the live clamp is avoided so deployed calibration doesn't shift.

Added to scope

  • CONTEXT.md: correct the guardrail range; add an uncalibrated state glossary entry; canonical business cycle wording (see the naming issue).
  • ADR (shared with #32): the calibration-honesty policy is recorded in the shared ADR "Published occupancy figures — measurement window & calibration honesty."

Acceptance criteria (supersede "Decision needed")

  • Seed k applied from day one from occupancy_config.initial_exit_multiplier; no hardcoded 1.1162 in function signatures.
  • Fresh-deployment config defaults to 1.0; this deployment's config carries 1.1162.
  • UNCALIBRATED/SIN CALIBRAR chip on calibrated tiles until 14 trusted business cycles.
  • Default variance when N < 2 derived from the guardrail (SD ≈ 0.11); CI renders wide.
  • Guardrail stays [0.90, 1.35]; CONTEXT.md updated to match.
  • Shared ADR authored/referenced.
  • Tests: an uncalibrated deployment renders tiles with the chip and a wide CI; a fresh deployment applies k = 1.0, not a borrowed constant.

Re-tagged ready-for-agent.

## ✅ Design settled (grilling session) **Note — reversed from this issue's original suggestion.** The issue proposed defaulting `k` to `1.0` (identity) for uncalibrated deployments. Decision: **do not** fall back to 1.0. Apply the **seeded `k` from day one** and mark the figures explicitly unreliable until calibration matures — one consistent formula, honesty via a state chip and a wide CI. ### Settled spec - **Seed applied from day one**, sourced from **`occupancy_config.initial_exit_multiplier`** (a visible per-deployment value), not from hardcoded function defaults. Remove the scattered `1.1162` literals (`occupancy_service.py:736,778`, `occupancy_repository.py:1006,1065`). **This** mall's config carries `1.1162`; a **fresh** deployment defaults to `1.0` and the operator sets a measured value when they have one. (Biased to this site for now — accepted.) - **`UNCALIBRATED` / `SIN CALIBRAR` state chip** on every calibrated tile (Calibrated Residual, Peak Occupancy, Mean Dwell, k̂ stability) until **14 trusted business cycles**. (`RFC-ARCH-2026-004` gains this state.) - **Wide default CI when `N < 2`**: derive from the guardrail — SD ≈ `(1.35 − 0.90) / 4 ≈ 0.11` — replacing the overconfident `0.0002`. Widest band precisely when the sample is empty. - **Guardrail range**: keep the **running `[0.90, 1.35]`** (`occupancy_service.py:753,941`); **fix CONTEXT.md:88's stale `[0.80, 1.30]`** to match. Resolves a live code/doc contradiction. Changing the live clamp is avoided so deployed calibration doesn't shift. ### Added to scope - **CONTEXT.md**: correct the guardrail range; add an **uncalibrated state** glossary entry; canonical **business cycle** wording (see the naming issue). - **ADR (shared with #32)**: the calibration-honesty policy is recorded in the shared ADR *"Published occupancy figures — measurement window & calibration honesty."* ### Acceptance criteria (supersede "Decision needed") - [ ] Seed `k` applied from day one from `occupancy_config.initial_exit_multiplier`; no hardcoded `1.1162` in function signatures. - [ ] Fresh-deployment config defaults to `1.0`; this deployment's config carries `1.1162`. - [ ] `UNCALIBRATED`/`SIN CALIBRAR` chip on calibrated tiles until 14 trusted business cycles. - [ ] Default variance when `N < 2` derived from the guardrail (SD ≈ 0.11); CI renders wide. - [ ] Guardrail stays `[0.90, 1.35]`; `CONTEXT.md` updated to match. - [ ] Shared ADR authored/referenced. - [ ] Tests: an uncalibrated deployment renders tiles with the chip and a wide CI; a fresh deployment applies `k = 1.0`, not a borrowed constant. Re-tagged `ready-for-agent`.
Author
Owner

Being addressed in draft PR #70, one of four [data-veracity] drafts declared on 2026-09-23 (#68, #69, #70, #71). Each will be triaged, reviewed and implemented in order; the PR description lists the open design points to settle first.

Being addressed in draft **PR #70**, one of four [data-veracity] drafts declared on 2026-09-23 (#68, #69, #70, #71). Each will be triaged, reviewed and implemented in order; the PR description lists the open design points to settle first.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
gabogg/hikcentral#36
No description provided.