fix(calibration): calibration audit trail and trust vocabulary (#34, #33) #71

Merged
gabogg merged 15 commits from fix/calibration-audit-trail into master 2026-09-24 22:33:11 +00:00
Owner

Closes #34
Closes #33

Implemented and ready for review.


Problem

#34: calibration log "maintenance" destroys and fabricates audit data. quarantine_and_deduplicate_calibration_logs_async (occupancy_repository.py:1525), with the same logic at startup in database.py:433-448, does three things:

  1. Hard-deletes all but the newest log per cycle_date, including operator MANUAL_OVERRIDE rows.
  2. Backfills cycle_date as DATE(timestamp − 86400, 'localtime') (:1543, database.py:433). That's wrong for any run that isn't the 04:15 nocturnal one, and it's in server time (#27).
  3. Writes the raw I/E ratio into computed_exit_multiplier wherever it is NULL or exactly 1.0. A genuine 1.000 convergence is indistinguishable from "unset", and the k̂ drift chart plots backfilled ratios as convergence results.

#33: the Trust Index mixes sensor faults with quiet business days. evaluate_cycle_integrity_async (occupancy_service.py:967) folds four rules into one is_trusted flag:

  • R2/R3 are data quality: counter flush, ratio bounds.
  • R1/R4 are business activity: short day, low footfall (5000 / 1000, :1003).

A quiet Tuesday is therefore excluded from k learning and dents the published Trust Index. All thresholds are hardcoded. total_in > 40000 (occupancy_repository.py:1564, database.py:448) auto-quarantines the busiest days as counter flushes.

Approach (from the issues' suggested fixes; not yet triaged)

#34

  • Never delete: a superseded_by / is_current column; retention (prune_retention_data_async, 365 days) stays the only deletion.
  • Never overwrite or supersede a MANUAL_OVERRIDE.
  • Resolve cycle_date through the business-cycle seam (#27), not timestamp − 86400.
  • Make computed_exit_multiplier nullable, stop backfilling, and put the raw ratio in its own column (raw_io_ratio).
  • Test: a MANUAL_OVERRIDE survives a dedup pass.

#33

  • Split flags into DATA_QUALITY_* (R2, R3, counter flush) and LOW_ACTIVITY_* (R1, R4); only data-quality flags clear is_trusted and gate EWMA learning.
  • Publish two figures, Data Trust and Cycle Completeness, instead of one tile.
  • Move every threshold into occupancy_config (current values as defaults), surfaced in the admin UI.
  • Make the 40000 rule relative (e.g. > 3σ above the trailing 30-cycle mean).

Implementation notes

Audit rows are retained and automatic rows can be superseded; manual override rows remain intact. Raw I/E ratio has its own nullable field. Quiet cycles and sensor faults have separate flags and scores. The API and admin controls expose thresholds and scores; separate Statistics Deck tiles remain assigned to PR #20.

Acceptance criteria

#34

  • Dedup never deletes; superseded rows are kept and marked.
  • A MANUAL_OVERRIDE survives any maintenance pass (test).
  • cycle_date comes from the business-cycle seam.
  • No backfilled ratio in computed_exit_multiplier; unset is distinguishable from 1.0.

#33

  • Data-quality and low-activity flags separated; only data-quality flags clear is_trusted and gate EWMA.
  • Data Trust and Cycle Completeness exposed separately.
  • All thresholds in occupancy_config with current defaults.
  • The busiest-day rule is relative to the site.

All

  • Full suite green.

Sequencing across the [data-veracity] drafts

Declared together on 2026-09-23; each one is triaged, then reviewed, then implemented, in this order:

  1. #68: ingestion honesty (#26, #31, #30). Owns the sync loop and the peak walk.
  2. #69: business cycles in facility time (#27). Introduces the cycle-boundary seam that #32 and #34 then use.
  3. #70: calibration honesty (#36, #32). Changes the occupancy formula that #30's peak walk (in #68) calls.
  4. #71: calibration audit trail and trust vocabulary (#34, #33). Builds on #31's FLAG_INGESTION_GAP (in #68) and on #27's seam for cycle_date.

Review the PRs in this order; no PR has been merged into master.

🤖 Generated with Claude Code

Closes #34 Closes #33 **Implemented and ready for review.** --- ## Problem **#34: calibration log "maintenance" destroys and fabricates audit data.** `quarantine_and_deduplicate_calibration_logs_async` (`occupancy_repository.py:1525`), with the same logic at startup in `database.py:433-448`, does three things: 1. **Hard-deletes all but the newest log per `cycle_date`**, including operator `MANUAL_OVERRIDE` rows. 2. **Backfills `cycle_date` as `DATE(timestamp − 86400, 'localtime')`** (`:1543`, `database.py:433`). That's wrong for any run that isn't the 04:15 nocturnal one, and it's in server time (#27). 3. **Writes the raw `I/E` ratio into `computed_exit_multiplier`** wherever it is `NULL` **or exactly `1.0`**. A genuine `1.000` convergence is indistinguishable from "unset", and the k̂ drift chart plots backfilled ratios as convergence results. **#33: the Trust Index mixes sensor faults with quiet business days.** `evaluate_cycle_integrity_async` (`occupancy_service.py:967`) folds four rules into one `is_trusted` flag: - **R2/R3 are data quality:** counter flush, ratio bounds. - **R1/R4 are business activity:** short day, low footfall (`5000` / `1000`, `:1003`). A quiet Tuesday is therefore excluded from `k` learning and dents the published Trust Index. All thresholds are hardcoded. `total_in > 40000` (`occupancy_repository.py:1564`, `database.py:448`) auto-quarantines the **busiest** days as counter flushes. ## Approach (from the issues' suggested fixes; not yet triaged) **#34** - Never delete: a `superseded_by` / `is_current` column; retention (`prune_retention_data_async`, 365 days) stays the only deletion. - Never overwrite or supersede a `MANUAL_OVERRIDE`. - Resolve `cycle_date` through the business-cycle seam (#27), not `timestamp − 86400`. - Make `computed_exit_multiplier` nullable, stop backfilling, and put the raw ratio in its own column (`raw_io_ratio`). - Test: a `MANUAL_OVERRIDE` survives a dedup pass. **#33** - Split flags into `DATA_QUALITY_*` (R2, R3, counter flush) and `LOW_ACTIVITY_*` (R1, R4); only data-quality flags clear `is_trusted` and gate EWMA learning. - Publish two figures, **Data Trust** and **Cycle Completeness**, instead of one tile. - Move every threshold into `occupancy_config` (current values as defaults), surfaced in the admin UI. - Make the `40000` rule relative (e.g. > 3σ above the trailing 30-cycle mean). ## Implementation notes Audit rows are retained and automatic rows can be superseded; manual override rows remain intact. Raw I/E ratio has its own nullable field. Quiet cycles and sensor faults have separate flags and scores. The API and admin controls expose thresholds and scores; separate Statistics Deck tiles remain assigned to PR #20. ## Acceptance criteria **#34** - [x] Dedup never deletes; superseded rows are kept and marked. - [x] A `MANUAL_OVERRIDE` survives any maintenance pass (test). - [x] `cycle_date` comes from the business-cycle seam. - [x] No backfilled ratio in `computed_exit_multiplier`; unset is distinguishable from `1.0`. **#33** - [x] Data-quality and low-activity flags separated; only data-quality flags clear `is_trusted` and gate EWMA. - [x] Data Trust and Cycle Completeness exposed separately. - [x] All thresholds in `occupancy_config` with current defaults. - [x] The busiest-day rule is relative to the site. **All** - [x] Full suite green. ## Sequencing across the `[data-veracity]` drafts Declared together on 2026-09-23; each one is triaged, then reviewed, then implemented, in this order: 1. **#68**: ingestion honesty (#26, #31, #30). Owns the sync loop and the peak walk. 2. **#69**: business cycles in facility time (#27). Introduces the cycle-boundary seam that #32 and #34 then use. 3. **#70**: calibration honesty (#36, #32). Changes the occupancy formula that #30's peak walk (in #68) calls. 4. **#71**: calibration audit trail and trust vocabulary (#34, #33). Builds on #31's `FLAG_INGESTION_GAP` (in #68) and on #27's seam for `cycle_date`. Review the PRs in this order; no PR has been merged into master. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
chore(wip): open draft for calibration audit trail and trust vocabulary (#34, #33)
All checks were successful
CI / lint-and-test (pull_request) Successful in 1m15s
b72e1d14d9
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
gabogg changed title from WIP: fix(calibration): calibration audit trail and trust vocabulary (#34, #33) to fix(calibration): calibration audit trail and trust vocabulary (#34, #33) 2026-09-24 14:25:42 +00:00
Author
Owner

Implementation is pushed at b5a649a. This branch includes the preceding data-veracity commits; review #68–#70 first. The API exposes Data Trust and Cycle Completeness and the admin UI exposes thresholds. Separate Statistics Deck tiles remain assigned to PR #20. No PR has been merged into master.

Implementation is pushed at `b5a649a`. This branch includes the preceding data-veracity commits; review #68–#70 first. The API exposes Data Trust and Cycle Completeness and the admin UI exposes thresholds. Separate Statistics Deck tiles remain assigned to PR #20. No PR has been merged into `master`.
Author
Owner

Standards

Documented Standard Violations (Hard)

  • app/db/occupancy_repository.py:1642 — Layer Separation Breach (Business Logic in Repository):
    • Standard: docs/standards/code-standards.md §1.1 & AGENTS.md §1 (Layer Separation & Boundaries: Services vs Repositories).
    • Rule: Services handle domain workflows and business calculations; repositories are strictly responsible for querying and persisting data into SQLite.
    • Breach: In quarantine_and_deduplicate_calibration_logs_async, the repository performs rolling statistical analysis (statistics.mean, statistics.pstdev), evaluates z-score spike thresholds, and assigns domain anomaly classifications. Domain calculations belong in OccupancyManager.
  • app/db/occupancy_repository.py (record_calibration_log_sync/_async) — Untyped Dict Mutation:
    • Standard: docs/standards/code-standards.md §2.3 (Pydantic Schemas: Avoid passing untyped, raw dictionaries between service layers).
    • Breach: When entry: CalibrationLogEntry is provided, it is dumped to an untyped dict via entry.model_dump(), mutated, and forwarded as raw **kwargs with entry = None.

Baseline Smells (Judgement Calls)

  • Feature Envy (app/db/occupancy_repository.py): Repository contains statistical anomaly evaluation methods instead of delegating to domain services.
  • Duplicated Code & Divergent Change (app/schemas/models.py vs app/schemas/occupancy_models.py): CalibrationLogEntry is redundantly declared in two separate schema modules. Field superseded_by was added only to models.py, while occupancy_repository.py imports CalibrationLogEntry from occupancy_models.py where superseded_by is missing.
  • Primitive Obsession (app/services/occupancy_service.py & app/db/occupancy_repository.py):
    • Uses raw string prefix checks (flag.startswith("DATA_QUALITY_")) rather than Enum member comparisons.
    • In occupancy_repository.py, uses "DATA_QUALITY_RATIO" string literal instead of CalibrationAnomalyFlag.DATA_QUALITY_RATIO_OUT_OF_BOUNDS.
  • Data Clumps (app/db/occupancy_repository.py): Over 10 trust threshold parameters travel together through configuration lookups and query arguments instead of grouping into a typed trust settings model.

Spec

Missing or Partial Requirements

  • Missing is_current Column (#34): Spec states: "superseded rows kept with superseded_by and is_current columns". Only superseded_by was added; is_current was omitted from the SQLite schema (database.py), models, and queries. Queries rely solely on superseded_by IS NULL.
  • Partial Data Trust & Cycle Completeness Publishing (#33): Spec states: "Publish two figures, Data Trust and Cycle Completeness, instead of one tile." Scores are stored in DB columns and returned in status endpoints, but are not rendered in the main dashboard UI (app.js / index.html) nor included in WebSocket broadcast telemetry.
  • Missing Runtime Busiest-Day Evaluation (#33): Spec states: "Make the 40000 rule relative (e.g. > 3σ above the trailing 30-cycle mean)". The 3-sigma check is implemented only in the batch deduplication maintenance script in occupancy_repository.py; it was omitted from nightly runtime evaluation in evaluate_cycle_integrity_async.

Scope Creep (Unasked Behaviour)

  • Patrol Guard Removal from Formula Display (#34/#33): Formula display in index.html and calibration_desk.js was updated to 𝒪(t) = max(0, round(I - k·E)) (addressing PR #70/#32 rather than #34/#33).
  • Stepping Bounds Alteration: Multiplier bounds altered from [0.800, 1.300] to [0.900, 1.350].
  • Uncalibrated Badge in app.js: Render logic for the UNCALIBRATED (<14 samples) badge in app.js and configuration disclosure panel added outside issue scope.

Incorrect Implementations

  • Ingestion Gaps/Stalls Clearing Trust (#33): Spec states: "Only data-quality flags clear is_trusted and gate EWMA learning." In evaluate_cycle_integrity_async, is_trusted is cleared by (gaps or stalls): is_trusted = not any(flag.startswith("DATA_QUALITY_") for flag in flags) and not (gaps or stalls). Ingestion gaps are cycle completeness issues, yet they clear is_trusted and halt EWMA.
  • Incorrect Anomaly Flag on Negative Net Flow (#33): Spec states: "Separate flags into DATA_QUALITY_ (R2, R3, counter flush) and LOW_ACTIVITY_* (R1, R4)"*. In quarantine_and_deduplicate_calibration_logs_async, negative net flow (< -1000) assigns DATA_QUALITY_RATIO instead of DATA_QUALITY_NEGATIVE_NET or DATA_QUALITY_COUNTER_BURST.
  • Zero-Variance False-Positive in Spike Detection (#33): In quarantine_and_deduplicate_calibration_logs_async, if trailing history is identical (pstdev == 0), spike = total_in > mean + sigma * deviation and total_in > mean triggers for any total greater than mean by even a single count.

Summary: 6 standards findings (worst: repository performing rolling statistical domain calculations, and divergent CalibrationLogEntry definitions); 9 spec findings (worst: ingestion gaps clearing is_trusted contrary to spec, and zero-variance spike false-positive).

## Standards ### Documented Standard Violations (Hard) - **`app/db/occupancy_repository.py:1642` — Layer Separation Breach (Business Logic in Repository)**: - Standard: `docs/standards/code-standards.md` §1.1 & `AGENTS.md` §1 (*Layer Separation & Boundaries: Services vs Repositories*). - Rule: Services handle domain workflows and business calculations; repositories are strictly responsible for querying and persisting data into SQLite. - Breach: In `quarantine_and_deduplicate_calibration_logs_async`, the repository performs rolling statistical analysis (`statistics.mean`, `statistics.pstdev`), evaluates z-score spike thresholds, and assigns domain anomaly classifications. Domain calculations belong in `OccupancyManager`. - **`app/db/occupancy_repository.py` (`record_calibration_log_sync`/`_async`) — Untyped Dict Mutation**: - Standard: `docs/standards/code-standards.md` §2.3 (*Pydantic Schemas: Avoid passing untyped, raw dictionaries between service layers*). - Breach: When `entry: CalibrationLogEntry` is provided, it is dumped to an untyped dict via `entry.model_dump()`, mutated, and forwarded as raw `**kwargs` with `entry = None`. ### Baseline Smells (Judgement Calls) - **Feature Envy** (`app/db/occupancy_repository.py`): Repository contains statistical anomaly evaluation methods instead of delegating to domain services. - **Duplicated Code & Divergent Change** (`app/schemas/models.py` vs `app/schemas/occupancy_models.py`): `CalibrationLogEntry` is redundantly declared in two separate schema modules. Field `superseded_by` was added only to `models.py`, while `occupancy_repository.py` imports `CalibrationLogEntry` from `occupancy_models.py` where `superseded_by` is missing. - **Primitive Obsession** (`app/services/occupancy_service.py` & `app/db/occupancy_repository.py`): - Uses raw string prefix checks (`flag.startswith("DATA_QUALITY_")`) rather than Enum member comparisons. - In `occupancy_repository.py`, uses `"DATA_QUALITY_RATIO"` string literal instead of `CalibrationAnomalyFlag.DATA_QUALITY_RATIO_OUT_OF_BOUNDS`. - **Data Clumps** (`app/db/occupancy_repository.py`): Over 10 trust threshold parameters travel together through configuration lookups and query arguments instead of grouping into a typed trust settings model. ## Spec ### Missing or Partial Requirements - **Missing `is_current` Column (#34)**: Spec states: *"superseded rows kept with superseded_by and is_current columns"*. Only `superseded_by` was added; `is_current` was omitted from the SQLite schema (`database.py`), models, and queries. Queries rely solely on `superseded_by IS NULL`. - **Partial Data Trust & Cycle Completeness Publishing (#33)**: Spec states: *"Publish two figures, Data Trust and Cycle Completeness, instead of one tile."* Scores are stored in DB columns and returned in status endpoints, but are not rendered in the main dashboard UI (`app.js` / `index.html`) nor included in WebSocket broadcast telemetry. - **Missing Runtime Busiest-Day Evaluation (#33)**: Spec states: *"Make the 40000 rule relative (e.g. > 3σ above the trailing 30-cycle mean)"*. The 3-sigma check is implemented only in the batch deduplication maintenance script in `occupancy_repository.py`; it was omitted from nightly runtime evaluation in `evaluate_cycle_integrity_async`. ### Scope Creep (Unasked Behaviour) - **Patrol Guard Removal from Formula Display (#34/#33)**: Formula display in `index.html` and `calibration_desk.js` was updated to `𝒪(t) = max(0, round(I - k·E))` (addressing PR #70/#32 rather than #34/#33). - **Stepping Bounds Alteration**: Multiplier bounds altered from `[0.800, 1.300]` to `[0.900, 1.350]`. - **Uncalibrated Badge in `app.js`**: Render logic for the `UNCALIBRATED` (<14 samples) badge in `app.js` and configuration disclosure panel added outside issue scope. ### Incorrect Implementations - **Ingestion Gaps/Stalls Clearing Trust (#33)**: Spec states: *"Only data-quality flags clear is_trusted and gate EWMA learning."* In `evaluate_cycle_integrity_async`, `is_trusted` is cleared by `(gaps or stalls)`: `is_trusted = not any(flag.startswith("DATA_QUALITY_") for flag in flags) and not (gaps or stalls)`. Ingestion gaps are cycle completeness issues, yet they clear `is_trusted` and halt EWMA. - **Incorrect Anomaly Flag on Negative Net Flow (#33)**: Spec states: *"Separate flags into DATA_QUALITY_* (R2, R3, counter flush) and LOW_ACTIVITY_* (R1, R4)"*. In `quarantine_and_deduplicate_calibration_logs_async`, negative net flow (< -1000) assigns `DATA_QUALITY_RATIO` instead of `DATA_QUALITY_NEGATIVE_NET` or `DATA_QUALITY_COUNTER_BURST`. - **Zero-Variance False-Positive in Spike Detection (#33)**: In `quarantine_and_deduplicate_calibration_logs_async`, if trailing history is identical (`pstdev == 0`), `spike = total_in > mean + sigma * deviation and total_in > mean` triggers for any total greater than mean by even a single count. **Summary**: 6 standards findings (worst: repository performing rolling statistical domain calculations, and divergent `CalibrationLogEntry` definitions); 9 spec findings (worst: ingestion gaps clearing `is_trusted` contrary to spec, and zero-variance spike false-positive).
Author
Owner

Addressed the review in ec07481 and 2632878 (including fixes from the preceding PR branches).

  • Superseded audit rows remain stored and now expose is_current; trusted history uses only current rows. The write contract is distinct from the response contract.
  • Classification and statistical spike checks moved into the service. The zero-variance spike rule uses a configurable relative margin, and runtime ingress now receives the same historical anomaly check.
  • Ingestion gaps and stalls affect cycle completeness without clearing sensor trust; negative net movement during a gap is not mislabeled a burst. Data Trust and Cycle Completeness are separate fields in the live API, WebSocket payload, and dashboard.
  • The full suite and commit gate pass: 261 tests; the JavaScript suites passed as well.

An enum for the legacy flag strings and a larger configuration-object refactor remain reasonable follow-up cleanup, but are not needed for the audited behavior. No merge was performed.

Addressed the review in `ec07481` and `2632878` (including fixes from the preceding PR branches). - Superseded audit rows remain stored and now expose `is_current`; trusted history uses only current rows. The write contract is distinct from the response contract. - Classification and statistical spike checks moved into the service. The zero-variance spike rule uses a configurable relative margin, and runtime ingress now receives the same historical anomaly check. - Ingestion gaps and stalls affect cycle completeness without clearing sensor trust; negative net movement during a gap is not mislabeled a burst. Data Trust and Cycle Completeness are separate fields in the live API, WebSocket payload, and dashboard. - The full suite and commit gate pass: 261 tests; the JavaScript suites passed as well. An enum for the legacy flag strings and a larger configuration-object refactor remain reasonable follow-up cleanup, but are not needed for the audited behavior. No merge was performed.
Author
Owner

Review Follow-up & Verification of Previous Findings

Previous Review Status

  • Addressed:
    • Layer separation breach: Statistical spike checks (mean, pstdev, z-score) moved to OccupancyManager.is_counter_spike and service methods (docs/standards/code-standards.md §1.1).
    • Untyped dict mutation: Input validated via CalibrationLogWrite Pydantic model (docs/standards/code-standards.md §2.3).
    • Missing is_current: Column added to schema, migrations, models, and queries.
    • Ingestion gaps clearing trust: is_trusted is now cleared strictly by DATA_QUALITY_* flags; gaps and stalls reduce cycle completeness without clearing sensor trust.
    • Runtime spike check: is_counter_spike evaluated in evaluate_cycle_integrity_async.
    • Zero-variance spike detection: Guarded by trust_spike_zero_variance_margin.
    • UI display: Data Trust and Cycle Completeness rendered in index.html and app.js.
  • Acknowledged / Retained by Author:
    • Loose string flags and .startswith("DATA_QUALITY_") checks deferred to future enum cleanup.
    • 11 trust threshold parameters traveling together deferred to future settings refactor.

Items Missed by Previous Review

  • Historical Ingestion Gap Quarantined as Negative Net Anomaly in Batch Maintenance: In quarantine_calibration_anomalies_async (occupancy_service.py:993), if int(row["raw_net_flow"]) < int(cfg.get("trust_negative_net_limit", -1000)): flags.append("DATA_QUALITY_NEGATIVE_NET"). Unlike runtime evaluation in evaluate_cycle_integrity_async:938 (which includes and not gaps), this maintenance routine omits the gap check (FLAG_INGESTION_GAP not in flags). Consequently, historical cycles with negative net flow caused by ingestion gaps are retroactively flagged with DATA_QUALITY_NEGATIVE_NET and quarantined during dedup maintenance runs.
  • Untyped Dictionary Rows in Repo Seam: app/db/occupancy_repository.py:1831 get_current_automatic_calibration_rows_async() -> list[dict[str, Any]] returns raw dictionaries across the repository-to-service boundary.
  • Filter Inconsistency: app/db/occupancy_repository.py:1876 (get_calibration_log_for_cycle_async) filters on superseded_by IS NULL rather than is_current = 1.

Standards

Documented Standard Violations (Hard)

  • app/db/occupancy_repository.py:1831 & app/services/occupancy_service.py:983 — Untyped Dictionaries Across Layer Boundary:
    • Standard: docs/standards/code-standards.md §2.3 (Pydantic Schemas / Avoid passing untyped raw dictionaries between layers).
    • Breach: get_current_automatic_calibration_rows_async returns untyped list[dict[str, Any]], which OccupancyManager.quarantine_calibration_anomalies_async consumes via raw dictionary subscripting (row["trust_status"], row["total_in"], etc.) instead of utilizing CalibrationLogEntry or another typed schema.

Baseline Smells (Judgement Calls)

  • Primitive Obsession (app/services/occupancy_service.py:910-955, 960-964, 990-1002):
    Anomaly classifications rely on string literals and prefix checks (flag.startswith("DATA_QUALITY_"), flag.startswith("LOW_ACTIVITY_")) rather than Enum members (CalibrationAnomalyFlag), and are stored in SQLite as comma-separated strings (anomaly_flags).
  • Data Clumps (app/db/occupancy_repository.py:250-261, 350-363, app/schemas/occupancy_models.py:231-241):
    11 trust threshold configuration settings (trust_min_active_hours, trust_max_hourly_share, trust_ratio_min, etc.) travel together across SQL update statements, repository defaults, and service calculations instead of being grouped into a TrustSettings schema.
  • Duplicated Code (app/db/database.py:220-230, app/db/occupancy_repository.py:250-260, app/schemas/occupancy_models.py:231-241):
    Default values and boundaries for trust settings are duplicated across DDL statements, repository dicts, and Pydantic schemas.

Spec

Missing or Partial Requirements

  • None. Core requirements from Issues #34 and #33 are represented in the codebase.

Scope Creep (Unasked Behaviour)

  • Formula Display & Multiplier Bounds Alteration: Formula display update to 𝒪(t) = max(0, round(I - k·E)) and bounds alteration [0.900, 1.350] carried over from PR #70 scope.

Incorrect Implementations

  • Batch Maintenance Gap Net-Flow Quarantining:
    • Quote: Issue #33 / PR #71 AC: "Only data-quality flags clear is_trusted and gate EWMA learning".
    • Finding: quarantine_calibration_anomalies_async omits the FLAG_INGESTION_GAP check when evaluating raw_net_flow < -1000. Historical cycles with negative net flow caused by ingestion gaps are retroactively flagged with DATA_QUALITY_NEGATIVE_NET and quarantined during dedup maintenance runs, penalizing sensor trust for gateway outages.

Summary: 4 standards findings (worst: untyped dictionary rows returned across repository boundary in get_current_automatic_calibration_rows_async); 2 spec findings (worst: batch maintenance quarantine_calibration_anomalies_async retroactively quarantining cycles with ingestion gaps as DATA_QUALITY_NEGATIVE_NET).

### Review Follow-up & Verification of Previous Findings #### Previous Review Status - **Addressed**: - Layer separation breach: Statistical spike checks (`mean`, `pstdev`, z-score) moved to `OccupancyManager.is_counter_spike` and service methods (`docs/standards/code-standards.md §1.1`). - Untyped dict mutation: Input validated via `CalibrationLogWrite` Pydantic model (`docs/standards/code-standards.md §2.3`). - Missing `is_current`: Column added to schema, migrations, models, and queries. - Ingestion gaps clearing trust: `is_trusted` is now cleared strictly by `DATA_QUALITY_*` flags; gaps and stalls reduce cycle completeness without clearing sensor trust. - Runtime spike check: `is_counter_spike` evaluated in `evaluate_cycle_integrity_async`. - Zero-variance spike detection: Guarded by `trust_spike_zero_variance_margin`. - UI display: Data Trust and Cycle Completeness rendered in `index.html` and `app.js`. - **Acknowledged / Retained by Author**: - Loose string flags and `.startswith("DATA_QUALITY_")` checks deferred to future enum cleanup. - 11 trust threshold parameters traveling together deferred to future settings refactor. #### Items Missed by Previous Review - **Historical Ingestion Gap Quarantined as Negative Net Anomaly in Batch Maintenance**: In `quarantine_calibration_anomalies_async` (`occupancy_service.py:993`), `if int(row["raw_net_flow"]) < int(cfg.get("trust_negative_net_limit", -1000)): flags.append("DATA_QUALITY_NEGATIVE_NET")`. Unlike runtime evaluation in `evaluate_cycle_integrity_async:938` (which includes `and not gaps`), this maintenance routine omits the gap check (`FLAG_INGESTION_GAP not in flags`). Consequently, historical cycles with negative net flow caused by ingestion gaps are retroactively flagged with `DATA_QUALITY_NEGATIVE_NET` and quarantined during dedup maintenance runs. - **Untyped Dictionary Rows in Repo Seam**: `app/db/occupancy_repository.py:1831` `get_current_automatic_calibration_rows_async() -> list[dict[str, Any]]` returns raw dictionaries across the repository-to-service boundary. - **Filter Inconsistency**: `app/db/occupancy_repository.py:1876` (`get_calibration_log_for_cycle_async`) filters on `superseded_by IS NULL` rather than `is_current = 1`. --- ## Standards ### Documented Standard Violations (Hard) - **`app/db/occupancy_repository.py:1831` & `app/services/occupancy_service.py:983` — Untyped Dictionaries Across Layer Boundary**: - Standard: `docs/standards/code-standards.md` §2.3 (*Pydantic Schemas / Avoid passing untyped raw dictionaries between layers*). - Breach: `get_current_automatic_calibration_rows_async` returns untyped `list[dict[str, Any]]`, which `OccupancyManager.quarantine_calibration_anomalies_async` consumes via raw dictionary subscripting (`row["trust_status"]`, `row["total_in"]`, etc.) instead of utilizing `CalibrationLogEntry` or another typed schema. ### Baseline Smells (Judgement Calls) - **Primitive Obsession** (`app/services/occupancy_service.py:910-955, 960-964, 990-1002`): Anomaly classifications rely on string literals and prefix checks (`flag.startswith("DATA_QUALITY_")`, `flag.startswith("LOW_ACTIVITY_")`) rather than Enum members (`CalibrationAnomalyFlag`), and are stored in SQLite as comma-separated strings (`anomaly_flags`). - **Data Clumps** (`app/db/occupancy_repository.py:250-261, 350-363`, `app/schemas/occupancy_models.py:231-241`): 11 trust threshold configuration settings (`trust_min_active_hours`, `trust_max_hourly_share`, `trust_ratio_min`, etc.) travel together across SQL update statements, repository defaults, and service calculations instead of being grouped into a `TrustSettings` schema. - **Duplicated Code** (`app/db/database.py:220-230`, `app/db/occupancy_repository.py:250-260`, `app/schemas/occupancy_models.py:231-241`): Default values and boundaries for trust settings are duplicated across DDL statements, repository dicts, and Pydantic schemas. --- ## Spec ### Missing or Partial Requirements - None. Core requirements from Issues #34 and #33 are represented in the codebase. ### Scope Creep (Unasked Behaviour) - **Formula Display & Multiplier Bounds Alteration**: Formula display update to `𝒪(t) = max(0, round(I - k·E))` and bounds alteration `[0.900, 1.350]` carried over from PR #70 scope. ### Incorrect Implementations - **Batch Maintenance Gap Net-Flow Quarantining**: - Quote: Issue #33 / PR #71 AC: *"Only data-quality flags clear is_trusted and gate EWMA learning"*. - Finding: `quarantine_calibration_anomalies_async` omits the `FLAG_INGESTION_GAP` check when evaluating `raw_net_flow < -1000`. Historical cycles with negative net flow caused by ingestion gaps are retroactively flagged with `DATA_QUALITY_NEGATIVE_NET` and quarantined during dedup maintenance runs, penalizing sensor trust for gateway outages. --- **Summary**: 4 standards findings (worst: untyped dictionary rows returned across repository boundary in `get_current_automatic_calibration_rows_async`); 2 spec findings (worst: batch maintenance `quarantine_calibration_anomalies_async` retroactively quarantining cycles with ingestion gaps as `DATA_QUALITY_NEGATIVE_NET`).
chore(wip): open draft for passenger flow ingestion honesty (#26, #31, #30)
All checks were successful
CI / lint-and-test (pull_request) Successful in 1m15s
0b6cd98cee
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
chore(wip): open draft for business cycles in facility time (#27)
All checks were successful
CI / lint-and-test (pull_request) Successful in 1m16s
054612b972
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
chore(wip): open draft for calibration honesty: seeded k, open-window dwell (#36, #32)
All checks were successful
CI / lint-and-test (pull_request) Successful in 1m15s
0025f3f3bc
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
fix(occupancy): surface ingestion stalls in cycle integrity
All checks were successful
CI / lint-and-test (pull_request) Successful in 1m14s
7cfdcbcf60
fix(occupancy): resolve business cycles in facility time
All checks were successful
CI / lint-and-test (pull_request) Successful in 1m14s
54bc51edef
fix(calibration): seed site multiplier and measure open-window dwell
All checks were successful
CI / lint-and-test (pull_request) Successful in 1m17s
3221435faa
test: isolate background polling from API database tests
All checks were successful
CI / lint-and-test (pull_request) Successful in 1m12s
6f48205a16
test: isolate background polling from API database tests
All checks were successful
CI / lint-and-test (pull_request) Successful in 1m13s
1463e8c7f5
test: isolate background polling from API database tests
All checks were successful
CI / lint-and-test (pull_request) Successful in 1m17s
3da2ec0209
fix(occupancy): address PR #68 review round 2
All checks were successful
CI / lint-and-test (pull_request) Successful in 1m13s
b2b1d2eefc
- Spread reconstructed gaps uniformly (#31). _spread_gap_delta cut gaps
  only at hour edges, so an outage inside one clock hour (the 20-minute
  case from #31) still became a single spike at its midpoint. It now
  slices on a 5-minute grid that divides the hour, allocates shares by
  length with largest-remainder rounding (shares sum exactly), and
  never crosses an hourly bucket.
- Type the persisted counter readings. get_passenger_flow_readings_async
  returned untyped dicts and readings travelled as anonymous 7-tuples
  unpacked by index; both now use a PassengerFlowReading model
  (code-standards 2.3).
- Name the reconciliation tuning values (gap threshold, spread step,
  drift cadence, cap floor and multiplier, cold-start limit, stall
  polls) instead of inline literals.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
# Conflicts:
#	app/db/occupancy_repository.py
#	app/schemas/occupancy_models.py
#	app/services/occupancy_service.py
#	tests/test_counting_integrity.py
fix(occupancy): address PR #69 review round 2
All checks were successful
CI / lint-and-test (pull_request) Successful in 1m12s
30cb43bd57
- Quiet-window auto-calibration no longer runs before the reset it
  depends on. With the production window 03:30-04:30 around a 04:00
  reset, a run at 03:30 evaluated a cycle still in progress and set the
  per-cycle guard, so the proper post-reset run never happened. The
  daemon now waits until the calibrated cycle has ended when its window
  straddles the reset.
- Route the quiet-window checks through the facility-time seam (#27).
  The daemon and the analytics calibration status both did their own
  hour*60+minute arithmetic; they now use facility_window_containing()
  and seconds_until_window_start(), which handle midnight wrap.
- Tests pin the server clock to UTC with the facility at UTC-4: no run
  at 03:35, a run at 04:15 facility time, and no run at 04:15 server
  time (00:15 facility time).
- Cycle bounds are a CycleBounds named tuple (still unpackable) with a
  next_start property and a named CYCLE_END_EPSILON instead of bare
  +/-0.001 at call sites.
- Rename midnight_epoch to cycle_start_epoch: it holds the cycle reset,
  not midnight.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
# Conflicts:
#	CONTEXT.md
#	app/controllers/analytics_controller.py
#	app/db/occupancy_repository.py
#	app/facility_time.py
#	app/schemas/occupancy_models.py
#	app/services/analytics_service.py
#	app/services/occupancy_service.py
#	tests/test_counting_integrity.py
#	tests/test_facility_time.py
fix(occupancy): resolve opening hours for the business day, not the calendar day
All checks were successful
CI / lint-and-test (pull_request) Successful in 1m13s
04dddc099f
Between midnight and the reset (00:00-04:00 with a 04:00 reset) an
instant still belongs to the previous day's business cycle, but
get_active_schedule_info_async looked up the calendar day's schedule.
From midnight on it returned tomorrow's opening hours: live open-window
dwell read 0 for those hours (found while checking PR #70's review),
and a venue open past midnight was reported closed. The holiday lookup,
weekday and open/close epochs now use the date of the business cycle
that contains the instant.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
fix(calibration): address PR #70 review round 2
All checks were successful
CI / lint-and-test (pull_request) Successful in 1m16s
3c6daebfad
- Live dwell after closing: the review reported it dropping to 0 once
  the venue closes. The evening case already held (the schedule flag is
  per day), but between midnight and the reset the schedule resolved to
  the next calendar day, so dwell read 0 until 04:00. Fixed in #69
  (opening hours resolved for the business day) and merged here; a test
  pins that dwell after midnight equals dwell after closing.
- Remove the dead patrol_guard_count parameter from
  calculate_proportional_occupancy, calculate_confidence_interval,
  calculate_occupancy and the repository peak/average queries, along
  with the unused baseline_offset those queries only used to compute it.
  Since #32 the guard count plays no part in published occupancy.
- One DEFAULT_MULTIPLIER_VARIANCE constant (0.1125^2, the guardrail-
  derived default) replaces 18 copies of the literal across schemas,
  repository defaults, the database seed and services; the guardrail
  bounds are named OPERATIONAL_GUARDRAIL_MIN/MAX where the clamp uses
  them.
- Dayparts read ingress and average occupancy over the same half-open
  slice [start, end).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Brings PR #68, #69 and #70 round-2 review fixes into #71. Built by
applying the old->new #70 delta (3da2ec0..3c6daeb) with a 3-way apply;
the only conflicts were in occupancy_repository.py, where #71's trust
threshold columns and #70's DEFAULT_MULTIPLIER_VARIANCE both apply.
#71's own changes relative to #70 are line-for-line identical before
and after.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
fix(calibration): address PR #71 review round 2
All checks were successful
CI / lint-and-test (pull_request) Successful in 1m19s
dce8088fd8
- The historical maintenance pass quarantined gap cycles as sensor
  faults. quarantine_calibration_anomalies_async flagged
  DATA_QUALITY_NEGATIVE_NET without the "unless an ingestion gap
  explains it" condition that the runtime evaluation applies, so cycles
  whose net departures came from a gateway outage were retroactively
  excluded. Both paths now call one is_negative_net_anomaly() rule.
- get_current_automatic_calibration_rows_async returns typed
  AutomaticCalibrationSample models instead of raw dicts
  (code-standards 2.3); the maintenance pass reads attributes.
- get_calibration_log_for_cycle_async filters on is_current = 1, like
  every other reader, instead of superseded_by IS NULL.
- Trust flags use the CalibrationAnomalyFlag enum and its
  is_data_quality() / is_low_activity() classifiers instead of string
  prefix checks scattered through the service. Stored values unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Author
Owner

Review round 2 addressed — dce8088 (plus #68–#70 fixes merged in via 163549f)

Spec

  • Maintenance pass quarantining gap cycles (incorrect): fixed. quarantine_calibration_anomalies_async flagged DATA_QUALITY_NEGATIVE_NET without the ingestion-gap condition the runtime evaluation applies. Both now call one is_negative_net_anomaly() rule. Test: a gap cycle with net −1500 stays trusted while an identical cycle without a gap is quarantined; a mutation check confirms it.

Standards

  • Untyped calibration rows (hard, §2.3): fixed. get_current_automatic_calibration_rows_async returns AutomaticCalibrationSample models.
  • Filter inconsistency: fixed. get_calibration_log_for_cycle_async uses is_current = 1 (test added).
  • String flags / prefix checks: fixed. The existing CalibrationAnomalyFlag enum is now used, with is_data_quality() / is_low_activity() classifiers; stored values are unchanged.
  • Left for round 3: grouping the 11 trust thresholds into a TrustSettings schema, and deduplicating their defaults across DDL, repository and schema. That's the single-source-of-defaults work tracked by #74.

Scope-creep finding (formula display and bounds carried from #70): inherited from #70 through stacking, not new behaviour in this PR.

How #71 was brought up to date: applying the old→new #70 delta with a 3-way apply. The only conflicts were in occupancy_repository.py, where #71's trust-threshold columns and #70's variance constant both apply. #71's own changes relative to #70 are line-for-line identical before and after.

Verification: pytest 271 passed / 1 skipped, node --test 67/67, ruff clean, pre-commit passed.

Stacking: #68 → #69 → #70 → #71. Each branch now merges the one below it, so fixes propagate by merge instead of by cherry-pick. Merge in that order.

## Review round 2 addressed — `dce8088` (plus #68–#70 fixes merged in via `163549f`) **Spec** - **Maintenance pass quarantining gap cycles (incorrect): fixed.** `quarantine_calibration_anomalies_async` flagged `DATA_QUALITY_NEGATIVE_NET` without the ingestion-gap condition the runtime evaluation applies. Both now call one `is_negative_net_anomaly()` rule. Test: a gap cycle with net −1500 stays trusted while an identical cycle without a gap is quarantined; a mutation check confirms it. **Standards** - **Untyped calibration rows (hard, §2.3): fixed.** `get_current_automatic_calibration_rows_async` returns `AutomaticCalibrationSample` models. - **Filter inconsistency: fixed.** `get_calibration_log_for_cycle_async` uses `is_current = 1` (test added). - **String flags / prefix checks: fixed.** The existing `CalibrationAnomalyFlag` enum is now used, with `is_data_quality()` / `is_low_activity()` classifiers; stored values are unchanged. - **Left for round 3:** grouping the 11 trust thresholds into a `TrustSettings` schema, and deduplicating their defaults across DDL, repository and schema. That's the single-source-of-defaults work tracked by **#74**. **Scope-creep finding** (formula display and bounds carried from #70): inherited from #70 through stacking, not new behaviour in this PR. **How #71 was brought up to date:** applying the old→new #70 delta with a 3-way apply. The only conflicts were in `occupancy_repository.py`, where #71's trust-threshold columns and #70's variance constant both apply. #71's own changes relative to #70 are line-for-line identical before and after. **Verification:** pytest **271 passed / 1 skipped**, `node --test` **67/67**, ruff clean, pre-commit passed. **Stacking:** #68 → #69 → #70 → #71. Each branch now merges the one below it, so fixes propagate by merge instead of by cherry-pick. Merge in that order.
Author
Owner

Code review — round 3 (pre-merge)

Reviewed this PR's own increment in the stack (#68 → #69 → #70 → #71) with two independent passes: Standards (docs/standards/code-standards.md, AGENTS.md, CONTEXT-FORMAT.md / ADR-FORMAT.md, a code-smell baseline) and Spec (the PR description and its implementation decisions, all previous rounds, and the originating issues). The Spec pass re-ran the suites in a throwaway worktree.

Policy for this round: P1 is fixed before merge; P2 is fixed or explicitly accepted; P3 goes to follow-up issues (inline production values → #74, language-specific text → #73).

Standards

Round-2 fixes hold. The migrations are idempotent and lose no rows on a DB built from master's schema.

  • P2: the ratio rule is duplicated, and the copies already disagree. Runtime uses total_vol > 0 and maintenance uses total_in > 0 and total_out > 0, so a cycle with 0 exits is flagged by one and not the other. This is the same drift round 2 fixed for negative net.
  • P3: score_cycle_flags still uses string flags; AutomaticCalibrationSample uses str where the trust enums exist; the cycle-date derivation is copied between sync and async record_calibration_log; three ALTER loops with two shapes; multiplier_provenance defaults differ between CREATE (COMPUTED) and ALTER (LEGACY_UNKNOWN), with no comment; the new panels are English text with no data-i18n (→ #73); the threshold inputs carry a 4th copy of the defaults (→ #74).

Spec

pytest 271, node 67. #34's hard requirement is met: the only calibration-log deletion is in retention pruning, and MANUAL_OVERRIDE survives.

  • P2, reproduced: the upgrade discards production's learned k. Every pre-upgrade row becomes LEGACY_UNKNOWN and is excluded from trusted history, so on a migrated DB the multiplier refreshed to 1.0. The learned 1.116 is lost, along with a legacy MANUAL_OVERRIDE (1.15). Rows whose k differs from their raw I/E can't have been backfilled, yet are dropped too.
  • P2, reproduced: the live Data Trust tile shows 50/100 every morning. It runs whole-cycle rules on the cycle still in progress (at 11:00, 2 active hours with 800 in / 200 out trips both burst and ratio), implying a sensor fault that doesn't exist, which is exactly what #33 set out to stop. It also costs 4 extra queries per status call.
  • P2: gap cycles feed k. They stay trusted, but their I/E ratio is biased when an outage dropped exits. The ratio rule also still charges gaps to Data Trust while the negative-net rule exempts them.
  • P3: the is_current lookup test passes even with the filter removed; missing cross-field config validation (ratio_min > max is accepted, and so is spike_min_samples > window); existing DBs keep computed_exit_multiplier DEFAULT 1.0.
  • Correction to the round-2 reply: the [0.900, 1.350] bounds and the formula text in the UI are added by this PR, not inherited from #70. On #70's branch the admin calibration desk still clamps manual k to [0.80, 1.30]. master ends up correct only because #70 and #71 merge back to back.

Standards: 7 findings, worst P2 (the duplicated ratio rule). Spec: 7 findings, worst P2 (the upgrade discards production's learned k).

Resolution (maintainer, 2026-09-24): all P2s are being fixed in this PR. Gap cycles will be excluded from k learning without being counted as sensor faults, consistent with #31. The ratio rule becomes one shared function with the same gap exemption as negative net. Legacy rows are kept unless they are provably backfilled. The live tile reports completed cycles only.

## Code review — round 3 (pre-merge) Reviewed this PR's **own increment** in the stack (#68 → #69 → #70 → #71) with two independent passes: **Standards** (`docs/standards/code-standards.md`, `AGENTS.md`, `CONTEXT-FORMAT.md` / `ADR-FORMAT.md`, a code-smell baseline) and **Spec** (the PR description and its implementation decisions, all previous rounds, and the originating issues). The Spec pass re-ran the suites in a throwaway worktree. Policy for this round: **P1** is fixed before merge; **P2** is fixed or explicitly accepted; **P3** goes to follow-up issues (inline production values → #74, language-specific text → #73). ## Standards Round-2 fixes hold. The migrations are idempotent and lose no rows on a DB built from `master`'s schema. - **P2: the ratio rule is duplicated, and the copies already disagree.** Runtime uses `total_vol > 0` and maintenance uses `total_in > 0 and total_out > 0`, so a cycle with 0 exits is flagged by one and not the other. This is the same drift round 2 fixed for negative net. - **P3:** `score_cycle_flags` still uses string flags; `AutomaticCalibrationSample` uses `str` where the trust enums exist; the cycle-date derivation is copied between sync and async `record_calibration_log`; three ALTER loops with two shapes; `multiplier_provenance` defaults differ between CREATE (`COMPUTED`) and ALTER (`LEGACY_UNKNOWN`), with no comment; the new panels are English text with no `data-i18n` (→ #73); the threshold inputs carry a 4th copy of the defaults (→ #74). ## Spec pytest 271, node 67. #34's hard requirement is met: the only calibration-log deletion is in retention pruning, and `MANUAL_OVERRIDE` survives. - **P2, reproduced: the upgrade discards production's learned `k`.** Every pre-upgrade row becomes `LEGACY_UNKNOWN` and is excluded from trusted history, so on a migrated DB the multiplier refreshed to **1.0**. The learned 1.116 is lost, along with a legacy `MANUAL_OVERRIDE` (1.15). Rows whose `k` differs from their raw I/E can't have been backfilled, yet are dropped too. - **P2, reproduced: the live Data Trust tile shows 50/100 every morning.** It runs whole-cycle rules on the cycle still in progress (at 11:00, 2 active hours with 800 in / 200 out trips both burst and ratio), implying a sensor fault that doesn't exist, which is exactly what #33 set out to stop. It also costs 4 extra queries per status call. - **P2: gap cycles feed `k`.** They stay trusted, but their I/E ratio is biased when an outage dropped exits. The ratio rule also still charges gaps to Data Trust while the negative-net rule exempts them. - **P3:** the `is_current` lookup test passes even with the filter removed; missing cross-field config validation (`ratio_min > max` is accepted, and so is `spike_min_samples > window`); existing DBs keep `computed_exit_multiplier DEFAULT 1.0`. - **Correction to the round-2 reply:** the `[0.900, 1.350]` bounds and the formula text in the UI are **added by this PR**, not inherited from #70. On #70's branch the admin calibration desk still clamps manual `k` to `[0.80, 1.30]`. `master` ends up correct only because #70 and #71 merge back to back. --- **Standards: 7 findings, worst P2** (the duplicated ratio rule). **Spec: 7 findings, worst P2** (the upgrade discards production's learned `k`). **Resolution (maintainer, 2026-09-24):** all P2s are being fixed in this PR. Gap cycles will be **excluded from `k` learning without being counted as sensor faults**, consistent with #31. The ratio rule becomes one shared function with the same gap exemption as negative net. Legacy rows are kept unless they are provably backfilled. The live tile reports completed cycles only.
fix(occupancy): address PR #68 review round 3
All checks were successful
CI / lint-and-test (pull_request) Successful in 1m20s
2e1641bffc
- P1: a reconstructed gap crossing the cycle rollover was published twice.
  After a rollover the expected cycle total included the slices stamped
  before the reset, so the next drift reconcile republished them at
  "now" (untagged) just after 04:00, the phantom spike #26 targets. The
  expected total now counts only the slices that land in the current
  cycle.
- P2: drift above the cap is no longer carried forward in cap-sized
  pieces. As #26 specifies, it is recorded as a reset anomaly and
  re-seeded, never published; drift within the cap is published once.
- P2: one gap anywhere no longer disables R2 for the whole cycle. The
  hourly trust input excludes reconstructed slices, so a gap cannot fake
  a burst, while the share is still taken over all passages, so it
  cannot hide a real one either.
- The hour-edge spread test asserted a condition that is always true; it
  now pins the exact slices.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
docs(occupancy): document the quiet-window placement constraint (PR #69 round 3)
All checks were successful
CI / lint-and-test (pull_request) Successful in 1m20s
884aed8e4b
Round 3 found that calibration_cycle_bounds picks the cycle by calendar
day, so a quiet window far from the reset calibrates a stale cycle and a
window wrapping midnight reads the cycle in progress. Production's
03:30-04:30 window around a 04:00 reset is unaffected and the behaviour
predates this PR; the maintainer accepted it for follow-up. Record the
constraint where it is configured and where it is implemented.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
fix(calibration): address PR #70 review round 3
All checks were successful
CI / lint-and-test (pull_request) Successful in 1m20s
5841296daf
- P2: upgrading an existing database copied the learned
  active_exit_multiplier into the new initial_exit_multiplier seed, so
  the EWMA was re-seeded with its own output. Existing installations now
  get LEGACY_IMPLICIT_SEED_MULTIPLIER (1.1162), the constant every
  function defaulted to before the seed was configurable (#36); fresh
  databases keep 1.0. Test upgrades a pre-seed database twice.
- P2 (accepted): ADR 0006 claimed the published result has a zero floor.
  The proportional model and its aggregates do, but the live two-phase
  presentation still adds a staff baseline during opening hours; the ADR
  now says so and points to follow-up.
- initial_exit_multiplier bounds use OPERATIONAL_GUARDRAIL_MIN/MAX.
- The dashboard maturity badge branches on the backend's
  sample_maturity.is_uncalibrated instead of re-deriving the threshold,
  which left two branches unreachable.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Brings the #68-#70 round-3 fixes into #71. The one conflict was the R2
burst rule: kept #71's DATA_QUALITY_* vocabulary and configurable share,
and took #68's removal of the blanket 'and not gaps' exemption. Tests
that asserted the pre-#71 flag name FLAG_BURST_COUNTER_FLUSH (and would
have become vacuous) now assert DATA_QUALITY_BURST_COUNTER_FLUSH.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
fix(calibration): address PR #71 review round 3
All checks were successful
CI / lint-and-test (pull_request) Successful in 1m21s
357fb63058
- P2: upgrading discarded production's learned k. Every pre-provenance
  calibration log became LEGACY_UNKNOWN and was excluded from trusted
  history, so the multiplier fell back to its seed (operator
  MANUAL_OVERRIDEs lost too). Legacy rows are now classified at startup:
  a multiplier equal to the raw I/E ratio (the old backfill's signature)
  is LEGACY_BACKFILLED and excluded; everything else is LEGACY_MEASURED
  and kept. Idempotent; also repairs databases already migrated by
  earlier builds of this branch.
- P2: the live Data Trust tile showed 50/100 every morning because it ran
  whole-cycle rules on the cycle in progress. The live tile and the
  analytics calibration status now report the last completed cycle's
  stored scores (with its date), or nothing before the first calibration.
  Also drops four queries per status call.
- P2 (maintainer decision): cycles with an ingestion gap stay trusted but
  no longer teach k, since their I/E is biased by what the outage dropped
  (consistent with #31).
- P2: the ratio rule was written twice and the copies disagreed (a cycle
  with zero exits was flagged at runtime only). One is_ratio_anomaly()
  rule serves both, with the same ingestion-gap exemption as negative
  net.
- The is_current lookup test could not fail; the superseded row now has
  the newest timestamp.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Author
Owner

Review round 3 addressed — 357fb63 (plus #68–#70 round-3 fixes merged in via 0bc4583)

  • P2, the upgrade discarding learned k: fixed. Pre-provenance rows are classified at startup instead of dropped: a multiplier equal to the raw I/E ratio (the old backfill's signature) becomes LEGACY_BACKFILLED and is excluded; everything else, operator overrides included, is LEGACY_MEASURED and kept. Idempotent, and it repairs databases migrated by earlier builds of this branch. Test covers backfilled, measured and manual legacy rows.
  • P2, live Data Trust 50/100 every morning: fixed. The live tile and the analytics calibration status report the last completed cycle's stored scores (with trust_scores_cycle_date), or nothing (— in the UI) before the first calibration. The cycle in progress is never scored for display, which also drops 4 queries per status call.
  • P2, gap cycles feeding k: fixed (maintainer decision). They stay trusted (not a sensor fault) but are excluded from k learning, since their I/E is biased by what the outage dropped (consistent with #31).
  • P2, the duplicated ratio rule: fixed. One is_ratio_anomaly() rule for runtime and maintenance, with the same gap exemption as negative net. A zero-exit cycle is now flagged by both.
  • The vacuous is_current test: now fails without the filter.
  • Merge note: the R2 conflict with #68 kept this PR's DATA_QUALITY_* vocabulary and configurable share, and took #68's removal of the blanket gap exemption. Three tests that asserted the pre-#71 flag name, and would have become vacuous, now assert DATA_QUALITY_BURST_COUNTER_FLUSH.
  • A mutation check: the 5 new behaviour tests all fail on the pre-fix code.
  • P3s → #78; the threshold-input defaults → #74; the untranslated panels → #73.

Verification: pytest 281 passed / 1 skipped, node 67/67, CI green on 357fb63.

## Review round 3 addressed — `357fb63` (plus #68–#70 round-3 fixes merged in via `0bc4583`) - **P2, the upgrade discarding learned `k`: fixed.** Pre-provenance rows are classified at startup instead of dropped: a multiplier equal to the raw I/E ratio (the old backfill's signature) becomes `LEGACY_BACKFILLED` and is excluded; everything else, operator overrides included, is `LEGACY_MEASURED` and kept. Idempotent, and it repairs databases migrated by earlier builds of this branch. Test covers backfilled, measured and manual legacy rows. - **P2, live Data Trust 50/100 every morning: fixed.** The live tile and the analytics calibration status report the **last completed cycle's** stored scores (with `trust_scores_cycle_date`), or nothing (`—` in the UI) before the first calibration. The cycle in progress is never scored for display, which also drops 4 queries per status call. - **P2, gap cycles feeding `k`: fixed (maintainer decision).** They stay trusted (not a sensor fault) but are excluded from `k` learning, since their I/E is biased by what the outage dropped (consistent with #31). - **P2, the duplicated ratio rule: fixed.** One `is_ratio_anomaly()` rule for runtime and maintenance, with the same gap exemption as negative net. A zero-exit cycle is now flagged by both. - **The vacuous `is_current` test:** now fails without the filter. - **Merge note:** the R2 conflict with #68 kept this PR's `DATA_QUALITY_*` vocabulary and configurable share, and took #68's removal of the blanket gap exemption. Three tests that asserted the pre-#71 flag name, and would have become vacuous, now assert `DATA_QUALITY_BURST_COUNTER_FLUSH`. - A mutation check: the 5 new behaviour tests all fail on the pre-fix code. - **P3s → #78**; the threshold-input defaults → #74; the untranslated panels → #73. **Verification:** pytest **281 passed / 1 skipped**, node 67/67, CI green on `357fb63`.
gabogg merged commit dc80674df5 into master 2026-09-24 22:33:11 +00:00
gabogg deleted branch fix/calibration-audit-trail 2026-09-24 22:33:11 +00:00
Sign in to join this conversation.
No description provided.