hikctl update apply reports success while a stray process owns APP_PORT #180

Open
opened 2026-09-28 13:40:27 +00:00 by gabogg · 0 comments
Owner

Problem

hikctl update apply can report success even though the service it restarted never took over the app port.

Seen on prod (Windows Server 2019):

  • 2026-09-18 13:14: someone started a uvicorn process outside SCM. Its parent process later exited, but the process kept serving 0.0.0.0:8888.
  • 2026-09-21: update apply ran 3fe5c8a → 1bb097d and wrote UPDATE_APPLIED ... "success": true to logs/lifecycle.log.
  • From then until 2026-09-28, users were still on the pre-9/18 code. The HikCentralGateway service was restart-looping about every 23 s. Each attempt ran the full lifespan startup against the live DB (webhook re-subscription, anomaly scan, calibration baseline), then died with [Errno 10048] error while attempting to bind on address ('0.0.0.0', 8888). That grew HikCentralGateway.err.log by about 12 MB a day.

The /health probe in app/cli/commands/cmd_update.py (around line 239) passed because the stray process answered it, not the new service.

Expected

  • Before stopping the service, update apply (and probably service start / doctor) finds out which process listens on APP_PORT. If that process doesn't belong to the managed service, it refuses and explains.
  • After starting the service, verification confirms that the listener PID is the service's own child and that /health answers from it. A health response alone isn't enough.

Notes

  • Fixed on prod by hand on 2026-09-28 (stray process killed, service now owns :8888).
  • Open design point: whether hikctl should offer to kill a stray listener or only report it.
## Problem `hikctl update apply` can report success even though the service it restarted never took over the app port. Seen on prod (Windows Server 2019): - 2026-09-18 13:14: someone started a uvicorn process outside SCM. Its parent process later exited, but the process kept serving `0.0.0.0:8888`. - 2026-09-21: `update apply` ran `3fe5c8a → 1bb097d` and wrote `UPDATE_APPLIED ... "success": true` to `logs/lifecycle.log`. - From then until 2026-09-28, users were still on the pre-9/18 code. The `HikCentralGateway` service was restart-looping about every 23 s. Each attempt ran the full lifespan startup against the live DB (webhook re-subscription, anomaly scan, calibration baseline), then died with `[Errno 10048] error while attempting to bind on address ('0.0.0.0', 8888)`. That grew `HikCentralGateway.err.log` by about 12 MB a day. The `/health` probe in `app/cli/commands/cmd_update.py` (around line 239) passed because the stray process answered it, not the new service. ## Expected - Before stopping the service, `update apply` (and probably `service start` / `doctor`) finds out which process listens on `APP_PORT`. If that process doesn't belong to the managed service, it refuses and explains. - After starting the service, verification confirms that the listener PID is the service's own child and that `/health` answers from it. A health response alone isn't enough. ## Notes - Fixed on prod by hand on 2026-09-28 (stray process killed, service now owns :8888). - Open design point: whether `hikctl` should offer to kill a stray listener or only report it.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
gabogg/hikcentral#180
No description provided.