hikctl update apply reports success while a stray process owns APP_PORT #180
Labels
No labels
blocked
bug
enhancement
high-priority
low-priority
needs-info
needs-triage
ready-for-agent
ready-for-human
referenced
research
wontfix
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
gabogg/hikcentral#180
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Problem
hikctl update applycan report success even though the service it restarted never took over the app port.Seen on prod (Windows Server 2019):
0.0.0.0:8888.update applyran3fe5c8a → 1bb097dand wroteUPDATE_APPLIED ... "success": truetologs/lifecycle.log.HikCentralGatewayservice was restart-looping about every 23 s. Each attempt ran the full lifespan startup against the live DB (webhook re-subscription, anomaly scan, calibration baseline), then died with[Errno 10048] error while attempting to bind on address ('0.0.0.0', 8888). That grewHikCentralGateway.err.logby about 12 MB a day.The
/healthprobe inapp/cli/commands/cmd_update.py(around line 239) passed because the stray process answered it, not the new service.Expected
update apply(and probablyservice start/doctor) finds out which process listens onAPP_PORT. If that process doesn't belong to the managed service, it refuses and explains./healthanswers from it. A health response alone isn't enough.Notes
hikctlshould offer to kill a stray listener or only report it.