動作
Bug #1456
已結束[Backend Bug] A17 CP001 重連遇 Boot cooldown 後四支狀態殘留 UNAVAILABLE
開始日期:
2026-09-01
完成日期:
預估工時:
概述
問題摘要¶
2026-08-29 檢查 CT1(寶台 A17)時,CP001 底下四個 connectors 在 Branch/使用者畫面全部顯示 UNAVAILABLE,無法開始充電;但 CP edge 四支皆為 AVAILABLE、Modbus meter read 持續更新,CP001 本身也持續 ONLINE。
本次不是四支硬體同時故障,而是 CP001 共用的 OCPP WebSocket 反覆斷線重連後,Branch connector projection 長期殘留 OCPP heartbeat stale。
與 #1436 的關係¶
- #1436 已修正 Branch watchdog:
lastHeartbeat與lastBootNotification任一仍在 330 秒 cutoff 內時,不得投影 OFFLINE/UNAVAILABLE。 - A17 已部署含此判斷的 Branch JAR;本次不是舊版未部署。
- #1436 的需求紀錄明確排除「Charge Point 重連後立即補送 Heartbeat」,若 reconnect 沒送出 fresh BootNotification,既有規則不提供 grace。
- 本次正是該排除路徑:第二次 reconnect 落在 BootNotification cooldown 內,Boot 被跳過;第一筆 post-reconnect Heartbeat 又晚於 watchdog cutoff。
- 因此另開本單追蹤 CP reconnect liveness 與 ONLINE 後 connector convergence,不回頭擴張已完成的 #1436 範圍。
環境與影響¶
- 案場:CT1(寶台 A17)
- Charge Point:
CP001 - Connectors:
CP001-001、CP001-002、CP001-003、CP001-007 - 實際發生:2026-08-27 06:10:54(Asia/Taipei)
- 發現及暫時復原:2026-08-29
- 使用者影響:同一 CP 的四支 connector 全部顯示無法使用,無法開始充電
- Branch watchdog:Heartbeat interval 300 秒、offline timeout 330 秒
四支會同時受影響,是因為它們共用 CP001 的單一 OCPP session;watchdog 判定 charge point 離線後,會一次投影所有沒有 active transaction 保護的 connectors。最終 watchdog 執行時四支都沒有 active transaction,因此同秒被改寫。
已確認的事件時序¶
- 2026-08-27 06:04:17:CP001 WebSocket 因 pong timeout 關閉。
- 06:04:47:重連成功,BootNotification Accepted;Branch 記錄 fresh Boot。
- 06:06:47:WebSocket 再次因 pong timeout 關閉。
- 06:07:17:再次重連,但距前次 Boot 未滿 300 秒,BootNotification 因 cooldown 被跳過。
- 06:07:37~06:08:35:CP001 各 connector 重新送出實際狀態,四支最後都曾回到
AVAILABLE。 - 06:10:53:watchdog 使用 cutoff
06:05:23.505;當時lastHeartbeat=05:39:48、lastBootNotification=06:04:47,兩者都已過期。 - 06:10:54:四支 connectors 同時被改成
UNAVAILABLE / OCPP heartbeat stale。 - 06:11:29:Heartbeat 成功並把 CP001 恢復為
ONLINE,但沒有恢復 connector operational status。 - 至 2026-08-29 21:18:Branch 四支仍是 stale
UNAVAILABLE;edge 四支皆為AVAILABLE且 meter read 新鮮。
根因判定¶
已確認¶
-
BootNotificationScheduler使用 300 秒 cooldown;快速重連時可跳過 Boot。 -
HeartbeatScheduler已啟動時,reconnect 不會重建排程或立即補送 Heartbeat,只等待原本的 fixed-delay 時點。 - 因此可能出現「Boot 被 cooldown 跳過 + 下一筆 Heartbeat 晚於 Branch 330 秒 cutoff」。
- Branch 的下一筆 Heartbeat 只恢復 charge point connection,不會重新同步已被 offline propagation 覆寫的 connector 狀態。
- edge 若一直維持
AVAILABLE、沒有新狀態變化,就不會再自然補送 StatusNotification,造成 Branch/edge 長期分裂。
尚待驗證的上游觸發因素¶
- 2026-08-27 只有 CP001 出現 pong timeout,共 68 次;CP002/CP003 都是 0。
- CP001 同期有兩筆長時間 transaction、大量 MeterValues 與 970 次
WebSocket not connectedqueue retry。 - 這些證據高度指向 CP001 OCPP session/message queue 壓力,但 pong 未回覆的最底層成因尚未直接證明;實作前需補 instrumentation 或可重現測試,不能直接當作定論。
已執行的暫時復原¶
2026-08-29 21:22 復原前確認:
- Branch linked ACTIVE transaction:0。
- Edge linked ACTIVE transaction:0。
- Branch/edge 四支
current_transaction_id全為 null。
只重啟 ems-cp-api(CP001),沒有重啟 Branch、CP002、CP003,也沒有直接修改 DB:
- 21:22:49:BootNotification Accepted、Heartbeat #1 成功。
- 21:22:59~21:23:04:四支初始
StatusNotification(AVAILABLE/NOERROR)全部成功。 - 21:23:Branch 與 edge 四支均確認為
AVAILABLE,CP001 為ONLINE。 - 沒有中斷任何充電交易。
此操作只恢復現場,沒有永久修補。
下週待確認的修正方向¶
- CP 每次成功 reconnect 時,即使 BootNotification 因 cooldown 略過,也應立即送出一筆 Heartbeat 或等價 liveness;需避免 heartbeat storm 與重複 scheduler。
- Branch 在 OFFLINE → ONLINE 後應有 connector convergence 機制,使狀態能與 edge 實況重新同步;不得以 Heartbeat 任意偽造
AVAILABLE或覆寫FAULTED/CHARGING。 - 調查 CP001 pong timeout 與 MeterValues queue/pending request 壓力,評估 backpressure、批次上傳、pending request timeout/cleanup 與 WebSocket thread isolation。
- 保留既有 active transaction 保護、TCP half-open 偵測與 #1436 Boot grace 行為。
以上只是候選方向,須下週依 Redmine SOP 完成需求確認與規劃 gate 後才能實作。
驗收條件¶
- 舊 Heartbeat + reconnect 落在 Boot cooldown 內時,必須在 watchdog cutoff 前建立 fresh liveness。
- 重複 WebSocket close/reconnect 不產生多個 Heartbeat scheduler,也不造成 Boot/Heartbeat storm。
- reconnect 後已上報
AVAILABLE的 connectors 不得再被舊 liveness 覆寫成 staleUNAVAILABLE。 - 即使 watchdog 先投影離線,下一筆 Heartbeat 使 CP 回到 ONLINE 後,Branch connectors 最終仍須與 edge 當前狀態收斂。
-
FAULTED、CHARGING、active transaction 與 current transaction 保護不得退化。 - CP002/CP003 與第三方 OCPP Charge Point 行為不得受 CloudLink-specific scheduler 修改影響。
- 補可控制時間的 reconnect/Boot cooldown/watchdog/first Heartbeat 回歸測試與 production-equivalent E2E。
E2E Impact¶
- 分類:Add
- 案例:
OCP-018 - Priority:P0
- 狀態:Gap
- Trigger:OCPP reconnect、Boot cooldown、Heartbeat scheduler、offline watchdog 或 ONLINE recovery 變更
- Release pack:Major / Charging / OCPP
- 相關既有案例:
OCP-001、OCP-005、OCP-007、OCP-008、OCP-011
目前狀態¶
- 現場已暫時恢復。
- 本單保持 New。
- 2026-08-29 只完成診斷、復原與建單,不修改程式、不建立實作 branch、不部署。
- 預計下週有空時再依 Redmine SOP 進入需求確認、規劃、實作與測試。
動作