Перейти к содержанию

BPMN — Realtime Service Incident

Обработка падения realtime-service (NFR-007).

flowchart TD
    Start([Alert WS errors spike]) --> Detect{Health check fail?}
    Detect -->|Yes| Restart[Restart realtime pods]
    Detect -->|No| Investigate[Check Redis NATS]
    Restart --> Recover{WS connections restore?}
    Recover -->|Yes| Notify[Notify ops resolved]
    Recover -->|No| Failover[Failover to standby]
    Investigate --> Fix[Fix root cause]
    Fix --> Restart
    Failover --> Notify
    Notify --> End([End])

RTO target: < 5 min (NFR-007)

Исходник: bpmn-incident.mmd