Production Reliability Dashboard

Generated 2026-03-28 20:14 from Slack #_alerts_prod and AWS SNS alert emails for 2026-03-23 07:00 to 2026-03-28 07:00.

How to use Pick an alert family, service/resource, or AWS alarm. Inside an alert family, click a target to scope trends, notes, discussion signal, and latest matching alerts. Open Global Evidence Explorer only when you need report-wide pivots.
Slack messages66alert and ops posts in channel history
Slack discussions3threads with human follow-up we could mine for signal
AWS alert emails0Inbox/Trash/Spam matches from SNS sender
Latest observed event2026-03-28 05:12most recent alert timestamp seen in either source

Top Slack Alert Families

Top Impacted Services / Resources

AWS Email Alarm Families

No data.

Investigation

Investigation

Choose one item above to start a scoped investigation, then narrow to a target when an alert family exposes multiple workloads.
Global Evidence Explorer

Global Evidence Explorer

Report-wide charts and tables stay here, separate from the active investigation scope.

Slack Alerts by Day

AWS Alert Emails by Day

Evidence

Slack Alert Families

AlertSeverityCountLast SeenStatusThreadsTop Impacted ServicesDiscussion SignalLatest Thread Note
KubeJobFailedwarning322026-03-28 05:12Seen today1grafana (26)download-album-29576545 (6)General investigation
production.ERROR: Allowed memory size of 134217728 bytes exhausted (tried to allocate 200704 bytes) {"exception":"[object] (Symfony\\Compon… | Ionut Ciolan a dat out of memory. E setat cumva in php.ini max memory la 128mb? Ai cum sa o cresti sa nu mai crape? | mai bine ar fi sa faci un pic refactor la cod. Pare ca inc…
TraefikServiceHighLatencywarning162026-03-27 17:21Recent (72h)1ai-api-svc-3900 (11)web-80 (5)subscriptions-api-svc-3400 (4)rooms-api-svc-3700 (1)General investigation
era un query la catalog care mergea greu
NodeSystemSaturationwarning62026-03-27 11:29Recent (72h)0grafana (6)None
KubeHpaMaxedOutwarning52026-03-27 13:05Recent (72h)0docgen2-api (5)None
NodeCPUHighUsagewarning32026-03-27 11:29Recent (72h)0grafana (3)None
KubeNodeEvictionwarning22026-03-26 08:38Recent (72h)0grafana (2)None
TraefikServiceHighErrorRatecritical22026-03-27 08:23Recent (72h)1uni-api-svc-4000 (1)web-80 (1)General investigation
Andrei Alexandru pare ca e o buba cu foreign keys pe tuiasi | Service method DisciplineServiceImpl.patchDisciplineByIdAndAcademicPlan failed after 12ms: could not execute statement [(conn=4451547) Cann… | ok, urmaresc si rezolv, am mai gasit ieri.o astfel de problema, o urmaresc si pe asta

Status is heuristic. Slack rarely posts explicit resolutions, so “Seen today” or “Recent” means the alert family still appeared in production recently, not that it is definitely unresolved.

Slack Impacted Service / Resource View

This view attributes alerts to the workload or resource named in the alert text. Grafana, Loki, and Tempo are treated as observability components and are excluded when a more specific impacted target is also present.

Impacted Service / ResourceCountLast SeenStatusTop Alert TypesDiscussion SignalLatest Thread Note
grafana372026-03-28 04:14Seen todayKubeJobFailed (26)NodeSystemSaturation (6)NodeCPUHighUsage (3)KubeNodeEviction (2)General investigation
production.ERROR: Allowed memory size of 134217728 bytes exhausted (tried to allocate 200704 bytes) {"exception":"[object] (Symfony\\Compon… | Ionut Ciolan a dat out of memory. E setat cumva in php.ini max memory la 128mb? Ai cum sa o cresti sa nu mai crape? | mai bine ar fi sa faci un pic refactor la cod. Pare ca inc…
ai-api-svc-3900112026-03-27 17:21Recent (72h)TraefikServiceHighLatency (11)None
download-album-2957654562026-03-28 05:12Seen todayKubeJobFailed (6)None
web-8062026-03-27 13:19Recent (72h)TraefikServiceHighLatency (5)TraefikServiceHighErrorRate (1)General investigation
era un query la catalog care mergea greu
docgen2-api52026-03-27 13:05Recent (72h)KubeHpaMaxedOut (5)None
subscriptions-api-svc-340042026-03-27 08:35Recent (72h)TraefikServiceHighLatency (4)None
rooms-api-svc-370012026-03-27 08:35Recent (72h)TraefikServiceHighLatency (1)None
uni-api-svc-400012026-03-25 10:58Seen this weekTraefikServiceHighErrorRate (1)General investigation
Andrei Alexandru pare ca e o buba cu foreign keys pe tuiasi | Service method DisciplineServiceImpl.patchDisciplineByIdAndAcademicPlan failed after 12ms: could not execute statement [(conn=4451547) Cann… | ok, urmaresc si rezolv, am mai gasit ieri.o astfel de problema, o urmaresc si pe asta

AWS Email Alarm Families

AWS AlarmEmailsALARMOKState FlipsFirst SeenLast SeenLatest StateStatus

“Flapping, latest OK” means the most recent email was an OK, but the alarm toggled repeatedly and is still a reliability concern.

Global Discussion-Derived Signal

Thread DateAlertSeverityServicesSignalKey Notes
2026-03-26 11:19TraefikServiceHighLatencywarningweb-80General investigation
era un query la catalog care mergea greu
2026-03-25 10:58TraefikServiceHighErrorRatecriticaluni-api-svc-4000General investigation
Andrei Alexandru pare ca e o buba cu foreign keys pe tuiasi | Service method DisciplineServiceImpl.patchDisciplineByIdAndAcademicPlan failed after 12ms: could not execute statement [(conn=4451547) Cann… | ok, urmaresc si rezolv, am mai gasit ieri.o astfel de problema, o urmaresc si pe asta
2026-03-24 10:51KubeJobFailedwarninggrafanaGeneral investigation
production.ERROR: Allowed memory size of 134217728 bytes exhausted (tried to allocate 200704 bytes) {"exception":"[object] (Symfony\\Compon… | Ionut Ciolan a dat out of memory. E setat cumva in php.ini max memory la 128mb? Ai cum sa o cresti sa nu mai crape? | mai bine ar fi sa faci un pic refactor la cod. Pare ca inc…