AI Observability & Monitoring
Real-time monitoring of AI model performance, automated alert routing, searchable event logs, and end-to-end data lineage across all AI-powered features of The Reading Dr. platform.
Simulate a school-wide morning login storm with concurrent login attempts from students and staff. Measures login throughput (logins/sec), latency (average and p95), success rate, error rate, and resource utilization. Identifies bottlenecks in authentication, directory, and backend services and produces a report with findings and actionable recommendations.
Load Scenario Presets
Throughput
847
logins/sec
Avg / P95 Latency
118ms
P95: 312ms
Success Rate
99.2%
992 succeeded
Error Rate
0.8%
8 failed
Resource Utilization
67%
Target: <80%
54%
Target: <85%
71%
Target: <80%
58%
Target: <80%
23 reqs
Target: <50 reqs
14 reqs
Target: <30 reqs
Bottleneck Analysis
P95 latency elevated during peak ramp-up phase
→ Scale auth service replicas during 8-9am window
Connection pool utilization approaching threshold
→ Increase pool size from 20 to 30 connections
Recommendations
- Authentication service handles peak load within SLO targets
- Consider pre-warming auth service replicas 15 minutes before school opening
- Directory lookup queue remains well below threshold — no action needed
- Backend services maintain 99.2% success rate under peak concurrent load
4
Scenarios Run
4
Passed
0
Failed
847
Peak logins/sec
Findings & Recommendations
- All scenarios passed SLO targets (P95 < 1000ms, success rate > 95%)
- Peak throughput of 2,105 logins/sec achieved under stress conditions
- Authentication bottleneck identified at 3000 users — scale horizontally for >2500 concurrent
- Directory service and backend DB remain within utilization targets across all scenarios
- Recommendation: Implement connection pooling auto-scaling for auth service during peak hours