SREGym: Can AI agents resolve production issues? Real-world SRE problems including metastable failures, misconfigurations, and many more. Live system environments. From the University of Illinois at Urbana-Champaign. To submit, open an issue with the submission label at github.com/SREGym/SREGym.
SREGym-Lite results
Top results on curated 21-fault cohort.
1 | Codex | GPT-5.6 Sol (max) | 95.2 | 85.7 | 81.0 | 211.0 | 397.0 | 1.42M |
2 | Claude Code | Claude Opus 5 | 92.1 | 82.5 | 76.2 | 210.1 | 419.8 | 1.64M |
3 | Codex | GPT-5.6 Terra (max) | 85.7 | 79.4 | 69.8 | 214.7 | 415.8 | 1.68M |
4 | Codex | GPT-5.6 Luna (max) | 87.3 | 79.4 | 68.3 | 284.0 | 492.9 | 2.74M |
5 | Codex | GPT-5.6 Sol (medium) | 77.8 | 71.4 | 58.7 | 108.2 | 270.6 | 0.77M |
Diag. Diagnosis success rate · Mit. Mitigation success rate · E2E End-to-end (both diagnosis and mitigation correct) · TTD Time-to-diagnose (seconds) · TTM Time-to-mitigate (seconds) · Tokens Mean token usage per run