SREGym: Can AI agents resolve production issues? Real-world SRE problems including metastable failures, misconfigurations, and many more. Live system environments. From the University of Illinois at Urbana-Champaign. To submit, open an issue with the submission label at github.com/SREGym/SREGym.
Interactive terminal available on desktop
SREGym-Lite-1004 results
Top results on the curated 17-fault cohort.
1 | Codex | GPT-6 Astra (max) | 96.1 | 100.0 | 96.1 | 173.1 | 287.6 | 719K |
2 | Codex | GPT-6 Astra (medium) | 100.0 | 90.2 | 90.2 | 92.0 | 138.4 | 377K |
3 | Codex | GPT-6.1 Sol (max) | 98.0 | 90.2 | 90.2 | 217.4 | 371.6 | 807K |
4 | Codex | GPT-6.1 Sol (medium) | 98.0 | 90.2 | 90.2 | 103.8 | 156.6 | 462K |
5 | Codex | GPT-5.6 Sol (max) | 94.1 | 82.4 | 76.5 | 225.1 | 409.7 | 1.54M |
Diag. Diagnosis success rate · Mit. Mitigation success rate · E2E End-to-end (both diagnosis and mitigation correct) · TTD Time-to-diagnose (seconds) · TTM Time-to-mitigate (seconds) · Tokens Mean token usage per run