SREGym: Can AI agents resolve production issues? Real-world SRE problems including metastable failures, misconfigurations, and many more. Live system environments. From the University of Illinois at Urbana-Champaign. To submit, open an issue with the submission label at github.com/SREGym/SREGym.

SREGym-Lite-1004 results

Top results on the curated 17-fault cohort.

1
Codex
OpenAI
GPT-6 Astra (max)
96.1100.096.1173.1287.6719K
2
Codex
OpenAI
GPT-6 Astra (medium)
100.090.290.292.0138.4377K
3
Codex
OpenAI
GPT-6.1 Sol (max)
98.090.290.2217.4371.6807K
4
Codex
OpenAI
GPT-6.1 Sol (medium)
98.090.290.2103.8156.6462K
5
Codex
OpenAI
GPT-5.6 Sol (max)
94.182.476.5225.1409.71.54M

Diag. Diagnosis success rate · Mit. Mitigation success rate · E2E End-to-end (both diagnosis and mitigation correct) · TTD Time-to-diagnose (seconds) · TTM Time-to-mitigate (seconds) · Tokens Mean token usage per run

view all results ↗