Show HN: AI SRE Arena, an Open Benchmark for AI SRE Agents on Kubernetes
Sentiment Mix
Geography
Expert Signals
Hacker News
source • 2 mentions
felix089
author • 1 mention
emrahsamdan
author • 1 mention
AI-Generated Claims
Generated from linked receipts; click sources for full context.
Openai just launched their decisions endpoint, cloudflare launched clef the other week, and many more jev alternatives are out there.We wanted to put the popular ones to the test and thought Pac-Man is a good benchmark for simple and fast decision making.So we let jev 1.13, kev, clef, clef flash, GPT-6 Luna and Laya play Pac-Man against bot ghosts.The low latency of these models allows for real time play.
Supported by 1 story
We had each model play 100 games, published a leader board and open-sourced the repo so anyone can run their own model and join the ranking.
Supported by 1 story
Link to repo: https://github.com/opper-ai/jevman-benchmark/blob/main/CONTR...You can also join the game and play as Pac-Man yourself, and the ghosts are the models, either a mix of models or all jev, kev, clef etc.
Supported by 1 story
A game costs about 2 cent, all models are running via my startup opper, and we added free credits for everyone to try.It's pretty fun to play and surprisingly difficult to beat jev's highscore.
Supported by 1 story
Any feedback is more than welcome!
Supported by 1 story
Related Events
Show HN: Edi Life OS – self-hosted life dashboard with an MCP server for AI
Uncategorized • 10/9/2026
Anthropic, Google and Mistral Unveil Their Latest AI Models: What’s New and Why It Matters - CNET
LLMs • 10/9/2026
OpenAI, Anthropic, and Meta AI Breaches Shared the Same Testing Vendor - eSecurity Planet
LLMs • 10/9/2026
OpenAI cannot make AI safe on its own [pdf]
LLMs • 10/9/2026
FTC Probing OpenAI and Anthropic Over Product Safety Concerns - Bloomberg.com
LLMs • 10/9/2026