Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

DeepSWE has Gemini 3.8 Flash up really high, too.


It does and that one also feels kind of saturated for measuring the most advanced models - opus, Gemini, astra, sol, fable, glm, kimi all scoring within statistical noise of each other. (74 +-3% down to 69% +-5% for kimi).

It's still providing strong discrimination between weaker models.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: