Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

This benchmark draws a very different picture having GPT5.5 on the very top with 70% and DeepSeek at 8%

https://deepswe.datacurve.ai



DeepSWE has been heavily criticized though. https://github.com/datacurve-ai/deep-swe/issues/21 Putting GPT 5.5 on top is the obviously correct part, but everything else about it makes very little sense.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: