Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

wow, not bad result on the computer use benchmark for the mini model. for example, Claude Sonnet 4.6 shows 72.5%, almost on par with GPT-5.4 mini (72.1%). but sonnet costs 4x more on input and 3x more on output


what's the point of this benchmark if sonnet is working great at my tasks and mini can't solve my tasks?




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: