Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

DeepSWE “feels” like the right benchmark in comparison to Artificial Analysis indices and other coding benchmarks. And by their metrics, GPT-5.5 is still king in token efficiency, speed, and overall intelligence per dollar.

https://deepswe.datacurve.ai/

Fable 5 is cool and all, but we have not yet seen GPT-5.6.



GLM5.2 isn't even on this benchmark


True. Z.AI ran that bench themselves and report 46.2, which is lower than GPT-5.5 and Opus 4.8, but crushing the other open weights models.

https://z.ai/blog/glm-5.2




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: