Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

DeepSWE seems to strongly, strongly prefer ChatGPT models. There were also major flaws in its methodology pointed out recently, that overlap strongly with the flaws OpenAI pointed out in its SWE Verified report.

I use both ChatGPT and Claude for engineering work on a daily basis, touching performance critical code to application backends to frontend work, and I've found that DeepSWE scores don't reflect my reality when I assess high quality output from the models/harnesses.

Not that Opus always beats GPT 5.5., but that 5.5 is ahead of Opus on a general benchmark smells off to me.



Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: