Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

that’s been wrong for a while, but affirmed with gpt-4o-mini beating out Sonnet 3.5. OpenAI fine tuned 4o and 4o-mini to provide answers that meaningfully improve model congeniality but trivially improve model intelligence.

Chatbot Arena ELO is a dead metric.



Wow, I overlooked GPT-4o-mini that far up.

But if you change the category to Math (or something else hard), mini drops way down and Claude 3.5 Sonnet goes to the top.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: