Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

At this point I barely put any value in any of the benchmarks. I just use the models for coding (and related things like software product design/planning/ideation/etc.) tasks and judge them subjectively, and also see how others judge them subjectively on HN and Twitter.


I use benchmarks…

…that are my own private internal suite on my own code bases where I can judge the output properly

I also measure wall clock time to completion which has been a surprising separator in practice.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: