Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

If you want to not quantize at all, you need to double it for fp16—16GB.


Yes, but I think it's standard to do inference at q8, not fp16.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: