Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I don't think the problem is that memory footprint can't be reduced. The problem is that the token generation rate is simply far too low.

OP suggests that the token rate of of their solution is ~5 per second. That's at least an order of magnitude slower than commercially available models.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: