Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

M3 Ultra has a big GPU with 819 GB/sec bandwidth.

LLM performance is twice as fast as RTX 5090

https://creativestrategies.com/mac-studio-m3-ultra-ai-workst...



> LLM performance is twice as fast as RTX 5090

your tests are wrong. you used MLX for Mac Studio (optimized for Apple Silicon) but you didn't use vLLM for 5090. There's no way a machine with half the bandwidth of 5090 delivers twice as fast tok/s.


Unless it’s a large model that doesn’t fit in the 5090, bust that’s no longer a $4k macstudio I think.


that's orthogonal to the speed discussion.

also, the GP was mostly testing models that fit in both 5090 and Mac Studio.


$4k will get you a 96 GB Mac Studio with M3 Ultra (819 GB/sec).

That's 3x the RAM of the 5090.


> That's 3x the RAM of the 5090

And a bit less than half the bandwidth (saying for completeness).


Yeah that's probably wrong. But the M3 Ultra is good enough for local inferencing, in any case.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: