LLM performance is twice as fast as RTX 5090
https://creativestrategies.com/mac-studio-m3-ultra-ai-workst...
your tests are wrong. you used MLX for Mac Studio (optimized for Apple Silicon) but you didn't use vLLM for 5090. There's no way a machine with half the bandwidth of 5090 delivers twice as fast tok/s.
also, the GP was mostly testing models that fit in both 5090 and Mac Studio.
That's 3x the RAM of the 5090.
And a bit less than half the bandwidth (saying for completeness).
LLM performance is twice as fast as RTX 5090
https://creativestrategies.com/mac-studio-m3-ultra-ai-workst...