Qwen3.8-27B, up to 144 tok/sec on M5 Max.
Let that sink in. ⚡️✨🚀
Excited to partner with @inco_ai to bring you their Splash inference engine in LM Studio on day 0!
Get it running now: https://lmstudio.ai/blog/splash-engine
LM Studio announced that Qwen3.8-27B runs at up to 144 tokens per second on an M5 Max MacBook Pro through its partnership with Inco Splash. The post claims up to 3× the decode speed of Ollama, 2× oMLX, and almost 4× when an agent fans out into sub-agents.
Inco Splash is described as an open-source inference engine built for the model and Apple silicon, available through the linked LM Studio blog.
Qwen3.8-27B, up to 144 tok/sec on M5 Max.
Let that sink in. ⚡️✨🚀
Excited to partner with @inco_ai to bring you their Splash inference engine in LM Studio on day 0!
Get it running now: https://lmstudio.ai/blog/splash-engine
Qwen3.8-27B at 144 tok/s on an M5 Max MacBook Pro ⚡ Meet Inco Splash: our open-source inference engine, built around the model and around Apple silicon. Up to 3× the decode speed of Ollama, 2× oMLX, and almost 4× when an agent fans out into sub-agents.View quoted post on X
Source: LM Studio · x.comPublished · added here