Sky Lab's FreeToken runs large LLMs on consumer GPUs locally
Original titleReally cool results from Sky Lab’s @Andy_ShuoYang to run massive LLMs on local GPUs!
AISummary
Sky Lab's FreeToken runs official checkpoints of large models such as Qwen3.6 35B on an 8GB RTX 4060 laptop at 39 tokens per second. The same approach reportedly serves DeepSeek-V4-Flash 284B at 22-25 tokens per second on an RTX 5090 desktop, and GLM-5.2 753B at 15 tokens per second on an RTX PRO 6000 workstation.
Source: Matei Zaharia · x.comPublished · added here