Skip to content
Read the original: Matei Zaharia· Published 38/100AI score38/100

Sky Lab's FreeToken runs large LLMs on consumer GPUs locally

Original titleReally cool results from Sky Lab’s @Andy_ShuoYang to run massive LLMs on local GPUs!

AISummary

Sky Lab's FreeToken runs official checkpoints of large models such as Qwen3.6 35B on an 8GB RTX 4060 laptop at 39 tokens per second. The same approach reportedly serves DeepSeek-V4-Flash 284B at 22-25 tokens per second on an RTX 5090 desktop, and GLM-5.2 753B at 15 tokens per second on an RTX PRO 6000 workstation.

Read the original x.com

Source: Matei Zaharia · x.comPublished · added here