STORY · VERKTOY_
Run Kimi K3 with 29 GB RAM at 0.50 tokens per second
WASTE is a C-based inference engine that enables running Kimi K3 – a model with 2.78 trillion parameters – on a 64 GB MacBook Pro at 0.49–0.54 tokens per second. The engine keeps the model trunk in memory, streams experts directly from disk on demand, and uses remaining RAM as a limited expert cache.
WHY IT MATTERS
This demonstrates that frontier-scale models can run locally without network dependency or cloud costs, opening up applications where data security and data control are critical. It shows that the boundary of what's possible on consumer hardware has shifted significantly.
SOURCES
MACHINE-GENERATED SUMMARY This summary is written by machine from the sources below. We sort and explain — but we are a way into the field, not the final word. Check the source when something matters to you.