STORY · VERKTOY_

Open source engine runs Gemma 4 26B with just 2 GB RAM on M-series Mac

A developer has built TurboFieldfare, an inference engine written in Swift and Metal that makes it possible to run a 4-bit-quantized Gemma 4 26B model on Mac with limited RAM. The engine uses a hybrid approach where it keeps the core and KV cache in RAM, while experts are streamed from SSD as needed, achieving 5–6 tokens/second on M2 and 31–35 tokens/second on M5.

WHY IT MATTERS

This opens up the possibility for powerful on-device AI on standard consumer Macs without cloud dependency, with implications for privacy, latency, and accessibility of advanced models locally.

SOURCES

MACHINE-GENERATED SUMMARY This summary is written by machine from the sources below. We sort and explain — but we are a way into the field, not the final word. Check the source when something matters to you.