STORY · MODELLER_
Maple-Preview – 20B MoE model runs at 120 tokens/second on iPhone
Maple-Preview is a ternary 20B MoE model that has been optimized to run at 120 tokens per second directly on iPhone. The project is being showcased on Hacker News as a demonstration of efficient model compression for mobile devices.
WHY IT MATTERS
It shows that large language models can now run locally on smartphones without internet connection, making AI more accessible and private for end users.
SOURCES
MACHINE-GENERATED SUMMARY This summary is written by machine from the sources below. We sort and explain — but we are a way into the field, not the final word. Check the source when something matters to you.