🏢 Big Tech / /via The Verge / updated Jul 14, 2026

Apple Unveils On-Device 120B Parameter Model for iOS 20

Apple introduced a 120 billion parameter on-device model on July 9 2026 for iOS 20. The model runs entirely on A20 silicon with 4-bit quantization and delivers 78.3 percent on MMLU. It powers new Siri features launching in September.

#Apple
~/ Big Tech/ Apple Unveils On-Device 120B Parameter Model fo...

Apple revealed a 120 billion parameter language model on July 9 2026 that runs fully on-device. The model uses 4-bit quantization and fits within 48 GB of unified memory on A20 chips. It scores 78.3 percent on MMLU while consuming under 6 watts during inference.

The model will ship with iOS 20 in September 2026. It enables offline Siri conversation, real-time photo editing suggestions, and on-device code completion in Xcode. Apple claims latency under 180 milliseconds for typical queries.

Development began in late 2024 after Apple acquired two AI startups focused on efficient inference. The model was trained on 9 trillion tokens with heavy emphasis on private user data that never leaves the device. Apple published a technical paper detailing its new sparse attention mechanism.

Previous on-device models in iOS 18 topped out at 3 billion parameters. The 40x scale increase required new hardware accelerators in the A20 neural engine. Apple will offer an optional cloud fallback for complex tasks using Private Cloud Compute.

Why this matters

Apple's move pressures Android vendors to match on-device performance without cloud dependency. Privacy-focused regulation gains a concrete example of viable local AI. Developers gain new APIs for building offline-first intelligent apps.

Apple will release a 30 billion parameter version for iPhone 15 and older devices in October 2026. The company plans to open parts of the model architecture to academic researchers by year end.

share
𝕏 FB
← cd ../news