Apple in Talks With PrismML to Squeeze Massive AI Models Onto iPhone
PrismML, a Khosla Ventures-backed spinout from the California Institute of Technology, publicly released compressed versions of Alibaba's open-source Qwen model on Tuesday. The release is sending shockwaves through the AI industry. The company reduced the model from roughly 54 GB to less than 4 GB, allowing all 27 billion of its parameters to run on an iPhone 15 or newer. PrismML CEO Babak Hassibi told CNBC that Apple and other companies have been evaluating the startup's models and measuring their speed, energy efficiency, and performance on devices. PrismML shrinks AI models by drastically simplifying how their internal information is stored — reducing each value from 16 bits to just one or three possible values — significantly cutting the memory required. The compressed models use between 10 and 15 times less memory, generate responses six to eight times faster, and consume three to six times less energy than conventional versions. Larger models running directly on iPhones would allow for more Apple Intelligence features to run on device instead of on Apple's Private Cloud Compute servers, which could reduce Apple's costs and further enhance user privacy. Analysts, however, urge caution. PrismML's claims still need to be proven outside controlled demonstrations, with performance on lengthy prompts, battery consumption, and reliability at scale all critical factors.
Why Inbenta

