Observed Signal · Jul 14, 2026 · Partnership · Source: CNBC Technology · Impact: 4/5 · Sentiment: Positive
Apple in talks with PrismML on on‑device AI
Apple is in early talks with PrismML, a Caltech spinout backed by Khosla Ventures, after the startup publicly released compressed versions of Alibaba’s open-source Qwen model that it says shrink the model from roughly 54 GB to under 4 GB. PrismML’s technique — reducing internal values to one or three possible states — aims to let a 27-billion-parameter model run on iPhone 15 or newer devices, improving speed, energy use and memory footprint at the cost of a small drop in some performance metrics. PrismML released two compressed model variants for free, has Caltech-licensed patents, and raised a $16.25 million seed round. Analysts noted the potential impact on Siri, device battery use, and global chip demand, while cautioning that real-world testing at scale will determine whether the efficiency claims hold up.
If validated and adopted, PrismML’s compression could enable large-model inference on consumer devices, affecting Siri capability, device battery/chip demand, cloud costs and privacy strategies for a major platform (Apple).
Track Apple Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- Apple is in early talks with PrismML to evaluate its compressed AI models for on‑device use.
- PrismML says it compressed Alibaba’s open-source Qwen model from about 54 GB to under 4 GB, enabling 27B parameters to run on iPhone 15 or newer.
- PrismML is a Caltech spinout; Caltech owns the underlying patents and licenses them exclusively to PrismML.
- PrismML raised a $16.25 million seed round in March backed by Khosla Ventures and other investors and publicly released two compressed model variants for free.
- PrismML claims compressed models use 10–15x less memory, respond 6–8x faster, and consume 3–6x less energy versus conventional versions, with a small trade-off in some performance areas.
Connected Companies & Entities
11 Entities mapped“Apple is in talks with a small Silicon Valley company that says it can shrink powerful artificial intelligence models enough to run directly...”
“PrismML, a Khosla Ventures-backed spinout from the California Institute of Technology, publicly released compressed versions of Alibaba’s op...”
“PrismML publicly released compressed versions of Alibaba’s open-source Qwen model on Tuesday....”
“Apple is trying to make Siri more competitive with assistants from OpenAI and Anthropic while keeping more personal information and AI proce...”
“Apple is trying to make Siri more competitive with assistants from OpenAI and Anthropic while keeping more personal information and AI proce...”
“Hassibi said Google’s open-source Gemma model is next in the pipeline, followed by much larger models, including those from frontier labs th...”
“PrismML is releasing two compressed versions of the model for free. They are designed to run on everyday devices, including iPhones, MacBook...”
“Micron shares plunged in March after Google published its TurboQuant paper on cutting memory use without hurting model performance, though t...”
“Phil Solis, who leads IDC’s research on client processors, said power consumption may be the biggest open question....”
“Tarun Pathak, research director at Counterpoint Research, said the model’s performance on lengthy prompts, battery consumption during multit...”
“Morgan Stanley estimates Apple’s average dynamic random access memory cost per bit could rise roughly 190% year over year in fiscal 2027......”
Ontology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
PrismML launches tiny Bonsai 2 27B LLM for on-device AI
AI startup PrismML has released Bonsai 2 27B, a compressed large language model that fits on PCs and potentially high-end smartphones. The model compresses Alibaba's Qwen3.8 27B model to 5.9 GB, a 9x to 10x reduction in memory, while retaining 98% of the original's benchmark performance. Founded by Caltech researchers and led by CEO Babak Hassibi, PrismML has raised a $22.25 million seed round from Khosla Ventures, Cerberus Capital, and Caltech. The company's ternary weight compression technique simplifies model weights to just +1, -1, or 0. PrismML plans to apply this technology to larger models in the coming months. The startup is reportedly in talks with Apple, though this has not been confirmed.
Alibaba’s Qwen AI to Integrate with Apple Intelligence
China’s Cyberspace Administration approved Apple Intelligence for the Chinese market after Apple reached deals to integrate Chinese AI capabilities. Alibaba confirmed its Qwen model will be integrated into Apple Intelligence across iOS, iPadOS, macOS and visionOS in China but gave no timeline. PrismML released compressed versions of open-source Qwen, reportedly shrinking it from roughly 54 GB to under 4 GB and enabling a 27‑billion‑parameter variant to run on iPhone 15 and newer devices. Baidu said it is working with Apple on Apple Intelligence features for iPhones in China, and reports say Baidu’s Kunlunxin AI‑chip unit is targeting a Hong Kong IPO that could value it near $50 billion. The announcements coincided with notable share gains for Alibaba and Baidu and follow earlier delays while Apple evaluated Chinese AI providers.
On‑Device AI Becomes Practical
Apple and Google delivered on-device foundation models in 2026 that materially change the economics, privacy and availability of LLM features. Apple's third‑generation Foundation Models (AFM 3) announced at WWDC (June 8) ships an on‑device model that stores ~20 billion parameters in flash while activating ~1–4 billion per request via a sparse activation strategy. Google’s Gemma 4 family (released April 2) uses per‑layer embeddings and mixture‑of‑experts designs for small active footprints (edge variants like E4B run with ~4.5B effective parameters; a 26B MoE only activates a fraction of experts per token). Hardware NPUs in phones and edge devices now run 4–8B class models at usable speeds; Google says Gemma 4 edge models run offline on devices including Raspberry Pi and NVIDIA Jetson Orin Nano. Apple also opened its Foundation Models framework to third‑party and open models and added agent primitives and on‑device semantic search, shifting many features from cloud‑billed inference to free, private local execution.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
