Observed Signal · Sep 17, 2026 · Product Launch · Source: PR Newswire: Advertising & Marketing · Impact: 1/5 · Sentiment: Positive
PrismML Launches Bonsai 2 27B, Its Most Capable Model Yet
PrismML announced the launch of Bonsai 2 27B, its new flagship ternary AI model based on Qwen3.8 27B. The model is compressed to 5.9 GB, more than 9x smaller than its full-precision counterpart, while retaining over 98% of the benchmark performance. It is designed for local deployment, supporting reasoning, coding, vision, and agentic tasks. The model achieves an aggregate score of 83.9 across 20 benchmarks, retaining 98.2% of Qwen3.8 27B's performance. It is available for free under the Apache 2.0 license starting September 17, 2026. PrismML aims to deliver high intelligence density, enabling powerful AI on consumer-grade hardware.
Relevant to AI industry but not directly AdTech or advertising; a niche model launch with limited direct impact on AdTech.
Track NVIDIA Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- PrismML launched Bonsai 2 27B, a ternary model based on Qwen3.8 27B.
- The model is 5.9 GB, offering a 9x memory footprint reduction compared to full-precision.
- Ternary Bonsai 2 27B retains 98.2% of the performance of Qwen3.8 27B.
- The model is available for free under the Apache 2.0 license.
- Bonsai 2 27B is optimized for local deployment on consumer GPUs and edge devices.
Connected Companies & Entities
5 Entities mapped“The model reaches up to 143 tokens/second on NVIDIA GeForce RTX 5090....”
“PrismML was founded with support from Khosla Ventures....”
“PrismML was founded with support from Google....”
“PrismML has continuing support from Samsung....”
Ontology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
PrismML launches tiny Bonsai 2 27B LLM for on-device AI
AI startup PrismML has released Bonsai 2 27B, a compressed large language model that fits on PCs and potentially high-end smartphones. The model compresses Alibaba's Qwen3.8 27B model to 5.9 GB, a 9x to 10x reduction in memory, while retaining 98% of the original's benchmark performance. Founded by Caltech researchers and led by CEO Babak Hassibi, PrismML has raised a $22.25 million seed round from Khosla Ventures, Cerberus Capital, and Caltech. The company's ternary weight compression technique simplifies model weights to just +1, -1, or 0. PrismML plans to apply this technology to larger models in the coming months. The startup is reportedly in talks with Apple, though this has not been confirmed.
Prism ML Releases Bonsai 27B, 1-bit On-Device LLM
Prism ML announced Bonsai 27B, a 27B-class large language model using true 1-bit binary transformer weights (effective 1.125 bits per weight) that reduces deployed footprint to ~3.9 GB, enabling on-device inference on high-end smartphones and laptops. Derived from a Qwen3.6-27B backbone with hybrid attention and optional vision tower, Bonsai 27B supports a 262K-token context via predominantly linear attention and 4-bit KV-cache quantization. The release includes custom 1-bit kernels for Apple MLX and CUDA and a DSpark speculative-decoding drafter that yields a 1.37x decode speedup on NVIDIA H100. Benchmarks show Bonsai retains 89.5% of FP16 reasoning performance (76.11 thinking average) while dramatically lowering memory and energy use. The model is distributed under the Apache 2.0 license.
Alibaba Releases Qwen3.8-27B Open-Weight Model
Alibaba's Qwen team released Qwen3.8-27B, a 27-billion-parameter, Apache 2.0‑licensed, vision-capable model with a 262,144‑token context window and weights that compress to about 17–18 GB at 4-bit quantization. The release (Aug 14, 2026) enables frontier-like coding and agent capabilities to run locally on consumer hardware (e.g., a single 24 GB GPU or mid-range Apple Silicon). Independent benchmarking from Artificial Analysis scores the model 52 on its Intelligence Index; vendor-reported Terminal-Bench 2.1 results also show a substantial step up from Qwen3.6-27B. The article is a technical guide focused on runtime settings, quantization, hardware tiers, and deployment steps for local inference.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
