VMTech
Discuss a project

PrismML Releases 5.9GB Bonsai 2 Model Based on Qwen3.8 27B

PrismML Releases 5.9GB Bonsai 2 Model Based on Qwen3.8 27B

PrismML has released Bonsai 2 27B, a compressed version of Alibaba’s open-source Qwen3.8 27B model that occupies 5.9GB. The company says this represents a nine- to tenfold reduction in memory requirements while preserving 98% of Qwen’s aggregate benchmark scores, a size that could allow the model to run on PCs and potentially high-end smartphones.

The release is the latest from the Caltech-founded AI startup, which has raised a $22.25 million seed round. PrismML is led by Babak Hassibi, a Caltech professor specialising in compression technologies, and counts Databricks co-founder Ion Stoica among its advisers. Its backers include Khosla Ventures, Cerberus Capital and Caltech.

Compression through ternary weights

PrismML’s approach focuses on the weights in a language model, the values that store what the model learned during training. In conventional representations, each weight normally requires 16 bits. PrismML instead uses what it calls ternary weights, limiting each value to +1, −1 or 0.

Reducing the amount of information stored per weight cuts the model’s footprint substantially. PrismML argues that this can make capable reasoning models practical on devices that would not accommodate their uncompressed versions. The company says Bonsai 2’s 98% aggregate benchmark result improves on the first Bonsai release, which matched 95% of the original model’s benchmark scores.

Local execution is the central use case

The first Bonsai model, released in March, has been downloaded more than 11 million times, PrismML says. Its smaller models have attracted another 2.6 million downloads. The figures indicate interest in model formats intended to work outside large cloud infrastructure, although benchmark parity does not by itself determine performance in a specific deployment.

Hassibi has said that some loss from compression is likely, while also noting that benchmark differences may not directly translate into meaningful changes in real-world use. He also pointed to the software harness around a model as an important part of operational accuracy.

Next target: models with hundreds of billions of parameters

PrismML plans to apply its compression technique to models in the several-hundred-billion-parameter range in the coming months. Hassibi expects it may be easier to retain intelligence in larger models because there is more room for compression without sacrificing capability.

For organisations, Bonsai 2 highlights a practical evaluation path: assess whether a compressed local model meets the required task quality, device-memory limits and privacy needs before assuming that every reasoning workload must be sent to the cloud. On-device execution can keep prompts on equipment the user already owns, while the actual suitability still depends on the workload and surrounding software.

#aihardware#llmcompression#ondeviceai#machinelearning
Open analytics
On the site 1 views
min read 3 17.09.2026
Instagram

PrismML Releases 5.9GB Bonsai 2 Model Based on Qwen3.8 27B

Open the post on Instagram ↗