PrismML Creates Tiny AI Models Designed to Run on PCs and Smartphones

PrismML’s Bonsai 2 compresses a large language model to run on PCs and potentially smartphones while maintaining most benchmark performance.

Sep 18, 2026 - 15:27
 2
PrismML Creates Tiny AI Models Designed to Run on PCs and Smartphones
Image Credit: TechAmerica.ai / AI-generated image

AI startup PrismML is taking a different approach to artificial intelligence development by focusing on smaller, high-performance reasoning models that can run directly on personal devices.

The company believes advanced AI models don’t need to be huge to perform well. Its goal is to compress powerful large language models into versions that can operate on PCs and potentially smartphones while maintaining most of their original capabilities.

PrismML recently released Bonsai 2 27B, the latest model in its Bonsai family. The system compresses Alibaba’s Qwen3.8 27B open-source model from a much larger format into a 5.9 GB version, reducing memory requirements by roughly 9 to 10 times.

Reducing AI Model Size Without Major Performance Loss

PrismML was founded by Caltech researchers and is led by Hassibi, a Caltech professor known for work in compression technologies. The startup also has Ion Stoica, co-founder of Databricks and director of Berkeley’s Sky Computing Lab, as an adviser.

The company says its compression technology allows models to become significantly smaller while keeping most of their original intelligence. Bonsai 2 reportedly achieves about 98% of Qwen’s aggregate benchmark performance, improving from the first Bonsai release, which reached about 95%.

PrismML says its earlier Bonsai model has already been downloaded more than 11 million times, while its smaller models have received millions of additional downloads.

How PrismML Compresses AI Models

The company’s approach focuses on reducing model-weight size, which stores the information learned during training. Traditional AI models often represent weights using 16-bit values, but PrismML uses a ternary approach that reduces them to three possible states: +1, -1, or 0.

By simplifying these values, the company can significantly reduce storage and memory requirements while attempting to preserve model capability. Developers can explore PrismML’s Bonsai models through its Hugging Face collection.

Moving AI Processing to Personal Devices

PrismML believes smaller AI models could help make advanced artificial intelligence more accessible by allowing more processing to happen directly on user devices instead of depending entirely on cloud infrastructure.

Running AI locally could offer benefits including improved privacy because user data does not need to be sent to remote servers. It could also make advanced AI features available on hardware that users already own.

The company is now working on applying its compression methods to larger models. PrismML says future releases could target models with hundreds of billions of parameters while continuing to focus on preserving their intelligence after compression.

As AI models continue to grow in size and complexity, companies like PrismML are exploring whether better compression techniques can make powerful AI systems smaller, faster, and more practical for everyday devices.

What's Your Reaction?

Like Like 0
Dislike Dislike 0
Love Love 0
Funny Funny 0
Angry Angry 0
Sad Sad 0
Wow Wow 0
Shivangi Yadav Shivangi Yadav’s current bio says she reports on technology-focused developments “in India”, but the same profile publishes stories about U.S. NHTSA investigations, Hugging Face, global AI startups and other international topics.