Prismml's compact llms aim to revolutionize AI accessibility
PrismML, a startup founded by Caltech researchers and led by professor Babak Hassibi, is developing highly compressed large language models (LLMs) that can run on PCs and smartphones. Its latest model, Bonsai 2 27B, reduces the memory footprint of Qwen3.8 27B from Alibaba by 9x to 10x, fitting into just 5.9 GB. Hassibi claims the model retains 98% of Qwen’s benchmark performance, up from 95% in the previous version. The startup also counts Ion Stoica, co-founder of Databricks, as an advisor. PrismML’s compression technique, called 'ternary' weights, simplifies model data storage by using three values instead of 16 bits. The company plans to apply this to larger models in the next few months, aiming for better performance retention. Stoica highlights the potential for on-device AI, emphasizing privacy and cost-efficiency.
Investors including Khosla Ventures and Caltech back the startup, which has raised $22.25 million in seed funding. While not the only company working on LLM compression, PrismML’s focus on maintaining performance sets it apart. Hassibi notes that while perfect benchmark parity remains elusive, even minor performance drops may not significantly impact real-world use. The startup’s smaller models have been downloaded over 13.6 million times, showcasing growing adoption.