Ant Group opens code of Ling-3.0-tiny model optimized for local inference
Ant Baili, Ant Group's AI division, launched the open-source version of the Ling-3.0-tiny model, a lightweight Mixture of Experts (MoE) architecture model with 7.9 billion total parameters, activating 1.3 billion per token. The model was specifically designed for low-cost inference and was tested on platforms like NVIDIA DGX Spark and Apple Silicon computers, achieving inference speeds of approximately 86-90 tokens per second on a MacBook Pro M4. The model is available in multiple weight formats (BF16, FP8, and INT4) to meet different deployment requirements.