lupAI
novos-desenvolvimentos

Prime intellect introduces prime inference for efficient model serving

Prime IntellectSource: MarkTechPost03/10/2026, 07:07
Prime Intellect has launched Prime Inference, a platform designed for serving frontier open-source models. The platform offers serverless endpoints and reserved GPU capacity across multiple datacenters. Before its public release, it processed nearly a trillion tokens daily, driven by RL rollouts, synthetic data generation, evaluations, and long-running coding agents. Prime Inference is part of Prime Intellect’s open training stack, which includes post-training tools like prime-rl, verifiers, and sandboxes. The serving layer enables deployed models to generate production traces that feed back into training. Prime reports its GLM-5.3 endpoint ranks among the fastest on OpenRouter, with a near-zero tool-call error rate and 100% uptime since launch. The platform integrates NVIDIA Dynamo, vLLM, Mooncake, and FlashInfer, with contributions from Inferact and NVIDIA. It is optimized for agentic workloads, where a typical agent turn adds about 6,000 tokens to a 140,000-token prompt. Prime benchmarks this with SemiAnalysis AgentX and injected cold arrivals. Prefill and decode processes run on separate GPU groups, with Dynamo handling routing and vLLM executing the model. Decoders pull computed KV through NIXL, achieving nearly 40% lower p90 inter-token latency. Cache-aware routing, managed by Dynamo’s KV-aware router, keeps sessions on the same decoder between turns, with Mooncake adding a second KV tier in host DRAM.
Prime intellect introduces prime inference for efficient model serving — lupAI