New distributed training architecture enables resilient AI model development
Google introduced Decoupled DiLoCo, a distributed architecture for training large language models across distant data centers with reduced bandwidth requirements and improved hardware resilience. By dividing training into decoupled compute islands with asynchronous data flow, the approach isolates hardware failures and enables flexible training across globally distributed infrastructure, addressing scalability challenges of frontier model training.