ByteDance unveils SeedRealtime, a full-duplex audio-visual LLM
ByteDance's Seed team introduced SeedRealtime, a native audio-visual large language model capable of real-time full-duplex conversation. The model fuses audio, video, and text in a single unified architecture, enabling simultaneous listening and speaking with natural conversational timing. Currently deployed in ByteDance's Doubao consumer assistant, SeedRealtime represents a shift from cascade architectures that chain separate components, moving perception and decision-making into parallel processing within a single end-to-end model.