AI assistants often stumble at two points: they answer before a user finishes, or they need a pipeline of speech recognition, vision, dialogue and voice synthesis before reacting to what the user is showing.
ByteDance Seed released SeedRealtime on August 5 and says the native audiovisual full-duplex model is now fully available in the Doubao app. Users can update Doubao, choose "call" in the chat box and enter a video-call interface where the model receives video, audio and text together.
The entry point matters
Phoenix Tech and IT Home both confirmed the launch status and the Doubao entry point. That separates this release from multimodal demos that stay on a project page or waiting list.
The official technical claim is a unified end-to-end architecture. Instead of chaining ASR, vision, a dialogue model, TTS and external VAD rules, SeedRealtime is designed to perceive, understand, decide and respond in a continuous multimodal stream. ByteDance says the model focuses on three abilities: audiovisual joint understanding, proactive interaction, and smoother conversational rhythm.
“We hope AI interaction will no longer be limited to question-and-answer turns in clean rounds.”
The clearest number is also bounded: ByteDance says an end-to-end human evaluation found that audiovisual dialogue rhythm problems fell by half compared with a cascaded system. That refers to interruptions, late replies and noise-triggered responses. It should not be read as a public latency benchmark or a 50% speed gain.
The hard part is silence
The official examples show why this is harder than image description. SeedRealtime is shown identifying speakers at a group dinner, explaining a Chinese menu for a foreign diner, watching for a museum exhibit, correcting a coffee-machine operation and stopping on a target section while a user flips through a ResNet paper.
All of those cases test the same thing: whether the model can connect what it sees with whether it should speak. In video calling, the assistant has to ignore background chatter, remember what was seen a moment ago and intervene only when the user needs it.
What to watch next
The Chinese market gives this launch practical weight. Users are more likely to point a phone at menus, appliances, schoolwork, road signs or work documents. If "see and talk" becomes reliable inside Doubao, ByteDance gets higher-frequency feedback than a text chatbot can provide.
The limits are visible too. ByteDance lists lower end-to-end latency, more proactive perception and decision-making, stronger handling of multi-person scenes, and tool use as future work. The real checks are latency, false triggers, speaker-reference stability and actual task completion.
This is not a story about Doubao beating another model on a chart. It is a product test: can a consumer AI assistant keep up when the camera moves, people interrupt and the environment is noisy?
Sources: ByteDance Seed technical blog, Phoenix Tech, IT Home, CocoLoop; verification covers the Doubao app entry point, unified end-to-end architecture, three core abilities, the human-evaluation claim that rhythm problems fell by half, and the stated optimization roadmap.