BL
MarkTechPost • 47일 전
바이트댄스, 실시간 시청·청취·대화 가능한 통합 멀티모달 AI 공개
IMP 8/10
핵심 요약
바이트댄스의 Seed 팀은 시각, 청각, 텍스트를 하나로 통합한 네이티브 오디오-비주얼 풀듀플렉스 대형 언어 모델(LLM)인 'SeedRealtime'을 공개했습니다. 이 모델은 한 번씩 주고받는 기존 방식을 넘어, 연속적인 멀티모달 스트림을 통해 실시간으로 상호작용할 수 있어 완벽한 옴니모달(Omni-modal) 인터랙션 구현에 한 걸음 다가섰다는 점에서 중요합니다.
번역된 본문
바이트댄스(ByteDance)의 Seed 팀은 네이티브 오디오-비주얼(Native Audio-Visual) 풀듀플렉스(Full-Duplex) 대형 언어 모델인 'SeedRealtime'을 공개했습니다. 이 모델은 단일 통합 아키텍처 내에 오디오, 비디오, 텍스트를 융합했습니다. 한 번에 한 턴씩 주고받는 대화가 아닌, 연속적인 멀티모달 스트림(Multimodal Stream)을 통해 실시간으로 상호작용합니다. Seed 팀은 이를 옴니모달(Omni-modal) 상호작용을 향한 발걸음으로 규정하며, 통합된 시청각 이해(Joint Audio-Visual Understanding) 등 세 가지 주요 혁신을 이루었다고 주장합니다. […] 이 글 '바이트댄스 Seed, 보고 듣고 말하는 하나의 모델인 네이티브 오디오-비주얼 풀듀플렉스 LLM SeedRealtime 소개'는 MarkTechPost에 처음 게재되었습니다.
원문 보기 (영어)
ByteDance’s Seed team has introduced SeedRealtime, a native audio-visual full-duplex LLM. The model fuses audio, video and text in a single unified architecture. It interacts in real time over continuous multimodal streams, rather than one turn at a time. Seed positions it as a step toward omni-modal interaction, and claims three breakthroughs: joint audio-visual understanding, […]
The post ByteDance Seed Introduces SeedRealtime: a Native Audio-Visual Full-Duplex LLM That Watches, Listens and Speaks in One Model appeared first on MarkTechPost.