메뉴
HN
Hacker News 54일 전

노트북으로 즐기는 오픈소스 실시간 AI 음악 모델

IMP
8/10
핵심 요약

구글의 Magenta 팀이 실시간으로 반응하는 오픈소스 AI 음악 모델인 Magenta RealTime 2(MRT2)를 공개했습니다. 이전 버전과 달리 고성능 TPU/GPU 없이도 Apple Silicon 기반 맥북에서 매우 낮은 지연 시간(200ms)으로 작동하며, MIDI, 텍스트, 오디오를 통한 정밀한 제어가 가능합니다. 또한 C++ 기반의 고속 추론 엔진과 Python 라이브러리를 함께 제공하여 누구나 자신만의 AI 악기를 구축하고 DAW에 통합할 수 있습니다.

번역된 본문

Magenta RealTime 2: 오픈소스 로컬 실시간 음악 모델 2026년 6월 4일

여러분이 노트북에서 AI 악기를 만들고 연주할 수 있게 해주는 최첨단 오픈 모델이자 효율적인 실시간 추론 엔진인 Magenta RealTime 2(MRT2)를 공개하게 되어 매우 기쁩니다! 시작하려면 MacBook(Apple Silicon 필요)용 앱을 다운로드하세요.

  • 플러그인 번들 (MacOS)
  • GitHub에서 보기
  • 모델

프롬프트를 트랙으로 변환하기 위해 오프라인에서 작동하는 다른 대규모 생성형 음악 모델과 달리, MRT2는 텍스트뿐만 아니라 MIDI와 오디오로도 제어할 수 있는 라이브 인터랙티브 모델입니다. 사용자의 입력에 즉각적으로 반응하기 위해 기기 내에서 저지연 추론을 수행합니다. 독립 실행형 앱으로 실행하거나, DAW(디지털 오디오 워크스테이션)에 플러그인으로 추가하거나, 다른 음악 소프트웨어에 통합할 수 있습니다.

오픈 가중치(Open-weights) 모델과 함께, MRT2로 제작한 연주 가능한 악기 및 체험 모음도 공개합니다. 이 저지연 음악 모델을 통해 사운드 클로닝, 스타일 혼합, 라이브 반주 만들기를 직접 실험해 볼 수 있습니다.

라이브 음악 모델이 악기로 가지는 잠재력을 탐구하기 위해 오늘 다음을 공개합니다:

  • Magenta RealTime 2: MIDI, 텍스트, 오디오를 통한 저지연 실시간 제어가 가능한 고품질 실시간 음악 합성을 지원하는 오픈 가중치 모델(24억 파라미터).
  • 모델과 함께 JAX / MLX 및 SequenceLayers를 사용한 추론을 제공하는 오픈소스 Python 라이브러리(pip install magenta-rt)를 공개합니다.
  • MLX를 통해 MacBook GPU에서 효율적인 스트리밍 오디오 생성을 가능하게 하는 C++ 기반의 추론 엔진.
  • 추론 엔진 기반으로 구축된 예제 애플리케이션 제품군. 이는 Magenta RealTime 2의 창의적 잠재력을 엿볼 수 있게 해주며, 새로운 악기 및 소프트웨어 통합 구축을 시작하는 데 도움이 되는 참고 자료로 활용할 수 있습니다.

지난 10년 동안 Magenta 팀은 AI가 음악가를 대체하는 것이 아니라 도구가 되어야 한다는 비전을 추구해 왔습니다. 2017년에 머신러닝을 연주 가능한 하드웨어에 탑재한 최초의 신경망 신디사이저인 NSynth를 공개했습니다. 이후 DDSP, Piano Genie, 그리고 다양한 음악 스타일을 생성하고 혼합할 수 있는 최초의 라이브 음악 모델인 1세대 Magenta RealTime 등의 프로젝트를 통해 계속해서 AI 악기를 만들어 왔습니다. MRT2는 1세대보다 약 15배 낮은 지연 시간을 달성했으며, 일반 하드웨어에서 작동하고 DAW에 직접 통합되어 이 라이브 모델을 진정한 악기로 만들어 줍니다.

더 낮은 지연 시간과 확장된 제어 기능을 갖춘 라이브 음악 모델

  • 모델: Magenta RealTime / Magenta RealTime 2
  • 라이브 음악 생성: 지원함 / 지원함
  • 필요 하드웨어: TPU/GPU / MacBook
  • 프레임 크기: 2초 / 40ms
  • 제어 지연 시간: 약 3초 / 약 200ms
  • 제어 방식: 텍스트, 오디오 / 텍스트, 오디오, MIDI
  • 모델 크기: 7억 6천만 / 2억 2천만 / 24억 / 2억 3천만

MRT와 MRT2는 모두 SpectroStream 코덱의 오디오 토큰 시퀀스에서 작동하는 코덱 언어 모델이지만, MRT2는 프레임 수준의 자기회귀와 프레임 정렬 조건화(frame-aligned conditioning)를 수행하여 더 낮은 지연 시간을 달성합니다. 풍부한 음악적 제어를 가능하게 하기 위해 MRT2는 오디오나 텍스트 형태의 스타일 프롬프트와 함께 MIDI 입력을 지속적으로 따르는 오디오를 모델링하도록 설계되었습니다. 프롬프트는 MusicCoCa를 통해 임베딩됩니다.

상호작용 지연을 최소화하기 위해 두 신호는 모든 생성 단계에서 프레임 정렬 조건화로 주입되어 모델이 단일 프레임(40ms, 경험적 지연의 추가 소스 포함, 아래 참조) 내에서 신호 변화에 반응할 수 있도록 합니다. 이 접근 방식의 핵심은 메모리 요구 사항을 제한하면서 연속적인 스트리밍 생성을 가능하게 하는 인과적 슬라이딩 윈도우 어텐션 메커니즘(causal sliding window attention mechanism)을 사용하는 것입니다. 이와 함께 임의 길이에 대한 일반화를 개선하고 긴 컨텍스트 생성 중 발생하는 컨텍스트 제거 아티팩트(예: 잔향 및 피드백)를 줄이기 위해 학습 가능한 어텐션 임베딩(learnable attention embeddings)도 통합되었습니다.

MLX를 통한 빠른 C++ 추론 엔진

원래의 Magenta RealTime은 고성능 GPU 또는 TPU가 필요했지만, Magenta RealTime 2는 실제 음악가들이 사용하는 하드웨어에 라이브 생성을 가능하게 합니다. 이를 달성하기 위해 MRT2가 Apple Silicon에서 기본적으로 실행될 수 있도록 MLX로 구동되는 C++ 추론 엔진을 구축했습니다. Apple의 MLX 프레임워크...

원문 보기
원문 보기 (영어)
Magenta RealTime 2: Open & Local Live Music Models Jun 4, 2026 We’re excited to share Magenta RealTime 2 (MRT2), a state-of-the-art open model and efficient real-time inference engine that enables you to build and play AI musical instruments on your laptop! To get started, download the apps on your MacBook (requires Apple Silicon). Plugin Bundle (MacOS) View on GitHub Models Unlike other large generative music models that work offline to turn a prompt into a track, MRT2 is a live, interactive model that you can control with MIDI and audio, in addition to text. It performs low-latency on-device inference to respond to your inputs instantly. You can run it as a standalone app, drop it into your DAW, or integrate it into other music software. In addition to the open-weights model, we are releasing a collection of playable instruments and experiences built with MRT2. Experiment with cloning sounds, blending styles, and creating live accompaniment with this low-latency music model. To explore the potential of live music models as instruments, today we are releasing: Magenta RealTime 2, an open-weights model (2.4B parameters) capable of high-quality real-time music synthesis with low-latency real-time controls via MIDI, text, and audio . Alongside our model, we release an open source Python library ( pip install magenta-rt ) offering inference via JAX / MLX using SequenceLayers . An inference engine written in C++, enabling efficient streaming audio generation on a MacBook GPU via MLX . A suite of example applications built on the inference engine. These offer a glimpse into the creative potential of Magenta RealTime 2, and serve as references to help you get started building new instruments and software integrations. For a decade, the Magenta team has championed a vision of AI as a tool for musicians, never a replacement. We released our first neural synthesizer, NSynth , back in 2017 which put machine learning into playable hardware . We continued creating AI Instruments with projects such as DDSP , Piano Genie , and the first version of Magenta RealTime , our debut live music model capable of generating and blending a wide range of musical styles. MRT2 achieves ~15x lower latency than version one, works on standard hardware and integrates directly into DAWs, making this live model a true musical instrument. A live music model with lower latency and expanded control Magenta RealTime Magenta RealTime 2 Live music generation ✅ ✅ Hardware required TPU/GPU MacBook Frame size 2s 40ms Control latency ~3s ~200ms Control modalities Text, Audio Text, Audio, MIDI Model sizes 760M / 220M 2.4B / 230M Both MRT and MRT2 are codec language models operating on sequences of audio tokens from the SpectroStream codec, but MRT2 achieves lower latency by performing frame-level autoregression with frame-aligned conditioning. To enable expressive musical control, MRT2 is designed to model audio that continuously follows MIDI inputs, alongside style prompts which can be either audio or text; prompts are embedded via MusicCoCa . For minimal interaction lag, both signals are injected as frame-aligned conditioning at every generation step, allowing the model to react to changes in the signal within a single frame (40 ms, plus additional sources of empirical latency, see below ). Key to this approach is the use of a causal sliding window attention mechanism to enable continuous streaming generation while bounding memory requirements. Alongside this, learnable attention embeddings are also incorporated to improve generalization to arbitrary durations and context eviction artifacts (e.g., ringing and feedback) during long-context generation. Fast C++ inference engine via MLX While the original Magenta RealTime required a high-power GPU or TPU, Magenta RealTime 2 brings live generation to the hardware musicians actually use. To achieve this, we built a C++ inference engine powered by MLX that allows MRT2 to run natively on Apple Silicon . Apple’s MLX framework provides the link between Python and C++. More specifically, we use MLX to compile the MRT2 model, implemented using the SequenceLayers library , into an .mlxfn file which is a model container that bundles the weights and computational graph. Our C++ inference engine loads that file and uses the MLX runtime to efficiently execute it on Apple Silicon GPUs. The inference engine handles other necessary infrastructure (model state, audio buffering / resampling, MIDI input) and can be embedded into many music application frameworks where C++ supported. MLX allows MRT2 to run on Apple Silicon (M-series): both model sizes can run offline (non-real-time) inference on any Apple Silicon Mac, while real-time streaming (generating audio faster than playback) is supported on the following devices: Model Platform Base (2.4B) MacBook M3 Pro (or higher) MacBook M2 Max (or higher) Small (230M) Any Apple Silicon MacBook , including MacBook Air A suite of example applications for musicians and developers A key goal of Magenta RealTime 2 is to allow musicians to integrate live music models within existing software, and help developers build custom applications. To help you get started, our codebase provides several examples , including standalone apps, plugins and extensions. What’s Next? Our team members have been building new instruments with machine learning for nearly 10 years , excitedly making unique and quirky sounds from statistical knowledge of music. With Magenta RealTime 2, AI instruments are finally starting to gain the controllability and immediacy we expect from music creation tools, but plenty remains to be explored. From even more interaction and lower control latency, to audio streaming inputs that can enable jamming and real-time audio control, we look forward to expanding the capabilities of live music models further. Stay tuned for future updates! And in the meantime, we are also excited to bring more features and example applications to MRT2 soon, including: Finetuning , allowing anyone to customize the model by directly training on their own data. Example performance tools created in collaboration with Manaswi Mishra . In the next few days, we will also be at the Music Technology Hackathon in Boston , where we are presenting a challenge centered around Magenta RealTime 2. We look forward to seeing what everyone will come up with! Citation Please cite our work as: Magenta Team. “Magenta RealTime 2: Open & Local Live Music Models”. https://magenta.withgoogle.com/magenta-realtime-2. June 2026 @article{mrt2, title = {Magenta RealTime 2: Open & Local Live Music Models}, author = {Magenta Team}, year = {2026}, note = {https://magenta.withgoogle.com/magenta-realtime-2} } Appendix: Technical Details Low-latency streaming generation Some background on Codec Language Modeling. A codec language model (LM) operates on discrete sequences of tokens from a neural audio codec. Here a codec refers to a pair of functions, an encoder and decoder, that convert audio to and from a discrete, compressed representation while minimizing distortion. More formally, the encoder is a function mapping raw stereo audio waveforms \(\textbf{a} \in \mathbb{R}^{T f_s \times 2}\) into matrices of discrete tokens \(\mathbf{x} \in \mathbb{V}_c^{Tf_k \times d_c}\) where \(T\) is the duration in seconds, \(f_s\) the audio sampling rate, \(f_k\) the token frame rate, \(\mathbb{V}_c\) the codec vocabulary, and \(d_c\) is the number of tokens per frame. In this case, \(d_c\) refers to the “depth” of the residual vector quantization algorithm, referring to the iterative quantization of continuous embeddings of each audio frame. The goal of the codec LM is to model these token matrices. For efficiency, an increasingly common approach is to adopt a hierarchical autoregressive framework using a pair of Transformers: one which compresses temporal history into fixed-length embedding vectors (\(\texttt{Temporal}_\theta\)), and another which iteratively decodes tokens depth-