메뉴
HN
Hacker News • 4일 전

M5 울트라 맥 스튜디오, 로컬 AI 에이전트용 꿈의 머신

IMP
7/10
핵심 요약

MacStories 리뷰어가 RAM 256GB를 장착한 M5 Ultra Mac Studio를 테스트한 결과, 클라우드 비용 없이 로컬 모델 기반 개인 비서 에이전트를 빠르게 구동할 수 있는 최적의 머신이라 평가했습니다. 메모리 대역폭에서는 RTX 5090이 여전히 앞서지만, 크기·발열·소음·통합 메모리 구조 측면에서 M5 Ultra가 압도적으로 실용적이라며 로컬 AI 활용 패러다임을 바꿨다고 강조했습니다.

번역된 본문

macOS 27: MacStories 리뷰. OS 27 리뷰 부가 자료: 전자책, 단축어, iOS 및 iPadOS 27 리뷰 제작기. 이번 주 스폰서: LookAway — Mac 사용 중 흐름을 깨지 않고 눈을 쉴 수 있도록 정기적인 휴식을 도와주는 앱입니다.

M5 Ultra Mac Studio. 지난 며칠 동안 저는 (현재) 최상위 사양인 RAM 256GB의 M5 Ultra Mac Studio를 테스트했습니다. 결론부터 말하겠습니다: M5 Ultra Mac Studio는 로컬 AI 에이전트를 위한 꿈의 머신입니다. 이 컴퓨터는 추가 클라우드 비용 없이도 뛰어난 성능으로 로컬 모델 기반 개인 비서를 구동할 수 있게 해줍니다. 로컬 모델이 클라우드 모델의 지능과 속도에 결코 미치지 못할 것이라는 이유로 OpenClaw나 Hermes Agent를 로컬 모델로 테스트하는 것에 회의적이었다면, 이 Mac은 그 생각을 바꿔놓을 것입니다.

지난 목요일부터 저는 이 Mac Studio를 전작인 RAM 512GB의 M3 Ultra, 그리고 RTX 5090을 장착한 제 데스크톱 게이밍 PC와 비교해왔습니다. 크기, 가격, 열 관리 성능 — 게다가 애플의 통합 메모리(unified memory) 접근 방식까지 고려하면 — M5 Ultra Mac Studio는 로컬에서 모델을 구동하는 것과 그것이 지금 가능하게 하는 일에 대한 제 생각을 근본적으로 바꿔놓았습니다. 물론 5090은 더 높은 메모리 대역폭 덕분에 여전히 M5 Ultra보다 우위에 있습니다. 하지만 제 PC 빌드의 거대한 크기, 발열, 소음을 고려하면 저는 언제라도 M5 Ultra Mac Studio를 선택할 것입니다. 게다가 이것은 Mac이기도 합니다. 보기 좋고 형편없지 않은 운영체제와 활기찬 앱 생태계까지 갖추고 있죠. (Windows 팬 여러분, 죄송합니다만, 마이크로소프트 소프트웨어는 결국 제 마음을 얻지 못할 겁니다.)

이 글에서 살펴보겠지만, M5 Ultra Mac Studio에서 최신 Qwen3.8-Flash-Next 모델을 구동한 경험이 너무 만족스럽고 빨라서, iOS용 Open Minis와 Hermes Agent 모두에서 기본 모델로 설정했습니다. 그렇습니다. 제가 가장 많이 사용하는 개인 비서 — 사실 Siri AI보다 더 자주 쓰는 — 는 이제 완전히 Mac Studio에서 로컬로 구동되는 모델로 작동합니다. 더구나 M5 Ultra의 더 빠른 GPU와 높은 메모리 대역폭 덕분에 이 에이전트들은 더 빠르게 응답을 시작하고, 더 큰 컨텍스트 윈도우에서도 빠른 속도를 유지하며, 세션이 길어져도 느려지지 않고 긴 다중 턴(multi-turn) 루프를 실행할 수 있습니다. 이 덕분에 Mac의 Codex 앱에서도 로컬 모델을 사용하고 있습니다 — 메인 스레드로든, GPT-6 Astra가 조율하는 서브에이전트로든 — 그리고 그 경험이 훌륭했습니다.

M5 Ultra Mac Studio의 Codex에서 로컬 서브에이전트를 구동하는 모습. 먼저 밝혀둘 점은, 저는 본업이 AI 개발자가 아니라는 것입니다. 모델을 학습하거나 파인튜닝하지 않습니다. 저는 본질적으로 만지고 실험하는 것을 좋아는 사람이고, 지금까지 1년 넘게 로컬 AI 모델을 가지고 놀아왔습니다. 이번 여름에는 제가 진행하던 큰 프로젝트를 위해 로컬 AI 활용에 올인했는데, 이는 다음 섹션에서 설명하겠습니다. 이 글의 목표는 두 가지를 혼합해서 전달하는 것입니다: 4일 동안 실행한 (수많은) 테스트에 기반한 수치와 시각화, 그리고 제 워크플로와 MacStories 업무 수행 방식에 로컬 AI를 적용한 실제 사용 사례 설명입니다. 자, 시작해봅시다.

왜 로컬 AI인가? 먼저 방 안의 코끼리부터 짚고 넘어갑시다: 클라우드 프론티어 모델이 더 좋고 종종 더 빠른데, 굳이 로컬 AI를 왜 신경 써야 할까요? 당연한 질문입니다. 이런 모델을 구동하려면 비싼 하드웨어가 필요하고, 투자비를 회수할 무렵이면 가장 비싼 Anthropic 구독을 몇 년 쓰고도 돈을 아꼈을 수 있으며, 성능도 더 좋았을 겁니다. 이 질문에 대한 답은 사람마다 다를 것입니다. 어떤 사람은 프라이버시 때문이라고 말할 겁니다. 민감한 데이터와 문서를 외부 클라우드에 업로드하기보다 로컬 지능에 맡기고 싶은 거죠. 다른 사람들은 그냥 멋있기 때문이라고 말할 수도 있습니다 — 저도 동의합니다. 어떤 사람에게는 업무상 필요한 일이기도 합니다. AI 개발자라면 자체 어댑터를 학습하거나 파인튜닝하기 위해 훌륭한 로컬 환경을 갖추는 것이 합리적입니다.

원문 보기
원문 보기 (영어)
macOS 27: The MacStories Review Our OS 27 Review Extras: eBooks, Shortcuts, and the Making of the iOS and iPadOS 27 Review iOS and iPadOS 27: The MacStories Review This Week's Sponsor: LookAway It helps you take regular breaks from your Mac to rest your eyes, without breaking your flow. The M5 Ultra Mac Studio. For the past few days, I’ve been testing the ( currently ) top-of-the-line M5 Ultra Mac Studio with 256 GB of RAM . I’ll cut to the chase: the M5 Ultra Mac Studio is a dream machine for local AI agents. This computer makes it possible to run personal assistants powered by local models with great performance and no additional cloud costs. If you’ve been skeptical of testing OpenClaw or Hermes Agent with local models because they’d never be even remotely near the intelligence and speed of cloud ones, this Mac will change your mind about that. Since last Thursday, I’ve been comparing this Mac Studio to its predecessor, the M3 Ultra with 512 GB of RAM , as well as my own desktop gaming PC with an RTX 5090 inside. For its size, price, thermal performance – not to mention Apple’s approach to unified memory – the M5 Ultra Mac Studio has fundamentally changed how I think about models running locally and what they can enable now. A 5090, of course, still has an edge over the M5 Ultra thanks to its higher memory bandwidth . But considering the sheer size of my PC build, as well as its heat and noise, I would prefer an M5 Ultra Mac Studio any day. It also happens to be a Mac, with an operating system that looks nice and doesn’t suck, plus a vibrant app ecosystem. (Windows fans, I’m sorry , but Microsoft software will never get my sympathy.) As I’ll explore in this article, running the latest Qwen3.8-Flash-Next model on the M5 Ultra Mac Studio has been so nice and fast, I’ve made it my default in both Open Minis for iOS and Hermes Agent. That’s right: the personal assistants I use the most – more than Siri AI , in fact – are now entirely powered by a model running locally on a Mac Studio. Furthermore, thanks to the M5 Ultra’s faster GPU and higher memory bandwidth, these agents start responding more quickly, stay fast at larger context windows, and can run long, multi-turn loops without slowing to a crawl as the session grows. Because of this, I’ve also been using local models in the Codex app on my Mac – either as main threads or subagents orchestrated by GPT-6 Astra – and I’ve had a great experience doing so. Local subagents running in Codex on the M5 Ultra Mac Studio. I should note upfront that I’m not an AI developer by trade: I do not train or fine-tune models. I’m a tinkerer at heart, and I’ve been playing around with local AI models for over a year at this point . This summer, I went all-in on local AI usage for a big project I was working on, which I will explain in the following section. My goal with this article is to provide you with a mix of two things: numbers and visualizations based on the (many) tests I’ve run over the course of four days, and an explanation of my practical use cases for local AI applied to my workflow and how I get things done for MacStories. Let’s dive in. Why Local AI? Let’s address the elephant in the room first: why bother with local AI at all when cloud frontier models are better and often faster? It’s a fair question. You need expensive hardware to run these models, and by the time you’ve repaid your investment, you could have used the most expensive Anthropic subscription for several years, still saved money, and got better performance in return. Different people will have different answers to this question. Some might say they use local models because of privacy: they’d rather rely on local intelligence for sensitive data and documents than upload anything to an external cloud. Others might argue that it’s simply cool – and I do not disagree. For some, it’s a work-related task: if you’re an AI developer, it makes sense to have a great local setup for training your own adapters or fine-tuning models. For me, the journey into local AI has been characterized by a mix of the “ cool, why not? ” factor of it all as well as considerations about privacy and costs. As I will share later this week with Club MacStories members , my research and writing setup for the iOS and iPadOS 27 review this summer has been powered and made possible by local AI. Back in June, I created an internal app, called Desk , to organize hundreds of notes, sessions, PDF documents, and clipped webpages related to iOS and iPadOS 27, as well as chapters of the review. By the end of the process, the project consisted of 310 documents. In Desk, a team of agents – all based on DeepSeek V4 Flash , plus olmOCR for PDFs – ran 24/7, for 99 days, to perform the following tasks: Transcribe my favorite WWDC sessions (using summarize plus LLM processing) Extract features of iOS and iPadOS 27 from clipped webpages, sessions, PDF guides, and my own notes Cross-reference features across different sources, and keep track of which features belonged to which chapter of the review Extract features and bugs from screenshots I uploaded Work with the Notion API to organize everything across multiple databases One of the views of Desk, the app powered by the Notion API and local AI I used for my iOS 27 review. The local AI agent runs in my custom Desk app. When I started working with this setup in early June, I quickly realized that relying on the OpenAI or Anthropic APIs for this kind of always-on, persistent background task would be…cost-prohibitive, to say the least. So I pivoted to local AI, and the result is the iOS and iPadOS 27 review you can read on MacStories . It was all written by me, the old-fashioned human way. But the entire research stack, deep-linking between notes, and keeping track of new features and betas were all performed by my agents, running locally on the Mac Studio, for a total cost of $0. If you don’t think that’s neat, or a powerful concept to explore, then this article probably isn’t for you – and I understand. Dealing with these models is fiddly, and it’s not something I would ever recommend to someone who (rightfully) just wants to pay $20 to use Claude Cowork. This kind of setup is, by definition, the bleeding edge of AI workflows at the moment. If you fall on the other end of the spectrum, though, and if you think this kind of stuff is neat…let me tell you: the M5 Ultra Mac Studio is a massive leap in performance for local models powered by MLX , and I have a few examples to prove it. A Leap for Prompt Processing and Generation As you may have seen from the announcement and my initial coverage , the M5 Ultra Mac Studio looks identical to the M3 Ultra model it replaces, but it comes with an all-new Apple silicon architecture that uses UltraFusion to connect two dual-die M5 Max chips to form a quad-die architecture, which is a first for the Apple ecosystem. As far as local AI workloads are concerned, there are two areas we have to pay attention to (and which I have been following since my coverage of the M5 iPad Pro for local AI last year): GPU and memory bandwidth. The M5 Ultra has a next-gen GPU with 80 cores, each with a Neural Accelerator that grants it up to 4.5× the peak GPU compute for AI compared to the M3 Ultra. As for memory, Apple’s unified memory architecture still tops out at 512 GB as before (although that model will come out in late October), but its bandwidth has jumped from 819 GB/s to 1.2 TB/s, or 50% higher than the M3 Ultra. With these numbers in mind, I started testing the M5 Ultra against the M3 Ultra with 512 GB of RAM and my RTX 5090. I’ll share more details on testing below, but the short version is this: with the M5 Ultra, you spend considerably less time waiting for a model to read your prompt and begin generating a response; and when it does start answering, text appears much faster than it used to on the M3 Ultra. These two improvements alone make the machine viable for modern agentic loops that require fast it