메뉴
BL
The Decoder • 45일 전

미스트랄, EU 데이터 처리 및 우선 접근 제공...중요한 제한도 동반

IMP
7/10
핵심 요약

미스트랄(Mistral)이 AI 요청을 EU 또는 미국 서버로 라우팅하는 리전별 추론 기능과, 트래픽 폭주 시 대기열을 우선 처리해 주는 Priority Tier(우선 등급)를 유료로 출시했습니다. 데이터 주권과 지연 시간 단축을 기대할 수 있으나, 에이전트나 파일 API 등 상태 저장 기능은 제외되고 가용성 SLA 보장 수치도 낮아 기업 실무자들은 도입 전 세부 제약을 면밀히 검토해야 합니다.

번역된 본문

미스트랄(Mistral)이 고객들에게 AI 요청을 유럽 또는 미국 서버를 거치도록 라우팅할 수 있는 옵션을 제공하며, 트래픽 최대치일 때 우선순위 대기열에 액세스할 수 있는 권한을 판매하기 시작했습니다. 두 가지 옵션 모두 추가 요금이 부과되며, 리전별 라우팅이 모든 기능이나 데이터를 아우르는 것은 아닙니다.

기업이 자체 제품에 미스트랄의 AI를 통합할 때, 요청은 해당 제공업체의 서버 중 한 곳으로 전달됩니다. 기업 고객에게는 두 가지 사항이 중요합니다. 서버가 실제로 어디에 위치하는가, 그리고 모두가 동시에 요청을 보낼 때 어떤 일이 발생하는가입니다. 미스트랄은 이제 이 두 가지에 대한 해결책을 판매하며, 이를 더 광범위한 유럽 AI 주권 전략의 일환으로 블로그 게시물을 통해 설명했습니다.

리전별 추론(Regional inference)은 이제 일반적으로 사용할 수 있습니다. 고객은 유럽 엔드포인트(api.eu.mistral.ai) 또는 미국 엔드포인트(api.us.mistral.ai)로 요청을 보낼 수 있으며, 처리 과정은 해당 지역 내에 머물게 됩니다. 이는 고객 데이터가 EU를 결코 떠나지 않는다는 것을 증명해야 하는 은행, 정부 기관 및 보험사에게 중요합니다. 더 짧은 네트워크 경로는 더 낮은 지연 시간(latency)을 의미하기도 합니다. 기본 엔드포인트를 사용하는 경우 요청이 처리되는 위치에 대해서는 어떠한 보장도 받을 수 없습니다. 리전별 라우팅에는 표준 가격에 10%가 추가됩니다.

EU 데이터 처리의 중요한 한계점 하지만 플랫폼의 부가 기능 도구 중 '함수 호출(Function calling, 모델이 외부 API를 트리거하는 기능)'만이 리전별 엔드포인트와 함께 작동합니다. 에이전트(Agents), 일괄 처리(Batch processing), 파일 관리는 해당 리전 엔드포인트에서 사용할 수 없습니다. 모델 선택도 지역에 따라 다릅니다. 미스트랄은 고정된 목록을 게시하지 않으며, 고객은 각 엔드포인트를 쿼리하여 어떤 모델이 있는지 확인해야 합니다.

이러한 제한의 예상되는 이유는 단순한 모델 쿼리와 중간 상태 저장 간의 차이 때문입니다. 표준 모델 호출에는 영구 스토리지가 필요하지 않습니다. 반면 에이전트, 일괄 작업 및 파일 저장은 단일 호출 이후까지도 중간 단계나 업로드된 문서와 같은 데이터를 유지합니다. 미스트랄은 이러한 기능을 "상태 저장(Stateful)"이라고 부릅니다. 회사가 확인하지는 않았지만, 아마도 이 기능들은 추가적인 온사이트(현지) 인프라가 필요할 것으로 보입니다.

또한 그 범위는 "주권(sovereignty)"이라는 단어가 암시하는 것보다 좁습니다. 문서에 따르면, 계정 설정, API 키, 결제 및 사용 통계는 선택한 리전 외부에서 여전히 처리될 수 있습니다. 블로그 게시물은 또한 하청업체에게 지역 외부로의 제한적이고 안전한 데이터 전송도 언급했습니다. 리전별로 적용되는 것은 플랫폼 전체가 아닌 연산(compute) 단계뿐입니다. 요청이 이후에 저장되거나 로깅되는지 여부는 'Zero Data Retention(데이터 보관 제로)'라는 별도의 설정에 따라 결정됩니다.

실제로 이것이 의미하는 바는, 모델에 직접 계약서 텍스트를 보내는 고객은 EU 엔드포인트를 통해 이를 실행할 수 있다는 것입니다. 하지만 Files API를 통한 에이전트나 파일 관리가 필요한 사람에게는 동일한 보장을 받을 수 없습니다.

기업이 더 빠른 대기열에 비용을 지불하는 이유 두 번째 제공 항목은 현재 오픈 베타 버전인 Priority Tier(우선 등급)입니다. 모든 고객은 동일한 데이터 센터를 공유합니다. 한 번에 많은 요청이 몰리면 응답 시간이 길어집니다. Priority Tier는 일종의 패스트 트랙입니다. 혼잡할 때 유료 고객의 요청은 일반 트래픽보다 먼저 처리됩니다. 미스트랄은 고객 서비스 챗봇이나 공장의 생산 시스템처럼 지연 시간이 실제 금전적 손실로 이어지는 사용 사례를 타겟팅하고 있습니다.

이 등급에는 99.5%의 가동 시간 SLA(서비스 수준 계약)가 포함됩니다. 이는 한 달에 대략 3시간 반의 다운타임을 허용하는 계약상 보장된 서비스 수준입니다. 미스트랄의 표준 등급에는 이러한 보장이 없습니다. 고객은 service_tier라는 단일 API 매개변수를 통해 우선 액세스를 활성화합니다. 이 값을 'auto'로 설정하면 용량이 있을 때 요청이 패스트 트랙을 통해 전달됩니다. 기본값은 'standard_only'이며, 이는 일반 경로를 택합니다. 각 고객은 분당 우선 처리되는 요청 수에 대해 개별적으로 협상된 속도 제한(Rate limit)을 받습니다. 이 제한을 초과해도 요청이 실패하지는 않습니다. 그저 표준 처리 방식으로 폴백(Fallback)될 뿐입니다. API 응답은 실제로 어떤 등급이 요청을 처리했는지 보여주므로 고객은 이를 확인할 수 있습니다.

원문 보기
원문 보기 (영어)
Mistral now offers EU data processing and priority access, but both come with important limits Maximilian Schreiner View the LinkedIn Profile of Maximilian Schreiner Aug 12, 2026 Mistral Mistral is giving customers the option to route AI requests through servers in either Europe or the US, and selling priority queue access during peak traffic. Both come with a surcharge, and the regional routing doesn't cover all features or data. When companies build Mistral's AI into their own products, requests go to one of the provider's servers. Two questions matter for enterprise customers: where that server actually sits, and what happens when everyone sends requests at once. Mistral is now selling answers to both, framing them in a blog post as part of a broader European AI sovereignty strategy. Regional inference is now generally available. Customers can send requests to a European endpoint (api.eu.mistral.ai) or a US endpoint (api.us.mistral.ai), and processing stays in that region. That matters for banks, government agencies, and insurers that need to prove customer data never leaves the EU. Shorter network paths also mean lower latency. Anyone using the default endpoint gets no guarantee about where their request is processed. Regional routing costs 10 percent on top of standard pricing. EU data processing comes with significant limits However, among the platform's add-on tools, only function calling works with regional endpoints, meaning the model's ability to trigger external APIs. Agents, batch processing, and file management aren't available at the regional addresses. Model selection varies by region, too. Mistral doesn't publish a fixed list, customers have to query each endpoint to see what's there. The likely reason is the gap between a simple model query and storing intermediate state. A standard model call needs no persistent storage. Agents, batch jobs, and file storage hold data beyond a single call, things like intermediate steps or uploaded documents. Mistral calls these features "stateful." They probably require extra on-site infrastructure, though the company hasn't confirmed that. The scope is also narrower than the word "sovereignty" suggests. Account settings, API keys, billing, and usage stats can still be processed outside the chosen region, according to the documentation. The blog post also mentions limited, secured transfers to subcontractors outside the region. What's regional is the compute step, not the whole platform. Whether requests get stored or logged afterward depends on a separate setting called Zero Data Retention. What this means in practice is that customers who send contract text directly to a model can run it through the EU endpoint. Anyone who needs agents or file management through the Files API won't get the same guarantee. Why companies would pay for a faster queue The second offering is the Priority Tier , currently in open beta. All customers share the same data centers. When lots of requests hit at once, response times go up. The Priority Tier is a fast lane: paying customers' requests get processed ahead of regular traffic when things get busy. Mistral is targeting use cases where latency costs real money, like a customer service chatbot or a production system on a factory floor. The tier includes an uptime SLA of 99.5 percent, a contractually guaranteed service level that allows roughly three and a half hours of downtime per month. Mistral's standard tier has no such guarantee. Customers activate priority access through a single API parameter called service_tier. Setting it to "auto" sends the request through the fast lane when capacity is available. The default value is "standard_only," which takes the regular path. Each customer also gets individually negotiated rate limits for how many requests per minute get priority treatment. Going over that limit doesn't fail the request. It just falls back to standard processing. The API response shows which tier actually handled the request, so customers can check whether they're getting what they pay for. Mistral charges 1.75x the standard price, a 75 percent surcharge. Discounts from prompt caching, where repeated text segments are stored and billed at lower rates, still apply. Those discounts can reach 90 percent and get calculated first according to the documentation, with the priority surcharge applied after. The Priority Tier isn't self-service. Customers have to sign a contract with Mistral's sales team. Third-party models join the platform Mistral is also opening its platform to open models from other providers. First up is GLM-5.2 from Chinese AI company Z.ai, which runs under the same regional rules and guarantees as Mistral's own models. To fund the compute capacity this requires, Mistral is collecting multi-year purchase commitments from large customers, packaged as European Compute Units. The idea is that building new data centers in Europe only pencils out if enough companies commit long-term. Mistral is a member of the Open Secure AI Alliance and Nvidia's Nemotron coalition. The company sees hosting third-party model weights on its platform as a natural extension of that work. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now --> Read on for the full picture. Subscribe for hype-free coverage. Full access to every article on THE DECODER No ads Join the comments and community discussions A weekly AI news recap via mail 6x/year: "AI Radar" — deep dives on the AI topics that matter most Daily AI news, always up to date Our full ten-year archive Covered by a team with 10+ years in AI Subscribe to The Decoder -->