메뉴
HN
Hacker News 7일 전

구글 제미나이 3.6 플래시 등 신규 AI 모델 3종 공개

IMP
8/10
핵심 요약

구글이 대규모 에이전트(Agent) 구축을 위한 효율성과 속도, 안정성을 갖춘 제미나이 3.6 플래시, 3.5 플래시-라이트, 3.5 플래시 사이버 등 3종의 새로운 AI 모델을 공개했습니다. 특히 3.6 플래시는 토큰 효율성과 코드 성능을 대폭 개선하여 비용을 절감했으며, 3.5 플래시-라이트는 초고속 처리에, 3.5 플래시 사이버는 보안 분야에 특화되어 실무자들의 AI 도입 및 운용 효율을 극대화할 수 있게 되었습니다.

번역된 본문

Gemini 3.6 Flash, 3.5 Flash-Lite, 3.5 Flash Cyber 공개

2026년 7월 21일 | x.com Facebook LinkedIn Mail Copy link 최신 Gemini 모델들은 대규모 AI 에이전트(Agent)를 구축하는 데 필요한 효율성, 낮은 지연 시간, 그리고 안정성을 제공합니다. 제품 관리 총괄 디렉터 Tulsee Doshi, Gemini 팀 대표

본 콘텐츠는 Google AI에 의해 생성되었습니다. 생성형 AI는 실험적인 기능입니다. (음성 듣기 속도 설정: 0.75X 1X 1.5X 2X)

상용 AI 에이전트를 구축하는 개발자와 고객들은 더 높은 토큰 효율성, 더 낮은 지연 시간, 그리고 더 안정적인 성능을 필요로 합니다. 저희의 Flash 모델 시리즈는 에이전트 워크플로우의 확장을 가능하게 하는 효율성과 품질의 최적점을 맞추도록 설계되었습니다. 기존 Gemini 3.5 Flash의 성공에 힘입어 다음과 같은 새로운 Gemini 모델을 소개합니다:

  • 3.6 Flash: 더 나은 코딩, 지식 노동, 그리고 멀티모달 성능을 제공하는 핵심 워크호스(주력) 모델입니다. Artificial Analysis Index에 따르면 3.5 Flash와 비교하여 출력 토큰 사용량을 17% 줄였으며, Datacurve의 DeepSWE와 같은 일부 벤치마크에서는 최대 65%까지 감소시켰습니다. 이 모든 것이 출력 토큰당 더 낮은 비용으로 이루어집니다.
  • 3.5 Flash-Lite: Artificial Analysis Index 기준으로 초당 350개의 출력 토큰을 생성하는 가장 빠르고 비용 효율적인 3.5 등급 모델입니다. 또한 에이전트 워크플로우에서 이전 세대 Flash-Lite 모델보다 월등히 뛰어난 성능을 보여줍니다.
  • 3.5 Flash Cyber in CodeMender: 성공적인 사이버 보안 애플리케이션은 에이전트 인프라와 함께 모델을 신중하게 조율하는 것을 필요로 합니다. 저희는 최고 수준의 경쟁력 있는 성능을 제공하는 새롭고 매우 효율적이며 보안에 특화된 모델과 CodeMender 코드 보안 에이전트의 조합을 소개하게 되었습니다.

오늘 발표된 모델 외에도, Gemini 3.5 Pro는 현재 파트너들과 테스트 중이며 준비가 완료되는 대로 광범위하게 공개할 계획입니다. 이와 병행하여 저희 팀은 이미 차세대 모델 구축에 집중하고 있습니다. 저희는 Gemini 4를 위한 가장 야심 찬 사전 학습을 시작했으며, 이 과정에서 큰 진전을 보이고 있어 매우 기쁩니다.

3.6 Flash: 3.5 Flash보다 더 효율적이고 향상된 품질

Gemini 3.6 Flash는 3.5 Flash에 대한 개발자 및 고객의 피드백을 직접적으로 반영하여 구축되었습니다. 3.6 Flash는 코딩 및 지식 노동에서 한 단계 발전한 성능을 제공할 뿐만 아니라, 이를 수행하면서도 토큰 효율성을 의미 있게 개선했습니다. 예를 들어, Artificial Analysis Index에 따르면 3.6 Flash는 3.5 Flash보다 17% 적은 출력 토큰을 소비합니다. 또한 다단계 워크플로우를 완료하기 위해 더 적은 추론 단계와 도구 호출(tool calls)을 필요로 합니다. 이러한 향상된 효율성은 3.5 Flash보다 낮은 가격과 결합됩니다. 입력 토큰 1백만 개당 $1.50, 출력 토큰 1백만 개당 $7.50의 가격으로 책정되어, 3.6 Flash는 에이전트 작업당 전체 비용을 줄여 에이전트를 구축하고 실행하는 데 더 많은 비용 효율성을 제공합니다.

(OSWorld 검증 작업(API)에서 3.6 Flash는 3.5 Flash보다 더 나은 토큰 효율성을 보여주며 불필요하게 긴 텍스트 생성이 줄어들었습니다.)

더 효율적으로 변했음에도 불구하고, 3.6 Flash는 다양한 사용 사례에서 3.5 Flash와 비교하여 성능 향상을 보여줍니다:

  • DeepSWE(49% vs. 37%)에서 볼 수 있듯, 원치 않는 코드 편집을 줄이고 실행 루프를 감소시켜 더 높은 정밀도를 제공합니다.
  • MLE Bench(63.9% vs. 49.7%)에서 입증된 바와 같이 기계 학습 연구(ML Research) 분야에서 큰 성능 향상을 보여줍니다.
  • OSWorld-Verified(83.0% vs. 78.4%)에서 확인된 바와 같이 컴퓨터 사용 기능이 향상되었습니다. 이제 컴퓨터 사용 기능은 Gemini API 및 Gemini Enterprise를 통해 기본 제공되는 클라이언트 측 도구가 되었습니다.
  • GDPval-AA v2(1421 vs. 1349)와 같은 벤치마크에서 나타나듯 지식 노동 분야에서 3.5 Flash를 능가합니다.

Hebbia 및 Harvey와 같은 고객들은 이 모델이 문서 파싱, 차트 및 데이터 분석, 보고서 초안 작성과 같은 멀티모달 작업에서 특히 뛰어난 능력을 발휘한다는 것을 확인했습니다. AIS의 관리형 에이전트(Managed Agents)를 사용하는 3.6 Flash는 3.5 Flash(AIS)보다 더 효율적이고 정확하게 금융 데이터 및 기록을 파싱하고 분석하는 데 도움을 줄 수 있습니다. 또한 3.6 Flash는 AGY에서 다중 에이전트 조정(multi-agent orchestration)을 사용하여 코드 마이그레이션을 더 낮은 지연 시간으로 실행합니다.

원문 보기
원문 보기 (영어)
Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber Jul 21, 2026 | x.com Facebook LinkedIn Mail Copy link Our newest Gemini models deliver the efficiency, latency, and reliability to build AI agents at scale. Tulsee Doshi Senior Director, Product Management, on behalf of the Gemini team Share x.com Facebook LinkedIn Mail Copy link . Inlining them here makes them available in the DOM for the page. --> Your browser does not support the audio element. Listen to article [[duration]] minutes This content is generated by Google AI. Generative AI is experimental Voice Speed Voice Speed 0.75X 1X 1.5X 2X Developers and customers building production AI agents need higher token efficiency, lower latency, and more reliable performance. Our Flash series of models is built to meet the sweet spot of efficiency and quality to enable scaling agentic workflows. Building on Gemini 3.5 Flash, we’re introducing new Gemini models: 3.6 Flash: Our workhorse model that delivers better coding, knowledge work, and multimodal performance. According to the Artificial Analysis Index , it reduces output token usage by 17% compared to 3.5 Flash, and in some benchmarks like DeepSWE by Datacurve , we observe up to 65%, all at a lower cost per output token. 3.5 Flash-Lite: Our fastest, most cost-effective 3.5-class model, delivering 350 output tokens per second according to the Artificial Analysis Index, also significantly outperforming prior Flash-Lite generations in agentic workflows. 3.5 Flash Cyber in CodeMender: Successful cybersecurity applications require careful orchestration of a model alongside an agent infrastructure. We’re introducing a combination of a new, highly efficient, specialized cyber-focused model paired with our CodeMender code security agent that delivers competitive performance at the frontier. Beyond today’s releases, Gemini 3.5 Pro is currently testing with partners and we plan to make it broadly available as soon as it’s ready. In parallel, our team is already focusing on building the next generation of models. We have started our most ambitious pre-training run yet, for Gemini 4, and are excited by the progress. 3.6 Flash: More efficient and better quality than 3.5 Flash Gemini 3.6 Flash builds directly on developer and customer feedback from 3.5 Flash. 3.6 Flash not only delivers a step up in coding and knowledge work, but it does this while meaningfully improving token efficiency. For example, on the Artificial Analysis Index, we see 3.6 Flash consuming 17% fewer output tokens than 3.5 Flash. It also takes fewer reasoning steps and tool calls to accomplish multi-step workflows. This enhanced efficiency is also combined with a lower price than 3.5 Flash. At $1.50/1M input tokens and $7.50/1M output tokens, 3.6 Flash reduces the overall cost per agentic task, making agents more cost-effective to build and run. 3.6 Flash shows better token efficiency and reduced verbosity than 3.5 Flash in an OSWorld verified task (API) Even while being more efficient, 3.6 Flash sees performance gains compared to 3.5 Flash across use cases: 3.6 Flash delivers higher precision with fewer unwanted code edits and reduced execution loops, as seen in DeepSWE (49% vs. 37%), and shows significant improvement in ML Research, as seen in MLE Bench (63.9% vs. 49.7%). It has improved computer use capabilities as seen in OSWorld-Verified (83.0% vs. 78.4%). Computer use is now a built-in client side tool via the Gemini API and Gemini Enterprise. It outperforms 3.5 Flash in knowledge work, as shown by benchmarks like GDPval-AA v2 (1421 vs. 1349). Customers like Hebbia and Harvey have found it particularly capable at multimodal tasks like document parsing, chart and data analysis, and report drafting. 3.6 Flash, using Managed Agents on AIS, can help parse through and analyze financial data and transcripts more efficiently and accurately than 3.5 Flash (AIS) 3.6 Flash executes code migrations,using multi-agent orchestration on AGY, with lower latency and higher quality than 3.5 Flash (AGY) 3.6 Flash helps develop a photographic texture extractor for 3D workflows, using canvas (Gemini App) Customers report 3.6 Flash is a step forward in both cost and quality, balancing token efficiency, accuracy, and speed across complex workflows and knowledge-based tasks: Built with safety 3.6 Flash is shipping with enhanced Frontier Safety safeguards in the domains of Chemical, Biological, Radiological, and Nuclear (CBRN) and cyber offense misuses. These safeguards make the model substantially more resistant to jailbreaks. At the same time, the model has been trained to minimize refusals for beneficial uses. For more information, see the 3.6 Flash model card. 3.5 Flash-Lite: Built to scale agentic workflows Beyond Flash, we’re also releasing Gemini 3.5 Flash-Lite, designed for both low-latency tasks and tasks where high throughput is critical for developers workflows, like agentic search and document processing. 3.5 Flash-Lite is the fastest model in the 3.5 series. As measured by Artificial Analysis , it runs at 350 output tokens/s. Priced at $0.3/1M input tokens and $2.5/1M output tokens and with significantly better quality than 3.1 Flash-Lite, 3.5 Flash-Lite offers a strong price-to-performance ratio for developers and customers running high throughput production traffic. 3.5 Flash-Lite executes high volume tasks at a lower latency than 3.5 Flash. 3.5 Flash-Lite enables efficient scaling for agentic systems. Across thinking levels, the model significantly outperforms 3.1 Flash-Lite. Depending on the workload, developers can configure the model to prioritize low-latency, low-cost execution for high-volume tasks with the minimal and low thinking levels, or engage higher thinking levels to process multi-step subagent workloads. The model now also has computer use as a built-in tool to reliably support these agentic tasks across surfaces. It’s a significant step up in coding and agentic tasks as seen in Terminal-Bench 2.1 (54% vs 31%), long context as seen in GDM-MRCR v2 (72.2% vs. 60.1%), and real-world task execution as seen in GDPval-AA v2 (1140 vs. 642). In fact, on many agentic and coding evals, 3.5 Flash-Lite even outperforms 3 Flash, including on SWE-Bench Pro (54.2% vs. 49.6%) and OSWorld-Verified (74.0% vs. 65.1%), making it a faster & more capable option for workloads on both 2.5 and 3 Flash. 3.5 Flash-Lite extracts product features from a massive e-commerce dataset and synthesizes it. Working alongside 3.6 Flash as the master agent, 3.5 Flash-Lite instantly generates 25 unique, ready-to-explore web design concepts. 3.5 Flash-Lite can scale receipt translation and summarization with its multimodal understanding. 3.5 Flash-Lite builds a game by instantly generating and iterating through multiple options. Early customers of 3.5 Flash-Lite are highlighting its unique combination of speed, intelligence, and cost efficiency for scaling agentic workflows and data processing tasks: For more information about the model, see the 3.5 Flash-Lite model card. 3.5 Flash Cyber in CodeMender: finding and fixing vulnerabilities efficiently AI models have become capable of finding security vulnerabilities faster than current systems can fix them. Tackling this growing threat requires an approach to securing software that is highly capable and efficient. Flash’s performance and efficiency makes it an ideal foundation to detect, validate, and patch code security issues at scale. Gemini 3.5 Flash Cyber is built on top of 3.5 Flash, and fine-tuned for finding and fixing cybersecurity vulnerabilities at a lower price per token than larger models. Within CodeMender, which uses multiple 3.5 Flash Cyber agents working together to produce a single combined report, 3.5 Flash Cyber reaches competitive performance at the frontier on the popular benchmark CyberGym. Given the dual-use nature of this technology, we have taken an intentional approach to deploying 3.5 Flash Cyber. The model will be exclusively ava