메뉴
HN
Hacker News • 23일 전

구글 제미나이 3.8 플래시 및 사이버 보안 버전 공개

IMP
8/10
핵심 요약

구글이 6주 만에 세 번째 플래시 모델인 제미나이 3.8 플래시를 공개했습니다. 3.7 플래시와 동일한 속도와 저렴한 가격($0.75/백만 입력 토큰)으로 소프트웨어 엔지니어링, 에이전트 작업, 다단계 추론에서 대형 프론티어 모델에 근접하는 성능을 제공합니다. 사이버보안 특화 버전인 3.8 Flash Cyber는 취약점 탐지와 자동 패치에서 최고 수준 성능을 보이며, 신뢰된 방어자들에게만 'Fairwind 프로그램'을 통해 제공됩니다.

번역된 본문

제미나이 3.8 플래시 및 3.8 플래시 사이버 소개 (2026년 9월 2일)

3주 전 출시된 3.7 플래시의 흐름을 이어받아 불과 6주 만에 세 번째 플래시 출시로, 오늘 우리는 지금까지 중 가장 뛰어난 추론 및 코딩 모델인 제미나이 3.8을 3.7과 동일한 속도와 저렴한 비용으로 선보입니다.

제미나이 3.8은 두 가지 변형으로 출시됩니다:

  • 제미나이 3.8 플래시: 가장 지능적인 실용형(workhorse) 모델로, 소프트웨어 엔지니어링, 에이전트 작업, 전문 분야의 중요한 다단계 추론 전반에서 3.7 플래시 대비 큰 향상을 보여줍니다. 3.7 플래시와 동일한 출시가로 백만 입력 토큰당 $0.75, 백만 출력 토큰당 $3.75에 이용할 수 있습니다.

  • 제미나이 3.8 플래시 사이버: 취약점 탐지 및 자동 패치에서 프론티어 수준의 성능을 갖춘 가장 강력한 사이버보안 모델로, 새로운 Fairwind 프로그램을 통해 신뢰된 방어자들에게 제공됩니다.

두 출시작은 서로 다른 배포 환경에 맞게 조정되었지만 동일한 기반 지능으로 구동되며, 기반 모델을 재귀적으로 평가하고 개선하도록 설계된 장기 실행 에이전트 루프로 더욱 가속화되었습니다. 이 공유 핵심의 코딩 및 추론 향상은 고도로 까다로운 사이버보안 분야에서의 엄격한 훈련 등 여러 혁신에서 비롯되었습니다.

제미나이 3.8 플래시: 장기적 코딩과 자율 에이전트를 위한 모델

제미나이 3.8 플래시는 3.7 플래시 대비 상당한 향상을 이루어, 종종 더 비싼 프론티어 모델에 근접하는 성능을 보입니다. DeepSWE v1.1(장기 소프트웨어 엔지니어링) 벤치마크에서 3.8 플래시는 복잡한 엔지니어링 문제를 종단간 자율적으로 해결하는 능력에서 대부분의 더 큰 프론티어 모델을 능가하며, 비용은 그 일부에 불과합니다. 또한 전문 지식 분야 전반에서 기업의 중요한 자율 작업에 필요한 신뢰성을 보여줍니다.

고급 분석과 보고가 필요한 정량·전문 분야에서 3.8 플래시는 Vals Finance Agent V2, Harvey의 법률 에이전트 벤치마크 등에서 3.7 플래시와 다른 프론티어 모델을 능가합니다. 또한 HLE-Verified에서 54.9%를 달성하며 STEM, 인문학, 전문 분야 전반의 다단계 추론 능력을 입증했습니다.

이러한 성능 향상은 핵심 설계 선택에서 비롯됩니다: 3.8 플래시는 더 열심히 작동합니다. 복잡한 작업에서 더 큰 근면함을 보이며, 추가 추론 단계를 실행하고 도구를 반복적으로 호출합니다. 때로는 성능을 극대화하기 위해 더 많은 토큰을 사용할 수 있으며, 특히 노력(effort) 수준이 높을 때 그렇습니다. 컴퓨팅 효율이 최우선 제약인 애플리케이션에서는 개발자가 낮은 노력 수준을 활용해 토큰 오버헤드를 최소화하거나, 효율 우선 워크로드를 위해 완전히 지원되는 제미나이 3.7 플래시를 계속 사용할 수 있습니다.

데모 사례:

  • 구글 안티그래비티(Antigravity)에서 루프 지시문이 포함된 단순한 프롬프트로 게임을 제작했습니다. 퍼즐, 환경 스토리텔링, 나노 바나나(Nano Banana)로 생성한 텍스처를 활용해 성을 탐험하는 마법사를 플레이하는 몰입형 3D 레벨을 구현했습니다.
  • 단일 프롬프트로 완전히 작동하는 DOS 버전 구글 지도를 만들었으며, 위치, 길찾기, 스트리트뷰를 모두 사용할 수 있습니다.
  • 미국 지질조사국(USGS)의 실제 데이터셋을 활용해 유명 지리 명소의 지형도에서 실시간 단면, 2D 투영, 과학적 설명을 탐색할 수 있습니다.
  • '하드웨어 해부(Hardware Anatomy)'는 구글 AI 스튜디오에서 제작된 인터랙티브 3D 시각화 도구로, 하드웨어 기기의 실제 비율 분해도를 사실적인 Three.js 렌더링으로 생성하며, 기기를 자동으로 레이어로 분해해 펼쳐볼 수 있습니다.
원문 보기
원문 보기 (영어)
Introducing Gemini 3.8 Flash and 3.8 Flash Cyber Sep 02, 2026 | x.com Facebook LinkedIn Mail Copy link Our newest Gemini models deliver next-generation intelligence for agentic workflows and cybersecurity. Tulsee Doshi Senior Director, Product Management Raluca Ada Popa Gemini Security Lead, Google DeepMind Share x.com Facebook LinkedIn Mail Copy link Building on the momentum of 3.7 Flash from three weeks ago and marking our third Flash release in only six weeks, today we’re introducing Gemini 3.8, our best reasoning & coding model yet, at the same speed and low cost of 3.7. Gemini 3.8 introduces 2 variants: Gemini 3.8 Flash: our most intelligent workhorse model, delivering significant improvements from 3.7 Flash across software engineering, agentic tasks, and critical, multi-step reasoning in specialized domains. It is available at the same introductory price 1 as 3.7 Flash at $0.75 per million input tokens and $3.75 per million output tokens. Gemini 3.8 Flash Cyber: our most capable cybersecurity model with frontier-level performance in vulnerability detection and automated patching, available to trusted defenders through our new Fairwind Program . While tailored for different deployment environments, both of today's releases are powered by the same foundational intelligence, and further accelerated by long-running agentic loops designed to recursively evaluate and refine the underlying models. The significant coding and reasoning gains across this shared core were driven by a number of innovations, including rigorous training in the highly demanding domain of cybersecurity. Gemini 3.8 Flash: built for long-horizon coding and autonomous agents Gemini 3.8 Flash delivers substantial gains from 3.7 Flash, often approaching the performance of higher-cost frontier models. On DeepSWE v1.1 (Long-Horizon Software Engineering ) 3.8 Flash outperforms most larger frontier models in autonomously solving complex engineering problems end to end, only at a fraction of the cost. Additionally, 3.8 Flash exhibits the dependability required for critical enterprise autonomy, across specialized knowledge domains . In quantitative and professional fields that require advanced analysis and reporting, 3.8 Flash outperforms 3.7 Flash and other frontier models in benchmarks like Vals Finance Agent V2 and Harvey's Legal Agent Benchmark . 3.8 Flash also achieves a 54.9% on HLE-Verified, demonstrating its ability to handle multi-step reasoning across STEM, humanities, and professional fields. These performance gains stem from a core design choice: 3.8 Flash works harder. On complex tasks, it exhibits greater diligence — executing extra reasoning steps, and calling tools iteratively. At times, the model might use more tokens to maximize performance, especially at higher effort levels. For applications where compute efficiency is the primary constraint, developers can utilize lower effort levels to minimize token overhead or continue to rely on Gemini 3.7 Flash, which remains fully supported for efficiency-first workloads. Gemini 3.8 Flash built this game with a simple prompt using a looping instruction in Google Antigravity. The game uses puzzles, environmental storytelling, and textures generated with Nano Banana to create an immersive 3D level in which you play a wizard navigating a castle. Gemini 3.8 Flash builds a fully functional DOS version of Google Maps in a single prompt in Google Antigravity that is fully playable with locations, directions, and Street View. Explore realtime cross-sections, 2D projections, scientific explanations in a topographic map of famous geographical sites built with Gemini 3.8 Flash in Google Antigravity using real datasets from the U.S. Geological Survey. Hardware Anatomy is an interactive 3D visualizer built with Gemini 3.8 Flash in Google AI Studio that generates realistic Three.js renderings of physically-proportioned teardowns for hardware devices. It automatically decomposes devices into layers you can explode and inspect with a deconstruction slider. Gemini 3.8 Flash Cyber: expert cyber performance Gemini 3.8 Flash Cyber, available to a set of trusted defenders via the Fairwind Program , provides a decisive advantage in today’s complex cybersecurity landscape, with the Flash speed and cost that enables quick iteration. Autonomous vulnerability discovery On the standard industry benchmark for finding vulnerabilities, CyberGym, Gemini 3.8 Flash Cyber demonstrates frontier-level performance in autonomous vulnerability discovery. It surpasses both 3.5 Flash Cyber as well as significantly larger frontier models. To better capture real-world defensive needs which are not limited to just C/C++ codebases like in CyberGym, we also evaluated Gemini 3.8 Flash Cyber against a comprehensive internal benchmark in which the model has to discover a wide range of vulnerabilities across complex codebases spanning 20 programming languages. Here, the model showcases an impressive leap over our previous models and reaches a success rate exceeding 70%. Automated patching With Gemini 3.8 Flash Cyber, we focused specifically on equipping defenders with expert capabilities that give them an advantage over attackers. This is why we have invested in vulnerability fixing from the start, and prioritized it over offensive capabilities like exploitation. CWE-Bench , run by Collinear, is a challenging external benchmark for patching capabilities. On this benchmark, Gemini 3.8 Flash Cyber is on the Pareto frontier: with a pass@1 of 47.2% compared to a leading frontier model at 47.8%, yet offered at a significantly lower cost. Real-world impact: securing Google’s code We’re already using Gemini 3.8 Flash Cyber to secure code across Google. For example: The Chrome Security team found that 3.8 Flash Cyber produced 2.6 times more correct patches to vulnerabilities in Chrome than the best commercial models that are much larger. Wiz found that Gemini 3.8 Flash Cyber achieves +7.5-9.7% higher recall on their internal penetration testing benchmark for a 2.3-5.2x lower cost compared to other leading frontier models. Google’s Cloud Vulnerability Research team leveraged the 3.8 Flash Cyber model to find a critical foundational vulnerability in less than 2 hours, a vulnerability for which research and discovery usually takes months. What our Fairwind Program partners are saying Built with safety in mind 3.8 Flash ships with safeguards against misuse in the domains of Chemical, Biological, Radiological, and Nuclear (CBRN) and cyber offense, while enabling beneficial use cases, as per our Frontier Safety Framework . 3.8 Flash Cyber ships with a more permissive set of mitigations for cybersecurity, and as such, is only available to trusted defenders who require a more comprehensive set of cyber capabilities. Gemini 3.8 models have also made a significant leap in prompt injection robustness as measured by Gray Swan, protecting Gemini model users from prompt-injection related malicious attacks. Gemini 3.8 Flash and Cyber: get started today Developers : Build with 3.8 Flash and explore agent-first workflows in Google Antigravity or start building today in the Gemini API via Google AI Studio and Android Studio , or generate UIs in Stitch . Get started with our developer docs . Enterprises : Access 3.8 Flash in Gemini Enterprise. Consumers : 3.8 Flash is available to Google AI Pro and Ultra subscribers across the Gemini app , AI Mode in Google Search and Gemini in Google Sheets . Cyber: Through our new Fairwind Program , we’re providing trusted government authorities, as well as critical infrastructure operators and software maintainers with prioritized access to Gemini 3.8 Flash Cyber. Apply for access . Posted in: