메뉴
HN
Hacker News • 23일 전

구글, 젬미니 3.8 플래시 모델 카드 공개

IMP
7/10
핵심 요약

구글 딥마인드가 2026년 9월 젬미니 3.7 플래시를 기반으로 한 젬미니 3.8 플래시 모델 카드를 공개했습니다. 소프트웨어 엔지니어링과 에이전트 기반 지식 작업 흐름 전반에서 성능이 향상되었으며, 품질·비용·지연 시간을 조절할 수 있는 노력 수준(effort level) 설정을 지원합니다. 최대 100만 토큰 컨텍스트 창과 6.4만 토큰 출력을 제공하며, 프로덕션급 에이전트의 비용 효율적 확장에 적합합니다.

번역된 본문

본문으로 건너뛰기 · 2026년 9월 2일 게시 · Gemini 3.8 Flash · PDF 버전 보기

모델 카드는 알려진 한계, 완화 접근 방식, 안전성 성과 등 젬미니 모델의 핵심 정보를 제공하기 위한 것입니다. 모델 카드는 모델이 개선·수정됨에 따라 업데이트된 평가 결과를 반영하는 등 수시로 갱신될 수 있습니다. 전체 모델 카드 목록은 Google DeepMind 사이트에서 확인할 수 있습니다.

게시 시기: 2026년 9월

목차: 모델 정보 / 모델 데이터 / 구현 및 지속가능성 / 배포 / 평가 / 의도된 용도 및 한계 / 윤리 및 콘텐츠 안전

모델 정보

설명: 젬미니 3.8 플래시는 젬미니 3 모델 패밀리의 차세대 버전으로, 젬미니 3.7 플래시를 기반으로 하며 소프트웨어 엔지니어링 및 에이전트 기반 지식 작업 흐름 전반에서 성능이 향상되었습니다. 품질, 비용, 지연 시간의 균형을 조절할 수 있는 사용자 지정 노력 수준(effort level) 기능을 계속 지원합니다.

모델 의존성: 젬미니 3.8 플래시는 젬미니 3.7 플래시를 기반으로 합니다.

입력: 텍스트 문자열(질문, 프롬프트, 요약할 문서 등), 이미지, 오디오, 비디오 파일. 최대 100만 토큰의 컨텍스트 창을 지원합니다.

출력: 텍스트, 최대 6.4만 토큰 출력.

아키텍처: 젬미니 3.8 플래시는 젬미니 3.7 플래시를 기반으로 합니다. 자세한 아키텍처 정보는 젬미니 3.7 플래시 모델 카드를 참고하세요.

모델 데이터

학습 데이터셋: 젬미니 3.8 플래시는 젬미니 3.7 플래시를 기반으로 합니다. 학습 데이터셋에 대한 자세한 내용은 젬미니 3.7 플래시 모델 카드를 참고하세요.

학습 데이터 처리: 학습 데이터 처리 관련 자세한 내용은 젬미니 3.7 플래시 모델 카드를 참고하세요.

구현 및 지속가능성

하드웨어: 젬미니 3.8 플래시는 젬미니 3.7 플래시를 기반으로 합니다. 하드웨어 및 지속가능한 운영에 대한 지속적인 노력은 젬미니 3.7 플래시 모델 카드를 참고하세요.

소프트웨어: 소프트웨어 관련 자세한 내용은 젬미니 3.7 플래시 모델 카드를 참고하세요.

배포

젬미니 3.8 플래시는 다음 채널을 통해 배포되며, 각 문서는 해당 링크에서 확인할 수 있습니다:

  • 젬미니 앱
  • Gemini Enterprise Agent Platform
  • Google AI Studio
  • Gemini API
  • Google AI Mode
  • Google Antigravity

본 모델은 API(애플리케이션 프로그래밍 인터페이스)를 통해 다운스트림 제공자에게 제공되며 관련 이용약관이 적용됩니다. 모델 사용에 별도의 하드웨어나 소프트웨어는 필요하지 않습니다. AI Studio 및 Gemini API는 Gemini API 추가 서비스 약관을, Gemini Enterprise Agent Platform은 Google Cloud Platform 서비스 약관을 참고하세요. 자세한 내용은 젬미니 모델 API 사용 안내 및 Gemini API 퀵스타트를 참고하세요.

평가

접근 방식: 젬미니 3.8 플래시는 코딩, 지식 작업, 멀티모달 기능, 롱 컨텍스트(long-context), 컴퓨터 사용(computer use), 과학적 추론을 포함한 다양한 벤치마크에서 평가되었습니다. 추가 벤치마크 및 접근 방식, 결과, 방법론에 대한 세부 사항은 deepmind.com/models/evals-methodology/gemini-3-8-flash에서 확인할 수 있습니다.

결과: 2026년 9월 기준 결과는 아래와 같습니다. 평가 방법론에 대한 자세한 내용은 deepmind.com/models/evals-methodology/gemini-3-8-flash를 참고하세요.

의도된 용도 및 한계

장점 및 의도된 용도: 젬미니 3.8 플래시는 일반 사용자, 개발자, 기업에 적합하며, 범용 프로덕션급 에이전트의 비용 효율적 확장을 위해 설계되었습니다. 주요 활용 사례로는 소프트웨어 엔지니어링, 에이전트 작업, 복잡한 지식 작업 흐름 등이 있습니다.

알려진 한계: 젬미니 3.8 플래시는 환각(hallucination)과 같은 기초 모델(foundation model)의 일반적인 한계를 보일 수 있습니다. 이 외에도 탈옥(jailbreak) 저항 성능 개선을 지속적으로 진행 중이며, 최근 프런티어 세이프티(Frontier Safety) 전반의 완화 조치를 강화했습니다. 간헐적인 속도 저하나 타임아웃 문제가 발생할 수 있으며, 특히 높은 노력 수준에서 성능을 극대화하기 위해 더 많은 토큰을 사용할 수 있습니다. 지식 컷오프 날짜...(이하 생략)

원문 보기
원문 보기 (영어)
Skip to main content Published 2 September 2026 Gemini 3.8 Flash View PDF version Model Cards are intended to provide essential information on Gemini models, including known limitations, mitigation approaches, and safety performance. Model cards may be updated from time to time; for example, to include updated evaluations as the model is improved or revised. See the Google DeepMind site for a comprehensive list of model cards. Published: September, 2026 Model Information Model Data Implementation and Sustainability Distribution Evaluation Intended Usage and Limitations Ethics and Content Safety Model Information Description Gemini 3.8 Flash is the next iteration in the Gemini 3 model family, building on Gemini 3.7 Flash, delivering performance advancements across software engineering and agentic knowledge workflows. It continues to support customizable effort levels to control the mix of quality, cost and latency. Model dependencies Gemini 3.8 Flash is based on Gemini 3.7 Flash. Inputs Text strings (e.g., a question, a prompt, document(s) to be summarized), images, audio, and video files, with a token context window of up to 1M. Outputs Text, with a 64K token output. Architecture Gemini 3.8 Flash is based on Gemini 3.7 Flash. For more information about the model architecture for Gemini 3.8 Flash, see the Gemini 3.7 Flash model card . Model Data Training Dataset Gemini 3.8 Flash is based on Gemini 3.7 Flash. For more information about the training dataset for Gemini 3.8 Flash, see the Gemini 3.7 Flash model card . Training Data Processing For more information about the training data processing for Gemini 3.8 Flash, see the Gemini 3.7 Flash model card . Implementation and Sustainability Hardware Gemini 3.8 Flash is based on Gemini 3.7 Flash. For more information about the hardware for Gemini 3.8 Flash and our continued commitment to operate sustainably , see the Gemini 3.7 Flash model card . Software Gemini 3.8 Flash is based on Gemini 3.7 Flash. For more information about the software for Gemini 3.8 Flash, Gemini 3.7 Flash model card . Distribution Gemini 3.8 Flash is distributed in the following channels; respective documentation shared in line: Gemini app Gemini Enterprise Agent Platform Google AI Studio Gemini API Google AI Mode Google Antigravity Our models are available to downstream providers via an application program interface (API) and subject to relevant terms of use. There is no required hardware or software to use the model. For AI Studio and Gemini API, see the Gemini API Additional Terms of Service ; for Gemini Enterprise Agent Platform, see Google Cloud Platform Terms of Service . For more information, see Gemini Model API instructions and Gemini API quickstart . Evaluation Approach Gemini 3.8 Flash was evaluated across a range of benchmarks, including coding, knowledge work, multimodal capabilities, long-context, computer use, scientific reasoning. Additional benchmarks and details on approach, results and their methodologies can be found at: deepmind.com/models/evals-methodology/gemini-3-8-flash . Results Results as of September, 2026 are listed below: For details on our evaluation methodology please see: deepmind.com/models/evals-methodology/gemini-3-8-flash Intended Usage and Limitations Benefit and Intended Usage Gemini 3.8 Flash is well-suited for users, developers, and enterprises, designed for cost-effective scaling of general-purpose, production-ready agents. Some use cases include: software engineering, agent tasks, and complex knowledge workflows. Known Limitations Gemini 3.8 Flash may exhibit some of the general limitations of foundation models, such as hallucinations. In addition to this, we are continually working to improve jailbreak resistance and have recently strengthened the mitigations across Frontier Safety. There may also be occasional slowness or timeout issues. At times, the model might use more tokens to maximize performance, especially at higher effort levels. The knowledge cutoff date for Gemini 3.8 Flash is March 2026 – users can expect updated information for some domains while in others they may experience the model’s knowledge is limited to January 2025 (in line with the Gemini 3 Model Family). For more information about known limitations, see the Gemini 3.7 Flash model card . Acceptable Usage For more information about the acceptable usage for Gemini 3.7 Flash, see the Gemini 3.7 Flash model card . Ethics and Content Safety Evaluation Approach For more information about the evaluation approach for Gemini 3.8 Flash, see the Gemini 3.7 Flash model card . Safety Policies For more information about the safety policies for Gemini 3.8 Flash, see the Gemini 3.7 Flash model card . Training and Development Evaluation Results Results for some of the internal safety evaluations conducted during the development phase are listed below. The evaluation results are for automated evaluations and not human evaluation or red teaming. Scores are provided as an absolute percentage increase or decrease in performance compared to the indicated model, as described below. Overall, Gemini 3.8 Flash performs similarly to Gemini 3.7 Flash across both safety and tone, with low unjustified refusals. Safety performance across non-English languages regressed slightly relative to 3.7 Flash. Evaluation Description Gemini 3.8 Flash vs. Gemini 3.7 Flash Text to Text Safety Automated content safety evaluation measuring safety policies -0.4pp Lower is better Multilingual Safety Automated safety policy evaluation across multiple languages +5.4pp Lower is better Image to Text Safety Automated content safety evaluation measuring safety policies 0.0pp Lower is better Tone 1 Automated evaluation measuring objective tone of model responses +0.2pp Higher is better Unjustified-refusals Automated evaluation measuring model’s ability to respond to borderline prompts while remaining safe +1.1pp Lower is better 1 For tone and instruction following, a positive percentage increase represents an improvement in the tone of the model on sensitive topics and the model’s ability to follow instructions while remaining safe compared to Gemini 3 Flash. We mark improvements in green and regressions in red. We continue to improve our internal evaluations, including refining automated evaluations to reduce false positives and negatives, as well as update query sets to ensure balance and maintain a high standard of results. The performance results reported below are computed with improved evaluations and thus are not directly comparable with performance results found in previous Gemini model cards. We expect variation in our automated safety evaluations results, which is why we review flagged content to check for egregious or dangerous material. Our manual review confirmed losses were overwhelmingly either a) false positives or b) not egregious. Human Red Teaming Results We conduct manual red teaming by specialist teams who sit outside of the model development team. High-level findings are fed back to the model team. For child safety evaluations, Gemini 3.8 Flash satisfied required launch thresholds, which were developed by expert teams to protect children online and meet Google’s commitments to child safety across our models and Google products. For content safety policies generally, including child safety, we saw similar or improved safety performance compared to Gemini 3.7 Flash. Additionally, the scope of red teaming covered potential issues outside of our strict policies, compared performance to Gemini 3.1 Pro, and found no egregious concerns. Frontier Safety Assessment Gemini 3.8 Flash is part of the Gemini 3 series of models. We evaluated Gemini 3.7 Flash as outlined in our latest Frontier Safety Framework (April-2026), and found that it did not reach any Tracked or Critical Capability Levels (T/CCLs). Our assessments have shown that Gemini 3.8 Flash does not have meaningful new capabilities or material increases in performance with respect to the domains out