Gemini API Safety Settings: A Comprehensive Guide to Content Moderation
In-depth discussion
Technical and instructional
0 0 5
本指南詳細介紹如何調整 Gemini API 的安全設定,以控制應用程式對騷擾、仇恨言論、煽情露骨內容和危險內容的處理方式。文章說明了內容安全篩選等級、每個要求的安全篩選機制,並提供了在 Google AI Studio 和程式碼中調整這些設定的具體步驟與範例。強調了根據使用情境調整安全設定的重要性,同時也提醒使用者注意服務條款的相關要求。
The Gemini API categorizes potentially unsafe content into four distinct harm categories, each with its own set of characteristics:
* **Harassment:** This category covers negative or harmful statements directed at individuals based on their identity or protected traits.
* **Hate Speech:** This encompasses content that is vulgar, disrespectful, or offensive in nature.
* **Sexually Explicit Content:** This includes discussions or references to sexual acts or other indecent matters.
* **Dangerous Content:** This category pertains to content that promotes, advocates for, or facilitates harmful actions.
These categories are defined within the `HarmCategory` enumeration. The ability to adjust filters for these categories allows developers to tailor the model's output to specific use cases. For instance, a game developer might choose to allow a higher tolerance for content flagged as 'dangerous' to better suit the game's narrative or theme. It's important to note that alongside these adjustable filters, the Gemini API incorporates built-in core safety protections, such as safeguards against child sexual abuse material, which are non-negotiable and always blocked by the system.
“ Content Safety Filtering Levels
Safety settings can be customized for each individual request made to the Gemini API. Upon receiving a request, the system analyzes the content and generates a safety score. This score represents the probability that the Gemini model identifies the content as belonging to a specific harm category. For example, if content is blocked due to a high probability of being classified as harassment, the returned safety score will indicate the category as `HARASSMENT` with a `HIGH` probability of harm. By default, the model's inherent safety features are active, and additional filters are disabled. If enabled, developers can set the system to block content based on its probability of being unsafe. The default model behavior is suitable for most applications, and adjustments should only be made when strictly necessary for the application's functionality. The table below outlines the adjustable blocking settings for each category. For instance, setting the 'Hate Speech' category's blocking threshold to 'Block only a few' will result in content with a high probability of being hate speech being blocked, while words with a lower probability may still be permitted. The available thresholds include: `OFF` (disables the safety filter), `BLOCK_NONE` (displays all content regardless of unsafe probability), `BLOCK_ONLY_HIGH` (blocks content with a high probability of being unsafe), `BLOCK_MEDIUM_AND_ABOVE` (blocks content with medium or higher probability of harm), and `BLOCK_LOW_AND_ABOVE` (blocks content with low, medium, or high probability of harm). `HARM_BLOCK_THRESHOLD_UNSPECIFIED` is used when no threshold is set, defaulting to the system's default blocking mechanism. For Gemini 2.5 and 3 models, blocking thresholds are disabled by default if not explicitly set. These settings can be configured for every request made to the generation service. Refer to the `HarmBlockThreshold` API reference for detailed information.
“ Interpreting Safety Feedback
Google AI Studio offers a user-friendly interface for adjusting safety settings. Within the 'Run settings' panel, navigate to 'Advanced settings' and then select 'Safety settings' to enter the 'Run safety settings' mode. Here, users can utilize sliders to modify the content filtering levels for each safety category. It is the developer's responsibility to ensure that the safety settings chosen for their intended use case comply with the Terms of Service. When a request is sent, such as a query to the model, and the content is blocked, a 'Content blocked' warning message will appear. To view more detailed information, hovering the cursor over the 'Content blocked' text will reveal the probability of the content belonging to specific categories and harm classifications.
“ Programmatic Adjustment of Safety Settings
To further enhance your understanding and implementation of safety features with the Gemini API, several resources are available. For a complete overview of the API's capabilities, refer to the API Reference. When developing with Large Language Models (LLMs), adhering to safety best practices is paramount; the Safety Guidelines provide essential information on this topic. For a deeper dive into evaluating probabilities and severity, the Jigsaw team's resources offer valuable insights. Additionally, products that aid in developing safety solutions, such as the Perspective API, can be explored. For instance, you can use these safety settings to build toxicity classifiers, and a classification example is available to get you started. Content on this page is licensed under Creative Commons Attribution 4.0 unless otherwise specified, and code examples are under the Apache 2.0 license. For more details, consult the Google Developers Site Policies. Java is a registered trademark of Oracle and/or its affiliates. The last update for this page was on June 1, 2026 (UTC).
We use cookies that are essential for our site to work. To improve our site, we would like to use additional cookies to help us understand how visitors use it, measure traffic to our site from social media platforms and to personalise your experience. Some of the cookies that we use are provided by third parties. To accept all cookies click ‘Accept’. To reject all optional cookies click ‘Reject’.
Comment(0)