Logo for AiToolGo

Gemini API Safety Settings: A Comprehensive Guide to Content Moderation

In-depth discussion
Technical and instructional
 0
 0
 5
本指南詳細介紹如何調整 Gemini API 的安全設定,以控制應用程式對騷擾、仇恨言論、煽情露骨內容和危險內容的處理方式。文章說明了內容安全篩選等級、每個要求的安全篩選機制,並提供了在 Google AI Studio 和程式碼中調整這些設定的具體步驟與範例。強調了根據使用情境調整安全設定的重要性,同時也提醒使用者注意服務條款的相關要求。
  • main points
  • unique insights
  • practical applications
  • key topics
  • key insights
  • learning outcomes
  • • main points

    • 1
      提供了 Gemini API 安全設定的全面概述,涵蓋了四個主要的危害類別。
    • 2
      詳細解釋了內容安全篩選等級(HIGH, MEDIUM, LOW, NEGLIGIBLE)及其在封鎖內容中的作用。
    • 3
      提供了在 Google AI Studio 和程式碼(Python 範例)中調整安全設定的具體操作指南。
  • • unique insights

    • 1
      強調了根據特定應用場景(如遊戲對白)調整安全設定的靈活性,同時保留了對兒童安全等核心危害內容的內建防護。
    • 2
      解釋了 API 封鎖內容的依據是內容不安全的機率,而非單純的嚴重程度,並以具體例子說明了這一點。
  • • practical applications

    • 該指南為開發者提供了在構建 Gemini API 應用時,如何有效管理內容安全性的實用方法,幫助他們平衡功能需求與用戶安全,並符合服務條款。
  • • key topics

    • 1
      Gemini API Safety Settings
    • 2
      Harm Category Thresholds
    • 3
      Content Moderation in AI
  • • key insights

    • 1
      Detailed explanation of how Gemini API's safety filters work based on probability.
    • 2
      Practical guidance on adjusting safety settings for specific use cases.
    • 3
      Code examples for programmatic control of safety configurations.
  • • learning outcomes

    • 1
      Understand the different types of harmful content Gemini API can filter.
    • 2
      Learn how to adjust safety settings in Google AI Studio and via API calls.
    • 3
      Be able to implement appropriate safety configurations for their specific AI applications.
examples
tutorials
code samples
visuals
fundamentals
advanced content
practical tips
best practices

“ Introduction to Gemini API Safety Settings

The Gemini API categorizes potentially unsafe content into four distinct harm categories, each with its own set of characteristics: * **Harassment:** This category covers negative or harmful statements directed at individuals based on their identity or protected traits. * **Hate Speech:** This encompasses content that is vulgar, disrespectful, or offensive in nature. * **Sexually Explicit Content:** This includes discussions or references to sexual acts or other indecent matters. * **Dangerous Content:** This category pertains to content that promotes, advocates for, or facilitates harmful actions. These categories are defined within the `HarmCategory` enumeration. The ability to adjust filters for these categories allows developers to tailor the model's output to specific use cases. For instance, a game developer might choose to allow a higher tolerance for content flagged as 'dangerous' to better suit the game's narrative or theme. It's important to note that alongside these adjustable filters, the Gemini API incorporates built-in core safety protections, such as safeguards against child sexual abuse material, which are non-negotiable and always blocked by the system.

“ Content Safety Filtering Levels

Safety settings can be customized for each individual request made to the Gemini API. Upon receiving a request, the system analyzes the content and generates a safety score. This score represents the probability that the Gemini model identifies the content as belonging to a specific harm category. For example, if content is blocked due to a high probability of being classified as harassment, the returned safety score will indicate the category as `HARASSMENT` with a `HIGH` probability of harm. By default, the model's inherent safety features are active, and additional filters are disabled. If enabled, developers can set the system to block content based on its probability of being unsafe. The default model behavior is suitable for most applications, and adjustments should only be made when strictly necessary for the application's functionality. The table below outlines the adjustable blocking settings for each category. For instance, setting the 'Hate Speech' category's blocking threshold to 'Block only a few' will result in content with a high probability of being hate speech being blocked, while words with a lower probability may still be permitted. The available thresholds include: `OFF` (disables the safety filter), `BLOCK_NONE` (displays all content regardless of unsafe probability), `BLOCK_ONLY_HIGH` (blocks content with a high probability of being unsafe), `BLOCK_MEDIUM_AND_ABOVE` (blocks content with medium or higher probability of harm), and `BLOCK_LOW_AND_ABOVE` (blocks content with low, medium, or high probability of harm). `HARM_BLOCK_THRESHOLD_UNSPECIFIED` is used when no threshold is set, defaulting to the system's default blocking mechanism. For Gemini 2.5 and 3 models, blocking thresholds are disabled by default if not explicitly set. These settings can be configured for every request made to the generation service. Refer to the `HarmBlockThreshold` API reference for detailed information.

“ Interpreting Safety Feedback

Google AI Studio offers a user-friendly interface for adjusting safety settings. Within the 'Run settings' panel, navigate to 'Advanced settings' and then select 'Safety settings' to enter the 'Run safety settings' mode. Here, users can utilize sliders to modify the content filtering levels for each safety category. It is the developer's responsibility to ensure that the safety settings chosen for their intended use case comply with the Terms of Service. When a request is sent, such as a query to the model, and the content is blocked, a 'Content blocked' warning message will appear. To view more detailed information, hovering the cursor over the 'Content blocked' text will reveal the probability of the content belonging to specific categories and harm classifications.

“ Programmatic Adjustment of Safety Settings

To further enhance your understanding and implementation of safety features with the Gemini API, several resources are available. For a complete overview of the API's capabilities, refer to the API Reference. When developing with Large Language Models (LLMs), adhering to safety best practices is paramount; the Safety Guidelines provide essential information on this topic. For a deeper dive into evaluating probabilities and severity, the Jigsaw team's resources offer valuable insights. Additionally, products that aid in developing safety solutions, such as the Perspective API, can be explored. For instance, you can use these safety settings to build toxicity classifiers, and a classification example is available to get you started. Content on this page is licensed under Creative Commons Attribution 4.0 unless otherwise specified, and code examples are under the Apache 2.0 license. For more details, consult the Google Developers Site Policies. Java is a registered trademark of Oracle and/or its affiliates. The last update for this page was on June 1, 2026 (UTC).

 Original link: https://ai.google.dev/gemini-api/docs/safety-settings?hl=zh-tw

Comment(0)

user's avatar

      Related Tools