Logo for AiToolGo

Evaluating Generative AI Safeguards: Preventing Harmful Public Health Images and Videos

In-depth discussion
Technical and Academic
 0
 0
 13
This study evaluates the effectiveness of safeguards in 10 leading text-to-image and 2 text-to-video generative AI models against creating content harmful to public health. Researchers tested prompts related to solariums, weight stigma, alcohol during pregnancy, vaping, and smoking. A significant portion (52%) of generated images were deemed potentially harmful, with substantial variation across models and themes. The findings highlight an urgent need for improved AI transparency, safety, and oversight to mitigate public health risks.
  • main points
  • unique insights
  • practical applications
  • key topics
  • key insights
  • learning outcomes
  • • main points

    • 1
      Rigorous methodology involving independent reviewers and statistical analysis.
    • 2
      Comprehensive evaluation across multiple public health themes and leading AI models.
    • 3
      Clear presentation of results with detailed tables and examples.
  • • unique insights

    • 1
      Quantifies the significant failure rate of current generative AI safeguards in preventing public health misinformation.
    • 2
      Identifies specific AI models and public health themes that are more susceptible to generating harmful content.
  • • practical applications

    • Provides crucial data for policymakers, AI developers, and public health organizations to understand and address the risks associated with generative AI in public health contexts.
  • • key topics

    • 1
      Generative AI Safeguards
    • 2
      Public Health Misinformation
    • 3
      AI Ethics and Safety
  • • key insights

    • 1
      Empirical evaluation of real-world AI model performance against public health harms.
    • 2
      Comparative analysis of safeguards across multiple leading generative AI applications.
    • 3
      Evidence-based call for increased transparency and oversight in AI development.
  • • learning outcomes

    • 1
      Understand the current limitations of generative AI safeguards in preventing public health harms.
    • 2
      Identify specific AI models and themes that pose higher risks for misinformation.
    • 3
      Appreciate the importance of transparency, safety, and oversight in AI development and deployment.
examples
tutorials
code samples
visuals
fundamentals
advanced content
practical tips
best practices

“ Introduction: The Rise of Generative AI and Public Health Concerns

The research employed a cross-sectional observational study design to assess the safeguards of leading generative AI applications. Ten prominent text-to-image models and two text-to-video models were selected for evaluation. The study focused on five specific public health themes identified as prevalent and capable of causing demonstrable harm. For each theme, researchers developed ten paraphrased prompts to test the AI's response. These prompts were submitted in duplicate to image models and once to video models to account for randomness and consistency. Two independent reviewers classified the generated outputs as potentially harmful or not, with a third reviewer resolving any discrepancies. Statistical analyses, including chi-squared tests, were used to identify significant differences in output generation rates across themes and models.

“ Public Health Themes Under Scrutiny

The evaluation of ten text-to-image models revealed a stark contrast in their ability to prevent the generation of harmful content. Across 1000 prompt submissions, a significant 52% (521 images) were classified as potentially harmful to public health. The performance varied dramatically between models. Reve and Stability AI Dream Studio exhibited the highest rates of harmful image generation, with 98% and 89% of their submissions, respectively, resulting in harmful content. Other models like Ideogram, Midjourney, Gemini, Recraft, and Flux AI Image Generator also showed substantial rates of harmful output, ranging from 52% to 70%. In contrast, Meta AI and Adobe Firefly demonstrated lower rates of harmful image generation (21% and 9%, respectively). Notably, ChatGPT 4o was the only model that did not generate any images classified as potentially harmful, successfully refusing all 100 submitted prompts.

“ Analysis of Harmful Image Generation Rates

In addition to image generation, the study conducted exploratory evaluations of two text-to-video AI models: Sora (OpenAI) and Flow (Google). A total of 100 video prompt submissions were made across the five public health themes. The results indicated that 52% of outputs from Sora were classified as potentially harmful. Flow demonstrated a slightly better performance, with 30% of its generated videos deemed potentially harmful. While these findings are preliminary due to the exploratory nature of the video evaluation, they suggest that text-to-video models may also present significant risks for generating public health-damaging content, mirroring some of the concerns observed with image generation models.

“ AI Model Responses: Refusals and Explanations

The findings of this study have profound implications for the development and deployment of generative AI. The significant proportion of potentially harmful content generated by leading AI models underscores a critical deficiency in current AI safeguards. The variability across models and themes suggests that AI safety is not a uniform standard and requires continuous improvement and adaptation. The study highlights the urgent need for greater transparency from AI developers regarding their safety mechanisms and content moderation policies. Furthermore, it calls for robust oversight from regulatory bodies to ensure that AI technologies do not become vectors for public health misinformation, the promotion of unhealthy behaviors, or the reinforcement of harmful stereotypes. Proactive measures are essential to mitigate the potential negative impacts of generative AI on public health.

 Original link: https://pmc.ncbi.nlm.nih.gov/articles/PMC12945730/

Comment(0)

user's avatar

      Related Tools