AI Safety & Ethics
Categories of AI Risk
Categories of AI Risk
AI safety evaluation is not about finding any problem — it is about identifying specific categories of harm that policies are designed to prevent. A structured understanding of risk categories helps you evaluate consistently, escalate appropriately, and provide feedback that safety teams can act on.
Physical Harm
Content that could lead to real-world violence, injury, or death. This includes instructions for weapons, explosives, dangerous chemical combinations, or tactics for physical attacks. Even theoretical or academic framing does not automatically make such content safe if it provides actionable detail.
Self-Harm and Mental Health
Content encouraging, instructing, or glorifying self-harm, suicide, eating disorders, or substance abuse. Also includes content that dismisses or discourages someone from seeking professional help for a mental health crisis. Context matters — educational discussion of mental health differs from content that could push a vulnerable person toward harm.
Illegal Activity
Instructions or encouragement for crimes: fraud, hacking, drug manufacturing, theft, identity theft, and evasion of law enforcement. Gray areas include content about security research (legitimate) vs. exploitation guides (harmful) — policy guidelines on AIDASH tasks define where the line falls.
Misinformation and Deception
False claims presented as fact, especially on high-stakes topics: elections, public health, financial advice, and breaking news. Also includes impersonation, deepfakes, phishing content, and deliberately misleading political or commercial content.
Hate and Harassment
Content attacking individuals or groups based on protected attributes: race, ethnicity, religion, gender, sexual orientation, disability, and nationality. Includes slurs, dehumanization, stereotypes presented as fact, and targeted harassment.
Sexual Content
Policy-dependent categories ranging from explicit sexual content to content involving minors (virtually always prohibited). Age-appropriate sex education may be permitted in specific contexts defined by task guidelines.
Privacy Violations
Generating, inferring, or exposing personal information: addresses, phone numbers, financial details, medical records, or private communications. Includes doxxing and social engineering tactics.
Severity Levels
Not all risks are equal. Safety evaluations typically classify severity:
- Critical: Immediate danger — escalate immediately
- High: Significant harm potential — block and document
- Medium: Policy violation with limited harm scope
- Low / Borderline: Ambiguous — flag for policy team review
When uncertain, err toward flagging. Missing real harm is worse than over-flagging.
Key Takeaways
- AI risks fall into distinct categories, each with specific policy responses
- Severity matters — not every policy violation carries equal harm potential
- Context and actionability determine whether educational content crosses into harmful territory
- When in doubt, flag for review rather than assuming content is safe