Safety & Alignment · since 2022
Constitutional AI
Training a model to critique and revise its own outputs against a written set of principles, reducing reliance on human feedback for harms. Anthropic's method (RLAIF) that scales alignment by substituting AI feedback for human labeling against an explicit constitution.
6
events traced
7
source records
30 Jun 2026
first signal
11 Jul 2026
last activity
Who drove it
Key movements
↓↓
human-label dependence
yes★
explicit written principles
2023↑
public-input variant
The story, event by event
Every point below is traced to a real source — nothing on this page is invented.
15 Dec 2022impact 78
Constitutional AI (RLAIF)
value: RLAIF; metric: method
17 Oct 2023impact 60
Collective Constitutional AI
value: public; metric: input
22 Sept 2025impact 60
Enhanced Frontier Safety Framework
new capability: Critical Capability Level (CCL) for harmful manipulation; risk assessment: expanded protocols for addressing misalignment risks
14 May 2026impact 65
ChatGPT integrates safety summaries for enhanced conversational safety
ChatGPT has been updated with "safety summaries" to improve its ability to recognize and respond to subtle or evolving signs of distress and harmful intent in sensitive conversations.
ChatGPT received safety updates to recognize emerging risks by identifying subtle or evolving cues and using context for safe responses.
Work focused on acute scenarios like suicide, self-harm, and harm-to-others, updating model policies and training with mental health experts.
Developed safety summaries: short, factual notes on earlier safety-relevant context for rare, high-risk situations, created by a model trained for safety reasoning tasks.
Safe-response performance improved by 50% in suicide/self-harm and 16% in harm-to-others for long single-conversation scenarios, and by 52% and 39% respectively on GPT-5.5 Instant.
This advancement demonstrates a significant step in scaling AI alignment by enabling models to internally reason about and adapt to complex safety contexts, reducing reliance on direct human oversight for critical harm prevention.
9 Jul 2026impact 40
Enhancement of Bio Bug Bounty Program
scope end date: 2026-07-27; previous reward: 25000; reward increase: 50000
11 Jul 2026impact 40
OpenAI's Shift to Family-Oriented AI Products
hiring focus: family-oriented experiences; safety needs: stronger controls for younger users; user demographics: {'older_users': '31%', 'parents_using_chatgpt': '24%'}
Lineage
Descended from
This is one node. The map holds the whole field.
Watch Constitutional AI — and everything it connects to — grow as real news threads onto the map every day.