Safety & Alignment · since 2022

Constitutional AI

Training a model to critique and revise its own outputs against a written set of principles, reducing reliance on human feedback for harms. Anthropic's method (RLAIF) that scales alignment by substituting AI feedback for human labeling against an explicit constitution.

6

events traced

7

source records

30 Jun 2026

first signal

11 Jul 2026

last activity

Who drove it

Anthropic62%
academia22%
others16%

Key movements

human-label dependence

yes

explicit written principles

2023

public-input variant

The story, event by event

Every point below is traced to a real source — nothing on this page is invented.

  1. 15 Dec 2022impact 78

    Constitutional AI (RLAIF)

    value: RLAIF; metric: method

  2. 17 Oct 2023impact 60

    Collective Constitutional AI

    value: public; metric: input

  3. 22 Sept 2025impact 60

    Enhanced Frontier Safety Framework

    new capability: Critical Capability Level (CCL) for harmful manipulation; risk assessment: expanded protocols for addressing misalignment risks

  4. 14 May 2026impact 65

    ChatGPT integrates safety summaries for enhanced conversational safety

    ChatGPT has been updated with "safety summaries" to improve its ability to recognize and respond to subtle or evolving signs of distress and harmful intent in sensitive conversations.

    ChatGPT received safety updates to recognize emerging risks by identifying subtle or evolving cues and using context for safe responses.

    Work focused on acute scenarios like suicide, self-harm, and harm-to-others, updating model policies and training with mental health experts.

    Developed safety summaries: short, factual notes on earlier safety-relevant context for rare, high-risk situations, created by a model trained for safety reasoning tasks.

    Safe-response performance improved by 50% in suicide/self-harm and 16% in harm-to-others for long single-conversation scenarios, and by 52% and 39% respectively on GPT-5.5 Instant.

    This advancement demonstrates a significant step in scaling AI alignment by enabling models to internally reason about and adapt to complex safety contexts, reducing reliance on direct human oversight for critical harm prevention.

  5. 9 Jul 2026impact 40

    Enhancement of Bio Bug Bounty Program

    scope end date: 2026-07-27; previous reward: 25000; reward increase: 50000

  6. 11 Jul 2026impact 40

    OpenAI's Shift to Family-Oriented AI Products

    hiring focus: family-oriented experiences; safety needs: stronger controls for younger users; user demographics: {'older_users': '31%', 'parents_using_chatgpt': '24%'}

Lineage

Descended from

This is one node. The map holds the whole field.

Watch Constitutional AI — and everything it connects to — grow as real news threads onto the map every day.