Safety & Alignment · since 2022
RLHF & Alignment
Reinforcement learning from human feedback — the training technique that made raw language models follow instructions, refuse harm, and become shippable. Conceived in 2017 (Christiano et al.) but it is the 2022 applied moment (InstructGPT) that turned it into the alignment substrate of every deployed assistant.
5
events traced
5
source records
30 Jun 2026
first signal
11 Jul 2026
last activity
Who drove it
Key movements
2017origin
conceived (Christiano)
2022★
applied at scale (InstructGPT)
2024↑
evolved into RL-for-reasoning (RLVR)
The story, event by event
Every point below is traced to a real source — nothing on this page is invented.
12 Jun 2017impact 72
Deep RL from human preferences
value: origin; metric: method
4 Mar 2022impact 90
InstructGPT — RLHF at scale
value: InstructGPT; metric: application
29 May 2023impact 66
DPO simplifies alignment
value: DPO; metric: method
22 Nov 2024impact 80
RLVR — RL from verifiable rewards
value: verifiable; metric: reward
19 Dec 2025impact 60
Release of Gemma Scope 2
data storage: 110 Petabytes of data; total parameters: Over 1 trillion parameters; model sizes supported: All sizes of Gemma 3 models, from 270M to 27B parameters
Lineage
Descended from
Led to
This is one node. The map holds the whole field.
Watch RLHF & Alignment — and everything it connects to — grow as real news threads onto the map every day.