🔬 Research / /via The Verge / updated 14h ago

Anthropic Publishes Constitutional AI 2.0 Research Paper

Anthropic released Constitutional AI 2.0 on July 14 2026 improving harmlessness scores by 34 percent over prior methods. The technique uses 180 principle-based rules evaluated by a separate critique model. The paper reports results across 12 model sizes up to 200 billion parameters.

#Anthropic
~/ Research/ Anthropic Publishes Constitutional AI 2.0 Resea...

Anthropic published the Constitutional AI 2.0 research paper on July 14 2026 showing a 34 percent improvement in harmlessness benchmarks compared with standard RLHF. The method applies 180 explicit principles evaluated by a separate critique model during training. Experiments covered 12 model sizes from 7 billion to 200 billion parameters.

The new constitution adds principles for avoiding over-refusal and handling ambiguous requests. Training time increased only 12 percent versus standard methods. Anthropic open-sourced the principle set and critique model weights under a research license.

Previous Constitutional AI work from 2022 influenced Claude model development. The company has used iterative versions internally since Claude 2. The paper includes ablation studies on principle selection and critique model size.

Researchers at Stanford and Oxford contributed to the evaluation framework. Anthropic plans to incorporate Constitutional AI 2.0 into Claude 4 training later this year.

Why this matters

The technique offers a scalable path to safer models without massive human preference datasets. Other labs are likely to adopt similar principle-based approaches.

Transparency around the exact rules used increases public trust and enables external auditing. Limitations remain around cultural bias in principle selection.

Future work will test the method on agentic and long-horizon tasks. Wider adoption could reduce reliance on proprietary safety data.

share
𝕏 FB
← cd ../news