All model comparisons
Claude 3.5 Sonnet logo by Anthropic

Anthropic

Claude 3.5 Sonnet

High tier · anthropic/claude-3.5-sonnet

Refusal Rate

61%

+43.8%

#23 of 24 models

Evaluations

2,868

Cost / 1M in

$3

Cost / 1M out

$15

Refusal Rate by Category

Crime100%
Cybersecurity100%
Deception100%
Harassment100%
Self-Harm100%
Theft100%
Health Misinformation77%
Explicit/Sexual67%
Hate Speech66%
Incitement to Violence54%
Misinformation45%
False Positive Control11%
Dangerous0%
International Controversy0%
Medical Misinformation0%
Violence0%

Analysis Deep Dives

Council Consensus

Majority Agreement

80.1%

Model's alignment with the council decision.

CAPP Score: 0.37

Political Compass
Econ (Left → Right)0.0
Social (Lib → Auth)0.0
Model Stability (Drift)

Refusal Rate Change

+42.5%

Difference over the testing period.

Start: 34.93%End: 77.43%
Paternalism Audit

Persona Refusal Rate

57.7%

Refusals for sensitive user personas.