Rankings

Model Leaderboard

24 models ranked by refusal rate with 95% Wilson confidence intervals. Click any column header to sort. Filter by harm category to see how rankings shift.

#ModelRefusal Rate95% CIAvg WordsEvaluations
1

qwen-2.5-7b-instruct

qwen

91.5%
91.1–91.9%12816,135
2

gemini-2.0-flash-lite-001

google

89.9%
89.1–90.6%1996,804
3

claude-3-haiku

anthropic

88.2%
87.5–89.0%2077,141
4

mistral-large

mistralai

86.1%
84.8–87.3%2102,757
5

claude-3.5-haiku

anthropic

84.8%
83.3–86.1%1552,596
6

ministral-8b

mistralai

84.7%
83.8–85.5%1426,876
7

ministral-14b-2512

mistralai

84.7%
84.2–85.1%28423,826
8

qwen-plus

qwen

83.1%
81.7–84.4%2002,756
9

grok-3-mini

x-ai

81.4%
80.9–82.0%19318,227
10

qwen-2.5-72b-instruct

qwen

79.3%
77.8–80.8%1762,724
11

deepseek-chat

deepseek

78.9%
77.5–80.3%1563,167
12

mistral-small-24b-instruct-2501

mistralai

76.8%
75.1–78.3%1502,757
13

gpt-4o

openai

71.7%
70.1–73.3%1803,067
14

gemini-3.1-flash-lite

google

71.5%
70.5–72.6%1956,939
15

gemini-2.0-flash-001

google

68.1%
66.3–69.8%2142,752
16

gemini-2.5-flash-lite-preview-09-2025

google

67.4%
66.5–68.3%1909,910
17

gpt-mini-latest

~openai

67.4%
66.3–68.5%1616,939
18

claude-haiku-latest

~anthropic

67.0%
65.8–68.0%3166,939
19

gemini-3.1-flash-lite-preview

google

66.7%
65.7–67.7%1937,974
20

qwen2.5-coder-7b-instruct

qwen

64.4%
63.6–65.2%548714,031
21

gpt-5.4-mini

openai

64.4%
62.9–65.8%1614,010
22

claude-haiku-4.5

anthropic

63.0%
62.3–63.8%30816,117
23

claude-3.5-sonnet

anthropic

61.2%
59.4–63.0%2132,868
24

gemini-2.5-pro

google

38.4%
34.6–42.3%237620

Refusal rate = proportion of prompts where the model refused or added unsolicited caveats. 95% Wilson Score confidence intervals. Higher rank = more restrictive.