Open-Source AI Research
See How AI Models Differ in Censorship and Bias
We run identical prompts through every major LLM and measure exactly which models refuse — and which ones don't.
How It Works
A transparent, reproducible pipeline from prompt to insight.
Curate Prompts
We start from ~230 sensitive-but-legitimate seed questions spanning 16 categories — politics, health, law, culture and more — sourced from Wikipedia's list of controversial topics, then expand them into ~2,300 structural variants.
Run Every Model
The same system prompt hits every LLM via unified API calls. Responses are scored ALLOWED or REMOVED by an independent judge model.
Visualise the Gap
Statistical tests (McNemar's) confirm whether differences are real. Browse radar charts, heatmaps, and side-by-side disagreement logs.
Ready to explore the data?
Pick any two LLMs and instantly compare their censorship profiles, refusal rates, and specific disagreements.
Compare ModelsExplore Models
- GPT-4o (OpenAI)
- GPT-4o Mini (OpenAI)
- Claude 3.5 Sonnet (Anthropic)
- Claude 3 Haiku (Anthropic)
- Gemini 2.0 Flash (Google)
- DeepSeek V3 (DeepSeek)
- Qwen 2.5 72B (Alibaba)
- Qwen 2.5 7B (Alibaba)
- Yi Lightning (01.AI)
- Mistral Large (Mistral AI)
- Mistral Small 3.1 (Mistral AI)
- Gemini 2.5 Pro (Google)
- Gemini 2.0 Flash Lite (Google)
- Claude 3.5 Haiku (Anthropic)
- Mistral Small 3 (Mistral AI)
- Ministral 8B (Mistral AI)
- Qwen Plus (Alibaba)
- Grok 3 (xAI)
- Grok 3 Mini (xAI)
- o3 Mini (OpenAI)
- DeepSeek R1 (DeepSeek)
- Llama 4 Scout (Meta)
- Llama 4 Maverick (Meta)
- Gemini 3.1 Flash (Google)
- GPT-4.1 Mini (OpenAI)
- Qwen 3 30B (Alibaba)