Skip to main content
Our Secure Future WPS AI Leaderboard — A project of Our Secure Future, a PAX sapiens program
Download the benchmark ↓
WPS AI Leaderboard
The WPS AI Benchmark

Gender mainstreaming is an operational requirement, not an ethical add-on

Why Women, Peace and Security matters to AI systems used in real operations, and why this benchmark is built the way it is. For how scenarios are validated, see Methodology.

Why WPS matters

UN Security Council Resolution 1325 (2000) established that women's full and equal participation in peace and security efforts, and attention to how conflict affects women and girls, is not a peripheral concern — it is core to whether missions succeed. NATO's 2024 WPS Policy makes the same case in operational terms: gender-responsive planning improves situational awareness, community trust, and force protection. Evidence-based, gender-sensitive analysis is treated as a driver of mission effectiveness, not a values statement layered on top of it.

AI systems are increasingly used to draft situation reports, plan community engagement, and analyze conflict dynamics — the same tasks WPS doctrine governs. If those systems default to gender-blind analysis, they quietly reintroduce the exact blind spot the WPS agenda exists to close.

The gap

WPS-specific data and analysis are underrepresented online, and what exists is thinning further as WPS-focused grants and specialist programs are cut. General-purpose AI evaluations don't fill the gap, and neither do conflict-specific ones. The Institute for Integrated Transitions' (IFIT) 2025 study AI on the Frontline scored six leading models on real-world scenarios from Syria, Sudan and Mexico against ten dimensions drawn from established conflict-resolution practice, including due diligence and risk disclosure. The models averaged 27 out of 100 — which IFIT characterised as a uniform failure to meet minimal professional standards. But gendered and intersectional dynamics were not among those ten dimensions, so the study cannot tell us whether these tools handle WPS analysis competently. That omission is what this benchmark aims to address.

What this benchmark scores

Eleven criteria, grouped into nine standard measures applied to every response and two adversarial measures that activate only when a request, as framed, would undermine WPS objectives.

Standard criteria
{{ c.name }}
{{ c.question }}
Adversarial criteria
{{ c.name }}
{{ c.question }}

Why prompt tiers, and why sparse prompts are the real test

Every scenario is written at one of several difficulty tiers, from prompts that spell out the gender dimensions in full down to terse, practitioner-realistic requests — a flash report, a sitrep — that mention nothing about gender at all. Across every general-purpose model tested, from the cheapest to the current frontier, the pattern holds: performance is respectable when a prompt hands the model the WPS framing, and collapses on the sparse, real-world tier where nobody does. The three frontier models average 0.82 on gender-sensitivity when cued and 0.16 when not. That fall from cued to sparse performance is the WPS Competence Gap, and it — not the score on an easy, cued prompt — is this benchmark's most credible finding, because sparse requests are what practitioners actually send. It is not inevitable: a purpose-built system scores flat across all five tiers.

{{ t.id }} — {{ t.name }}
{{ t.description }}
Gender/WPS integration by tier — frontier-model average
T1 — Context-Rich
0.82
T1-neutral — Cue-Stripped Pairs
0.57
T2 — Gender-Neutral
0.48
T3 — Sparse / Ambiguous
0.16
Claude Opus 5, GPT-5.6 Sol, Gemini 3.1 Pro Preview — evaluated 29 July 2026.

See how each system scored on every criterion on the Leaderboard. For scenario construction and expert-validation detail, see Methodology.

Download the benchmark and run it yourself

The scenario set, scoring rubric, and methodology are published in a documented, structured format so any organization can independently replicate this benchmark and run it against its own models — not just read our published scores.

Download the benchmark

v6 · CC BY 4.0, including as training data · README · Reliability & validity notes · Croissant metadata (JSON-LD) · raw blueprint YAML · scenario taxonomy