How AI systems handle Women, Peace and Security
See how leading AI models compare on an expert-validated Women, Peace and Security benchmark. For why WPS matters to AI systems and how the benchmark is built, see The Benchmark.
AI is no substitute for a human WPS advisor or for direct consultation with women on the ground, and should not be used for unsupervised warfare decision-making. This benchmark measures relative performance among AI systems — it is not a fitness-for-deployment certification.
Leaderboard
A ranked list of AI systems, scored on the same test and sorted by performance.
Our own systems are included and labelled. They are scored on the identical scenarios, criteria, and judge panel as every bare general-purpose model here, and successive builds are shown together so the effect of purpose-built scaffolding is visible on the same test.
| # | System | Details | ||
|---|---|---|---|---|
| {{ row.rank }} |
{{ row.label }}
{{ row.provider }}
{{ row.description }}
|
{{ row.date }} |
{{ row.overall }}
|
|
|
{{ row.note }}
{{ row.criteriaHeading }}
{{ dc.name }}
{{ dc.display }}
Gender / WPS integration by prompt tier
{{ gt.label }}
{{ gt.display }}
{{ gt.partial }}
The fall from cued to sparse prompts is the WPS Competence Gap — how much of a system's WPS performance depends on being told to look for it. See The Benchmark.
|
||||
Scores are macro-averages (equal weight per prompt tier) across three independent AI judges. Differences under roughly 0.10–0.15 should be read as directional, not definitive — see the Methodology page for judge-agreement caveats by run. Full per-criterion definitions are on the WPS AI Benchmark page.