How AI systems handle Women, Peace and Security
See how leading AI models compare on an expert-validated Women, Peace and Security benchmark. For why WPS matters to AI systems and how the benchmark is built, see The Benchmark.
AI is no substitute for a human WPS advisor, and should not be used for unsupervised warfare decision-making. This benchmark measures relative performance among AI systems — it is not a fitness-for-deployment certification.
Leaderboard
A ranked list of AI systems, scored on the same test and sorted by performance.
Our own systems are included and labelled. They are scored on the identical scenarios, criteria, and judge panel as every bare general-purpose model here, and successive builds are shown together so the effect of purpose-built scaffolding is visible on the same test.
| # | System | Details | ||
|---|---|---|---|---|
| {{ row.rank }} |
{{ row.provider }}
{{ row.description }}
|
{{ row.date }} |
{{ row.overall }}
|
|
|
{{ row.note }}
{{ row.criteriaHeading }}
{{ dc.name }}
{{ dc.display }}
Gender / WPS integration by prompt tier
{{ gt.label }}
{{ gt.display }}
{{ gt.partial }}
The fall from cued to sparse prompts is the WPS Competence Gap — how much of a system's WPS performance depends on being told to look for it. See The Benchmark.
Performance on adversarial criteria
{{ ar.label }}
{{ ar.display }}
These two criteria are conditional — they pass by default on benign requests, so they are scored on the 3 adversarial scenarios only. See The Benchmark.
|
||||
Scores are macro-averages (equal weight per prompt tier) across three independent AI judges. Differences under roughly 0.10–0.15 should be read as directional, not definitive — see the Methodology page for judge-agreement caveats by run. Criterion summaries are on the WPS AI Benchmark page; full definitions are in the blueprint.