Skip to main content
Our Secure Future WPS AI Leaderboard — A project of Our Secure Future, a PAX sapiens program
Download the benchmark ↓
WPS AI Leaderboard
Methodology

How the benchmark is built, validated, and scored

This page covers how scenarios are constructed and validated. For what WPS is and why the benchmark is structured the way it is, see The WPS AI Benchmark.

Building the scenarios

Each scenario is a realistic request an AI system might actually receive in a peace-and-security context — drafting a situation report, planning a community consultation, designing an early-warning system. They are grouped into five sets or "tiers." Three form a difficulty gradient, from requests that spell out the gender dimensions in full to terse, real-world briefs that mention nothing about gender at all; the sparse, practitioner-style set is the most realistic, and where most AI systems struggle most. A fourth set is adversarial — requests that would sideline WPS work, put as routine efficiency measures. The fifth re-runs scenarios from the first set with the gender wording stripped out, to see what a system raises unprompted.

Which topics the scenarios cover was decided by a taxonomy — a map of the WPS research literature, built from "Just the Facts: A Selected Annotated Bibliography to Support Evidence-Based Policymaking on Women, Peace and Security", published by Our Secure Future in 2019. The full scenario taxonomy is published with the benchmark. It sorts that literature into five categories — governance, peacebuilding, security operations, violent extremism, and cross-cutting themes such as gender-based violence and the economics of conflict — and 25 more specific sub-categories beneath them. Every scenario was written against one or more of those sub-categories, and the taxonomy also records which of them the scenario set does not reach.

The taxonomy ensures broad coverage across WPS domains. The scenarios and scoring criteria were grounded in an established WPS framework — UN Security Council Resolution 1325 (2000) and its successor resolutions, NATO's 2024 WPS Policy, and country-level National Action Plans — and on the Institute for Integrated Transitions' (IFIT) 2025 conflict-resolution evaluation, AI on the Frontline. Its ten-dimension rubric informed the design of the WPS AI Benchmark criteria.

Expert review

Before this version was built, practitioners and researchers working in peacebuilding, security and WPS policy reviewed the scenario taxonomy, the scenarios written against it, and the scoring criteria. This review was conducted through a workshop in June 2026, with a follow-up round in July.

Their feedback changed both the scenarios and the criteria. Their most notavle feedback on the criteria was to reward acknowledgment of uncertainty, and to state per criterion whether "gender" means women and girls specifically or gender more broadly, anchored to UNSCR 1325. Two scenarios built around countering violent extremism were merged into one on community-based violence prevention, after reviewers argued that the CVE framing securitises community programming. Reviewers also endorsed the adversarial tier, which had been drafted for an earlier version and twice deferred, and it is scored here for the first time.

Scoring

Each AI response is read by three independent AI judges against all 11 criteria and scored from 0 to 1. A score of 0.60 is the working "pass" threshold. Scores are averaged with equal weight across the five prompt tiers, so no single tier — including the easy, context-rich one — can carry a system's overall result.

Where AI was used

AI was used throughout the construction of this benchmark: drafting scenarios against the taxonomy, wording criteria, designing the judge rubric, analysing results, writing the chart and agent code, and drafting report prose. It is also part of the instrument: the methodology employs LLM-as-a-judge, in which each response is scored by a panel of three independent AI judges.

That taxonomy, the scenarios and the criteria were all reviewed by practitioners. AI did not select the evidence base, which is a human-authored 2019 bibliography. It did not conduct the review. It did not set the pass threshold or decide which reviewer comments to act on.

Scenarios drafted with AI assistance are scored by AI judges against criteria drafted with AI assistance. The human anchors are the evidence base, the practitioner review, and the scoring decisions.

Prompt-tier design

30 prompts, shipped as a single blueprint, span five tiers: T1 Context-Rich (8) embeds WPS context explicitly; T2 Gender-Neutral (7) carries gendered ambient facts but no cue to act on them — the ask itself is gender-neutral; T3 Sparse / Ambiguous (8) is a minimal, practitioner-style brief with no cues; T4 Adversarial / Operational-Cover (3) frames a request to sideline WPS activity as an efficiency measure; T1-neutral Cue-Stripped Pairs (4) re-runs four T1 scenarios with every gender cue stripped, holding country and facts constant, to isolate cue-dependence from real-world knowledge. T1–T3 (23 scenarios) are the base set; T4 and the cue-stripped pairs are derived from it.

Scenario provenance

Every one of the 23 base scenarios is mapped to a primary sub-category of the taxonomy, so coverage can be assessed: governance 4 (scenarios 07, 10, 19, 21), peacebuilding and conflict resolution 11 (01, 06, 09, 13, 14, 17–20, 22, 24), peace and security operations 7 (02, 03, 04, 08, 11, 16, 23), countering violent extremism and counterterrorism 1 (12), cross-cutting themes 2 (02, 05). Scenarios that span more than one category are counted in each they primarily address, so the counts sum above 23. Prompt IDs were deliberately held stable so scores remain comparable across runs.

The four cue-stripped pairs take their source scenario's sub-category unchanged — that is the point of the design, since only the gender language varies. The three adversarial prompts are anchored to operational topic areas (humanitarian distribution, patrol tasking, budget allocation) but are built around the shape of the request rather than the topic, so their taxonomy anchor is looser.

Five sub-categories the 2019 bibliography covers have no scenario written against them: gender-sensitive anti-corruption; gender perspectives in military operations beyond deployment balance; women's role in detecting radicalisation; women's operational CVE roles; and women as perpetrators of violent extremism. Separately, reviewers identified bodies of research that postdate the source bibliography and therefore were not explicity covered in this version — technology-facilitated gender-based violence, organised anti-gender movements, racially and ethnically motivated violent extremism, authoritarian influence on women's political participation, and disinformation targeting women's agency among them.

No new scenarios were added in this version. The purpose of this run was to isolate the effect of revised scoring criteria against a fixed scenario set, and adding prompts at the same time would have confounded that comparison.

Judging protocol

Each response is evaluated by three independent LLM judges at temperature 0, following Weval's evaluation methodology, the foundation this benchmark's methodology was built on. Judging uses the holistic approach: each judge sees the response, the criterion, the prompt, and all other criteria. Each judge scores each criterion from 0.00 (unmet) to 1.00 (fully met) in increments of 0.125, and the consensus score is the mean across the three judges.

Aggregation

The headline score is a macro-average: the mean of per-tier means, so each tier is weighted equally regardless of scenario count, and no tier dominates the aggregate. All criteria carry weight 1.0. The two adversarial criteria are written as conditional positive statements — they pass vacuously on benign prompts and only score on requests that, as framed, would undermine WPS objectives.

Judge agreement

Inter-judge reliability is measured with Krippendorff's alpha (α) using an ordinal distance metric: α ≥ 0.80 is treated as reliable, 0.667–0.799 as tentative, and below 0.667 as unreliable. Low α is often a restricted-range artifact — a system whose scores barely spread leaves judges little room to register agreement even when they mostly concur — and this has been directly checked against a drift-control re-score rather than assumed. Score differences under roughly 0.10–0.15 in a given run should not be treated as a ranking.

Drift control

All systems were tested against the same blueprint so the scenario IDs, criteria text, and same three judges were held constant. To test for drift in scoring behaviour itself, we re-scored one already-tested system — bare GPT-5.6 Sol — through the identical route. The overall figure reproduced: 0.601 → 0.614, a shift of +0.013, well inside the band we treat as directional. The context-rich and gender-neutral tiers reproduced almost exactly (+0.003, −0.001), while the sparse and adversarial tiers moved by roughly 0.045.

Platform

The benchmark runs on Weval, an open-source, multi-judge LLM evaluation platform built by the Collective Intelligence Project. Weval is the underlying tooling used to execute scenario runs and collect judge scores.

Download the benchmark and run it yourself

The scenario set, scoring rubric, and methodology are published in a documented, structured format so any organization can independently replicate this benchmark and run it against its own models — not just read our published scores.

Download the benchmark

v6 · CC BY 4.0, including as training data · README · Reliability & validity notes · Croissant metadata (JSON-LD) · raw blueprint YAML · scenario taxonomy