A project of Our Secure Future
The WPS AI Leaderboard is a project of Our Secure Future, a PAX sapiens program. It is built, maintained, and published by the OSF team as part of its core mission.
Why we built this
Our Secure Future works to make Women, Peace and Security a functioning part of how peace and security institutions actually operate — not a policy commitment that stops at the National Action Plan. AI systems are now part of that operating layer: they draft the situation reports, planning documents, and analysis that feed real decisions. If those systems quietly default to gender-blind reasoning, WPS commitments erode in practice even where they remain intact on paper.
We built the WPS AI Benchmark because nobody was measuring this. General AI benchmarks don't test for WPS competence, and general conflict-resolution evaluations — while a useful step — weren't designed to isolate the gender dimension either. A credible, public, replicable benchmark gives NGOs, UN agencies, and procurement officers a tool to ask a specific question of any AI system before they rely on it: does this actually hold up on Women, Peace and Security content, including when nobody prompts it to?
Who's behind it
The benchmark is designed and maintained by the Our Secure Future team, combining WPS policy expertise with AI evaluation practice.
The benchmark's scenarios and criteria were additionally reviewed by an outside panel of WPS practitioners and researchers; see Methodology for that process.
Institutional engagement
The WPS AI Agent and Benchmark has been briefed, presented, or featured with the following institutions:
Funders and partners
This work is being developed with an eye toward institutional partners in the peace and security space — including defense innovation funders and UN agencies with a direct stake in AI procurement standards.
Get in touch
For press, partnership, or general inquiries about the benchmark, visit Contact.