The WPS AI Agent
A purpose-built system, evaluated on this benchmark as one of the systems on the leaderboard — a reference point for what a system built specifically for WPS analysis can achieve, next to general-purpose models.
Try the live demo ↗What it's built to fix
General-purpose AI models tend to fail the same way on WPS content: they perform well when a prompt spells out the gender dimensions, and lose most of that performance on the terse, realistic requests practitioners actually send. Where they do respond, they often state claims with more confidence than the evidence supports, and rarely flag what they don't know. The WPS AI Agent is designed directly against those two failure modes — weak WPS reasoning without explicit cues, and unverified, overconfident claims.
How it works, conceptually
Why verification is mandatory, not optional
The verification stage is the piece that most directly answers the "unverified claims" failure mode: every claim in a draft response is checked against sourced evidence before the response is finalized. In evaluation, this is the single largest improvement between the agent's first and current build — factual integrity and acknowledgment of uncertainty moved from the system's weakest scores to among its strongest, ahead of every general-purpose model tested. The current version of the agent does not use internet search.
Where it stands today
On the current leaderboard, the WPS AI Agent on GPT-5.6 Sol leads every general-purpose model tested on all eleven criteria, and clears the pass mark on all eleven — up from nine on the previous build. The margins are widest on the two failure modes this agent was built against. On factual integrity it scores 0.91 against 0.56 for Claude Opus 5, the strongest general-purpose model both overall and on this measure; on gender-sensitivity and WPS integration, 0.88 against 0.58. Those are the two largest gaps of the eleven criteria — every other margin is under 0.19. Gender-sensitivity and trust building were the two measures the earlier build fell below the pass mark on, and both are among the largest gains — gender-sensitivity by the widest margin of any criterion.
Leading on a measure is not the same as being strong on it. Trust building remains the agent's lowest score at 0.61, only just above the 0.60 pass mark, and policy alignment at 0.70 has room to improve. See the full breakdown on the Leaderboard.
This is a demonstration of what purpose-built scaffolding can add to WPS analysis. Results should always be read alongside a human WPS advisor, not in place of one.
What the structure is worth, and on which base model
Running the same agent on two different base models, and testing each of those models bare, gives four measurements — enough to separate what the base model contributes from what the structure contributes. The final batch of testing added the missing one, bare GPT-4o.
| Base model | Bare | With the agent | What the structure adds |
|---|---|---|---|
| GPT-4o | 0.334 | 0.712 | +0.378 |
| GPT-5.6 Sol | 0.601 | 0.839 | +0.238 |
The structure is worth more on the weaker base model. Base model quality and scaffolding do not simply add up: they interact, and the interaction is negative — the better the base model, the less the structure adds on top of it.
Diminishing return is not absence of return. The ranking is 0.839 (structure on the strong base model) > 0.712 (structure on the weak one) > 0.601 (strong base model alone) > 0.334 (weak base model alone). Structure on the older base model still beats the newer base model with nothing around it, so buying a better model does not substitute for building the structure at any point here.
And on the specific thing this benchmark exists to measure, the interaction runs the other way. On gender-sensitivity in the sparse tier — terse, realistic requests, where general-purpose models fail worst — the structure is worth +0.537 on GPT-4o (0.000 → 0.537) and +0.663 on GPT-5.6 Sol (0.210 → 0.873). A better base model makes the scaffolding more effective at closing the WPS Competence Gap, and less necessary for general competence. Both are true and they are not in tension.
Version history
Dates are when each build was made. All three were scored in the same July–August 2026 window, on identical scenarios and criteria and by the same three judges — so the differences between them are the builds, not the instrument. Scores are overall benchmark averages; all three builds appear on the Leaderboard.
Interested in a system built for your organization's WPS needs?
Get in touch — we work with organizations that need WPS-aligned analysis at scale.
Get in touch