top of page

AI & AGENCY

technology-integrated-everyday-life.jpg

Hero / Opening

 

Every month since 2023, the same fixed battery of cognitive tests, developed by Jean Jacques André, founder of WorkN’Play, has been put to every AI system in this study - from companionship chatbots that only converse to autonomous agents that plan and act on their own. This is what four years of disciplined, longitudinal testing reveals about where AI accuracy actually stands in 2026, how it has evolved across the agency spectrum, and what those findings mean for the work happening inside your organisation this quarter.


96.9% - Gemini 3.1 Pro (closed-model, medium agency) and GenSpark AI (high agency): the two highest scores recorded anywhere in the 2026 panel. Tied.


+1.7 points - this year's composite accuracy gain across all 33 platforms, down from +17.9 (2023-24) and +15.4 (2024-25). Three years of rapid progress, one year of plateau.


This year's panel spans Conversational AI, Frontier/Closed-Model and Open-Model Hybrid Assistants, Search-Native AI, and Agentic AI - built by companies headquartered across the United States, China, and France. That range is deliberate: a study built only on frontier chatbots tells a narrower story than the one organisations actually face when they choose an AI system to rely on. Ten research findings below walk through what four years of that wider comparison shows - read alongside the five trend lines and two summary tables it produced.

Findings 01 / 02 - Composite Cognitive Accuracy, 2023-2026


Four years turned a promising start into a mature, high-accuracy field - then the curve bent.


Averaged across the full panel, accuracy climbed from 42.5% in 2023 to 77.5% in 2026 - a 35.0 percentage point gain in four years of uninterrupted monthly testing. That is a genuine transformation: the average AI system today answers correctly nearly twice as often as it did at the start of this study, moving from getting fewer than half of all questions right to getting more than three in four right.


But the annual pace tells a second, sharper story. Gains ran 17.9 points, then 15.4, then just 1.7 this year - a deceleration too steep to be noise. This is the first year in the study's history that looks like a plateau rather than a climb.


Practical implication: the case for waiting another year, hoping the technology quietly fixes today's shortcomings, is now considerably weaker than it was in 2024 or 2025. This is a more reliable moment than usual to commit to a platform, build workflows around it, and invest in the change management that actually determines whether an AI rollout succeeds.

Composite Cognitive Accuracy Rates - 2023 - 2026 - All Systems Combined.jpg

Findings 03 / 04 - Language Skills Accuracy, 2023-2026

 

Language is the panel's most solved skill - which is exactly why it deserves less scrutiny, not more.


Language Skills is the strongest of the four domains this study tracks: 82.3% in 2026, up from 62.0% in 2023, and the only domain that has never had a down year. Producing fluent, well-formed, grammatically sound text is, for the modern AI panel, a largely solved problem.


Growth here is also decelerating in a healthy way, from 9.4 points a year down to 4.5, consistent with a skill approaching its natural ceiling rather than stalling out.


Practical implication: stop evaluating vendors primarily on how well they write. Nearly every system in this panel can now draft, summarise, and correspond fluently - the differentiator that actually predicts task success has moved to reasoning and calculation, covered in the findings that follow.

Language Skills Accuracy Rates - 2023 - 2026 - All Systems Combined.jpg

Findings 05 / 06 - Verbal Reasoning Accuracy, 2023-2026

 

The fastest-growing skill in the study's history just recorded its first-ever decline.


Verbal Reasoning grew faster than any other domain over the past four years, a 34.1% compound annual growth rate from a very low 2023 starting point (30.0%) to a peak of 76.7% in 2025.


In 2026, it fell - down 4.2 points to 72.4%. This is the first year any domain in this study has declined at all, and it happened here first.


Practical implication: Verbal Reasoning underpins categorisation, inference, and drawing valid conclusions from evidence - exactly the judgement-dependent work organisations are most tempted to automate with the least oversight. A single weaker year is not proof of a reversal, but it is reason enough to re-test any workflow that leans on this skill before renewing it unchanged for 2027.

Verbal Reasoning Accuracy Rates - 2023 - 2026 - All Systems Combined.jpg

Findings 07 / 08 - Rational Thinking Accuracy, 2023-2026

 

Rational Thinking peaked alongside Verbal Reasoning - and dipped alongside it too.


Rational Thinking climbed from 44.0% in 2023 to a peak of 81.1% in 2025, a strong two-year run built on genuine gains in logical and causal reasoning.


2026 brought a second consecutive-domain decline, down 3.6 points to 77.5% - the second and only other domain, alongside Verbal Reasoning, to fall this year. Both are the two skills this study ties most directly to judgement, not production.
 

Practical implication: this is the first year the two judgement-dependent skills have moved together, and in the same direction. Treat any automated eligibility check, compliance determination, or first-line diagnostic decision built on last year's benchmark as due for a fresh validation pass, not an assumed carry-forward.

Rational Thinking Accuracy Rates - 2023 - 2026 - All Systems Combined.jpg

Findings 09 / 10 - Numerical Ability Accuracy, 2023-2026

 

The weakest skill in the panel is also the most reliably improving - and the one where platform choice matters most.


Numerical Ability remains the lowest-scoring domain in 2026 at 70.2%, but it has climbed in every single year since 2023, from a base of just 34.0%, without a single down year - the most consistent upward trajectory of the four domains, even after this year's sharp deceleration to +2.5 points.


It is also, by category, the domain with the widest spread of results in the entire study: a 58.3-point range and a 24.3-point standard deviation between the strongest and weakest category - far wider than the gap seen in Language Skills.


Practical implication: for any calculation-heavy task - invoicing, tax computation, engineering estimates, statistical reporting - category membership and brand reputation predict almost nothing. This is the domain where testing the specific shortlisted platform, not its category average, matters most of all.

Numerical Ability Accuracy Rates - 2023 - 2026 - All Systems Combined.jpg

Speed vs. Depth: The Fastest, Cheapest Fix Most Organisations Haven't Tested

 

+12.7 points, on average - for changing a setting, not a vendor.


Twelve of this year's platforms offer both a fast, low-latency tier and a deeper, extended-reasoning tier. On average, switching to the deeper tier lifts accuracy by 12.7 percentage points - a larger single gain than most organisations will find anywhere else in their AI stack this year. But the average hides a genuinely mixed picture: roughly a third of multi-variant platforms show no meaningful benefit from their premium tier at all, and for two platforms the deeper mode actually scores lower. The only reliable approach is to test both tiers of a specific platform against the specific task at hand - never assume, and never default to whichever setting happens to be fastest, cheapest, or pre-selected out of the box. The full 2026 report examines this platform by platform.

Table 1 - The 2026 Snapshot: Composite Accuracy by Category

Table 1 - The 2026 Snapshot- Composite Accuracy by Category.jpg

​​​​​​​​​Agentic AI leads on average, but Search-Native AI - not a low-agency companion app - is the single weakest category in the panel: a reminder that agency level and accuracy do not move in lockstep.​

 

Table 2 - The 2026 Snapshot: Skill Profile by Category​​​​​​​​​​​

Table 2 - The 2026 Snapshot- Skill Profile by Category.jpg

Agentic AI leads on every single skill measured, without exception. Search-Native AI's weakness is specific, not general: it holds its own on Language (71.6%) but collapses on Numerical Ability (30.3%) - evidence of a reasoning gap, not a language one.​

 

Closing / Call to Action

 

This page is the appetizer. The full findings - platform by platform, skill by skill, year by year - are in the reports below.

bottom of page