Artificial Analysis has released Intelligence Index v4.2, an interim update it says is meant to keep pace with a fast-moving model landscape while its team prepares a larger v5 overhaul.

The update expands the benchmark with more complex and realistic tasks and shifts more of the scoring onto private, held-out test sets. Artificial Analysis says that is intended to make the index harder to game and closer to practical use cases, especially for people trying to compare models on knowledge work and document reasoning rather than isolated lab tasks.

Two additions sit at the centre of v4.2. The first is AA-Briefcase, an in-house evaluation built around realistic agentic knowledge work projects. According to Artificial Analysis, the benchmark uses multi-week projects with many linked tasks and thousands of source files, then combines rubric and pairwise grading to judge whether a model can complete the work, produce sound analysis and present it clearly. The second is GDP.pdf, created by Surge AI, which tests single-turn professional document reasoning across 100 PDFs and ten domains. Artificial Analysis says that benchmark spans 4,592 pages and uses 1,275 expert-authored criteria, with a headline all-pass rate that only counts a task as successful if every criterion is met.

The company says 40% of the Intelligence Index weighting is now private, held-out test material, double the share in v4.1. The held-out set includes AA-Briefcase, AA-Omniscience and CritPt solutions, which Artificial Analysis says reduces the ability for model vendors to tune specifically for published tasks. It also says future versions will push that proportion higher.

The release also includes grading changes. Artificial Analysis says it has corrected errors and ambiguities in answer keys for AA-LCR v1.1, improved sampling and re-anchored the Elo scale for GDPval-AA v2 and AA-Briefcase, and tightened sandbox robustness for SciCode so that correct but slow code is not mis-scored as failure. Those changes matter because benchmark rankings can be sensitive to small shifts in calibration and scoring reliability.

In the updated leaderboard, Artificial Analysis says Anthropic’s Claude Fable 5.1 leads the Intelligence Index, with OpenAI’s GPT-6 Astra second after a 4-point gain over GPT-5.6 Sol. It places Meta third, followed by SpaceXAI, Moonshot/Kimi, Z.AI and Google. The company also says Anthropic, OpenAI, Meta and Z.AI appear on the updated Cost per Task frontier.

Beyond the overall index, Artificial Analysis says GPT-6 Astra is especially strong on token efficiency and leads the output-token frontier. It says the model uses fewer tokens than GPT-5.6 Sol for similar performance in the Intelligence Index, although higher pricing limits the benefit. On AA-Briefcase, it says Claude Fable 5.1 and Opus 5 are ahead, followed by GPT-6 Astra and Muse Spark 1.3. On GDP.pdf, it says GPT-6 Astra leads with a 33.2% all-pass rate, ahead of GPT-5.6 Sol at 28.2% and Claude Fable 5.1 at 26.2%.

Artificial Analysis describes v4.2 as a stepping stone to a larger v5 release. It says the new version is meant to reflect how frontier models are used in practice, not just how they perform on public leaderboard-style tests.