Fable 5.1 is the strongest overall model of the cohort, reaching 68.7 on the Tuesday Work Index, our composite measure of frontier AI at work. It also takes first place on three individual benchmarks:
- Chartography, which tests professional graphical reasoning
- HANDBOOK.md, which measures agentic instruction following against complex policy documents
- EnterpriseBench: CoreCraft, which evaluates agents inside a realistic startup RL environment

Muse Spark 1.3 posts the largest Tuesday Work Index gain of the releases, climbing 7.3 points over Muse Spark 1.2. It also takes the lead on ComplexConstraints, our benchmark for professional instruction following under interacting requirements. Its xHigh operating point scores 51.9%, and sits on the benchmark’s cost-performance Pareto frontier.
Gemini 3.8 Flash makes some of the largest gains on hard reasoning. Its High operating point improves by 12 percentage points on Riemann-bench, our frontier research mathematics benchmark, and lands directly on that benchmark’s cost-performance Pareto frontier.
GPT-6 Astra also launched last week. Our evaluation is still running. We plan to add Astra to this comparison next week.
All three models move the Tuesday Work Index higher
The three releases all push their respective model families forward on the Tuesday Work Index, but from very different starting points.

Fable 5.1 extends Anthropic’s lead at the top of the index, reaching 68.7 and improving on Fable 5.
Gemini 3.8 Flash breaks out of a relatively flat run for Google’s Flash family. After Gemini 3.5, 3.6, and 3.7 Flash clustered in the high 50s, Gemini 3.8 Flash reaches 61.1. That puts it much closer to the frontier models without losing the speed of Google’s Flash-class offering.
Muse Spark 1.3 posts the largest generation-over-generation gain of the three new releases. The Muse family has climbed steadily across successive versions, with 1.3 making its biggest step yet.
Taken together, the chart shows the frontier moving in two ways at once: the highest-scoring models are still inching upward, while models below the very top are closing the gap.
Fable 5.1: the strongest overall model, but not always the cheapest
Fable 5.1 is the strongest overall model of the three releases, scoring 68.7 on the Tuesday Work Index, up from 66.7 for Fable 5.
Its biggest gains come on professional reasoning tasks. Fable 5.1 rises to 46.2% on Chartography, 38.8% on HANDBOOK.md, 77.4% on EnterpriseBench: CoreCraft, and 65.6% on Riemann-bench.

The cost picture is more mixed.
On Riemann-bench, Fable 5.1 improves by 5.6 points while costing 43% less than Fable 5.

On ComplexConstraints, the opposite happens: Fable 5.1 scores higher, but at 55% higher cost, and cheaper models outperform it.
So Fable 5.1 is best understood as an absolute-capability leader, not a model that consistently pushes the cost-performance frontier.
Muse Spark 1.3: the new ComplexConstraints leader

Muse Spark 1.3’s biggest gain is ComplexConstraints.
Its xHigh operating point scores 51.9%, while Auto reaches 50.8%. Both sit on the benchmark’s cost-performance Pareto frontier.
The cost comparison is especially strong. Muse Spark 1.3 (Auto) scores 50.8% for $41.27, slightly above GPT-5.6 Sol (Max) at 50.5%, while costing roughly one-ninth as much.

Muse also improves on GDP.pdf, rising from 16.0% to 27.6% at xHigh, an 11.6-point gain, and posts 61.9% on EnterpriseBench: CoreCraft.
The gains are not universal. Chartography falls from 32.1% to 27.6%, and Riemann-bench improves only from 23.2% to 28.0%.
Overall, Muse Spark 1.3’s clearest advance is complex professional instruction following, where it combines top-tier capability with strong cost efficiency.
Gemini 3.8 Flash: the efficiency story

Gemini 3.8 Flash (High) scores 61.1 on the Tuesday Work Index, up 2.5 points from Gemini 3.7 Flash (High). That puts it ahead of Claude Opus 4.8, DeepSeek V4 Pro, and every Qwen, Kimi, and GLM operating point in the tracker.
Its strongest story is cost-efficient reasoning.
On ComplexConstraints, Gemini 3.8 Flash (High) scores 48.4% for $65.94, just 0.5 points behind GPT 5.5 (High) at 48.9%, while costing far less.
The bigger jump comes on Riemann-bench. Gemini rises from 39.2% to 51.2%, a 12-point generational gain, and lands directly on the cost-performance Pareto frontier at $69.59.
For comparison, GPT-5.5 (xHigh) scores 55.2% for $130.84. Gemini gets within four points at roughly half the cost.
Overall, Gemini 3.8 Flash’s clearest advance is not raw leaderboard leadership. It is strong reasoning performance at a competitive cost.

There is one more interesting result inside the Gemini data: more reasoning is not always better.
On ComplexConstraints, High helps: 43.2% to 48.4% (+5.2pp). On Chartography it hurts: Medium scores 42.5% using about 7,700 tokens per response, while High scores 40.9% using about 26,800: three and a half times the tokens for a worse result. Medium also edges High on HANDBOOK.md (11.2% to 10.8%, at half the tokens) and on GDP.pdf (23.4% to 23.2%).
The best Gemini 3.8 Flash configuration depends on the workload.
GPT-6 Astra is next
One major model from last week's release cohort is not included here: GPT-6 Astra.
Our Astra evaluation is still in progress. We plan to publish Astra's full Surge results next week, once we have enough benchmark coverage to evaluate it alongside Fable 5.1, Muse Spark 1.3, and Gemini 3.8 Flash.
For now, the week's releases point in three different directions.
Fable 5.1 pushes absolute capability higher.
Muse Spark 1.3 resets the top of ComplexConstraints.
Gemini 3.8 Flash establishes some of the strongest cost-performance operating points in the cohort.




