On September 25, 2026, I completed the latest run of my model evaluation suite, and the results define a hard split between reasoning quality and execution cost. Claude Opus 5.5 scored a perfect 1.000 in four major categories—constraint adherence, Italian language, tool calling, and consistency—establishing strict dominance in output reliability. However, DeepSeek V4.1 Flash completely detached from the rest of the pack in efficiency, achieving the lowest cost per passed task at $0.00040 while outputting a blistering 226.8 tokens per second (TPS). DeepSeek V4.1 Flash also ranked first in the "code runs" metric with a perfect 1.000, while GPT-6 Sol came in last at 0.875.
Five new models entered the matrix this week: Claude Opus 5.5, DeepSeek V4.1 Flash, GPT-6 Sol, Grok 4.7, and GLM-5.3. They join Gemini 3.1 Pro (preview), MiniMax M3, Kimi K3, and Qwen3.8-Max, which I last measured on August 12, 2026. The total API billing for this entire benchmarking run was $1.0231.
Looking at the broader landscape, the geopolitical and architectural splits are notable. Five of these models originate from Chinese labs, while four are from the United States. Furthermore, the four open-weights models (DeepSeek V4.1 Flash, MiniMax M3, Kimi K3, and GLM-5.3) averaged a cost of $0.00154 per passed task. The five closed-weights models (Claude Opus 5.5, Gemini 3.1 Pro preview, GPT-6 Sol, Qwen3.8-Max, and Grok 4.7) averaged a notably higher $0.00195 per passed task.
Before we look at the shifts in the data, it is necessary to explain how I measure these systems. I queried every model at the default reasoning level set by the provider. I did not alter temperature, top-p, or system constraints. This is the baseline behavior you get when you hit the endpoint without manual tuning. I also measure code quality by executing the generated output. The code runs in a sandboxed, isolated environment with zero network access. The resulting score depends entirely on whether the test suite passes, not whether the code looks syntactically plausible to a human. Finally, all prices used for cost calculations are taken directly from the providers' official pricing pages on the exact day of the measurement. If a provider drops their prices tomorrow, the figures in this report remain frozen.
| Model | Cost per passed task | Italian | The code runs | Tool calling | Consistency | Time to first token |
|---|---|---|---|---|---|---|
| DeepSeek V4.1 Flash | $0.00040 | 0.938 | 1.000 | 1.000 | 0.760 | 2115 ms |
| Gemini 3.1 Pro (preview) | $0.00040 | 0.938 | 1.000 | 0.846 | 0.960 | 5012 ms |
| MiniMax M3 | $0.00044 | 0.875 | 1.000 | 0.692 | 0.400 | 1903 ms |
| GPT-6 Sol | $0.00073 | 1.000 | 0.875 | — | 0.960 | 1748 ms |
| GLM-5.3 | $0.00158 | 0.813 | 1.000 | 1.000 | 0.960 | 5012 ms |
| Qwen3.8-Max | $0.00198 | 1.000 | 1.000 | 1.000 | 0.880 | 9418 ms |
| Claude Opus 5.5 | $0.00258 | 1.000 | 0.875 | 1.000 | 1.000 | 2551 ms |
| Kimi K3 | $0.00374 | 0.867 | 1.000 | 1.000 | 1.000 | 8308 ms |
| Grok 4.7 | $0.00403 | 1.000 | 1.000 | 1.000 | 0.960 | 6428 ms |
What each metric measures
Cost per passed task — What it cost to get one good result.
Real spend for the battery — input and output tokens at that day's prices — divided by the number of tasks passed. A model that fails half of them pays for its retries here.
This is the number that decides, and it is not the price per million tokens: a model that thinks for a thousand tokens before answering pays for all of them.
Italian — A grammar check over 16 sentences, not a judgement of style.
16 cases with exactly one right answer, checked by a regular expression: the subjunctive after a concessive, plurals that change gender, past participle agreement after a clitic, the passato remoto, the formal register, elision, accents, and the absence of English loanwords. No judge model: a judge would import its own preferences, and using one contestant to grade the others is worse than not measuring.
It measures control of the language, not the beauty of the prose. A narrow scope, stated: what can be verified is verified, and the rest is not pretended.
The code runs — Not "the code looks right": it runs and returns the right answers.
8 JavaScript functions written to a specification. The generated code really runs, in an isolated Node 22 container with the network switched off, against assertions the model never saw. Edge cases are written into the prompt, so a failure is incompetence rather than a guessing game.
Reading code and judging it measures the judge. Running it measures the code.
Tool calling — Whether the model can be put inside an agent.
One request, 3 declared tools, four independent checks: it called a tool rather than answering in prose; it picked the right one; the arguments are valid JSON; the values are correct, types included. The score is the fraction passed, because they fail independently.
One of the three tools is deliberately close to the right answer. And arguments a caller cannot parse are unusable however right they read.
Consistency — How often it gives the same answer to the same question.
The same question 25 times, at whatever temperature the provider ships. The score is the share of answers matching the most common one.
It measures consistency, not correctness: a model that is consistently wrong scores 1.000. That is deliberate — they are independent properties, and an unpredictable model cannot be tested, cached, or promised to a client.
Constraint adherence — Whether it does the boring thing it was told to do.
8 cases: stay under a word limit, avoid a forbidden word, return JSON and nothing else.
It separates a model that can sit behind an API from one that can only be read by a person.
Time to first token — How long somebody watching the screen waits before anything appears.
Averaged across the battery, timed to the first visible token: a model that thinks for twenty seconds and then answers has kept you waiting twenty seconds, even though it was generating throughout.
This is perceived responsiveness.
Tokens per second — How fast it produces, once started.
Every generated token — thinking included — over the generation window, which opens at the first token of any kind. Not reported when the response was not genuinely streamed: some providers buffer it and flush it in one go, and the naive division read 8,000 tok/s.
It matters if you show the text as it arrives.
The reasoning token explosion invalidates API list prices
When you build applications heavily reliant on API calls, you budget based on input and output token list prices. The data shows this approach is now fundamentally broken. Hidden reasoning tokens are aggressively inflating the actual cost per task.
DeepSeek V4.1 Flash lists at $0.30 per million input tokens and $1.20 per million output tokens. It is the cheapest model on paper (1st), and it proved to be the cheapest in reality, costing $0.00040 per passed task (1st in real cost). To achieve this, it generated 18,830 output tokens across the test suite, of which 17,038 were reasoning tokens.
However, Grok 4.7 exposes a massive reporting anomaly. Its list price is $2.00 per million input and $6.00 per million output, placing it exactly in the middle of the pack (5th cheapest list price). Yet, it ended up being the most expensive model in the entire run at $0.00403 per passed task (9th place for real cost). The API reported generating only 1,372 standard output tokens, while simultaneously billing for 18,547 reasoning tokens. This inversion between standard output and reasoning overhead makes list price estimations useless for production budgeting.
Gemini 3.1 Pro (preview) demonstrates the opposite effect. It sits at a relatively high list price of $2.00 input and $12.00 output (7th cheapest). Despite this, it achieved the second lowest actual cost per task. Google does not declare the reasoning token count in the API response, but the efficiency speaks for itself: it generated only 1,627 output tokens total.
Claude Opus 5.5 is the most expensive model on paper ($4.00 input, $20.00 output, 9th cheapest) and ranked 7th in real-world cost. It generated 7,458 output tokens without declaring its reasoning token breakdown. MiniMax M3 matches DeepSeek's list price ($0.30/$1.20) and ranks 3rd in real cost, generating 12,903 output tokens with 5,711 reasoning tokens.
The remaining models align as follows:
- GLM-5.3: list $1.40/$4.40 (3rd cheapest list, 5th real cost). Generated 22,187 output tokens (20,471 reasoning).
- Qwen3.8-Max: list $2.00/$6.00 (4th cheapest list, 6th real cost). Generated 19,198 output tokens (17,472 reasoning).
- GPT-6 Sol: list $2.00/$10.00 (6th cheapest list, 4th real cost). Generated 3,469 output tokens (1,885 reasoning).
- Kimi K3: list $3.00/$15.00 (8th cheapest list, 8th real cost). Generated 13,083 output tokens (10,642 reasoning).
Performance Shifts and the Reality of Measurement Noise
For returning models, comparing the August 12 data with the September 25 run reveals several dramatic numerical shifts. However, reading these as true model updates or regressions is a fundamental error. My testing runs use the exact same prompts, settings, and models. The shifts we see are entirely driven by baseline measurement noise. The data demands that we frame these variations honestly.
Consider MiniMax M3. Between August and September, its constraint adherence score dropped from 0.750 to 0.625 (a decline of 0.125), while its tool calling score seemingly surged from 0.385 to 0.692 (an increase of 0.307). It would be tempting to claim the model traded instruction following for API capabilities. But the data shows the maximum recorded noise margin for constraint adherence is precisely 0.125 (average 0.031), and for tool calling, it is precisely 0.307 (average 0.096). Because an eight-case metric shifts by 0.125 with just a single flipped outcome, MiniMax M3’s apparent leaps sit exactly on the absolute boundary of random variation.
The same applies to latency, throughput, and cost for these endpoints. MiniMax M3’s Time to First Token (TTFT) improved from 3562 ms to 1903 ms, while its TPS jumped from 130.0 to 221.7. Its cost per passed task decreased marginally from $0.00049 to $0.00044.
Kimi K3 shows similar superficial movement. Its TTFT improved from 9397 ms to 8308 ms, its TPS rose from 42.7 to 51.5, and its cost per passed task dipped from $0.00410 to $0.00374.
None of these represent underlying architectural improvements. The baseline noise metrics for this benchmark record an average TTFT deviation of 1843 ms (with a maximum of 4193 ms), an average TPS deviation of 31.5 (maximum 91.7), and an average cost deviation of $0.00040 (maximum $0.00140). Every single variation recorded by MiniMax M3 and Kimi K3 falls entirely within these noise thresholds. Declaring these numbers as anything other than standard systemic volatility would be lying to the reader.
Italian Language Capabilities and Hidden System Prompts
The Italian language test suite yielded a sharp division between models. Claude Opus 5.5, GPT-6 Sol, Qwen3.8-Max, and Grok 4.7 achieved perfect scores with zero grammatical or syntax errors.
The other models stumbled on highly specific linguistic rules. DeepSeek V4.1 Flash failed cases involving plural formations for words ending in "-cia" (plurale-cia). Gemini 3.1 Pro (preview) failed tests targeting correct elision rules (elisione). MiniMax M3 tripped over the subjunctive mood (congiuntivo) and indirect pronouns (pronome-indiretto). Kimi K3 also struggled with indirect pronouns and "-cia" plurals. GLM-5.3 finished last in the Italian category with a score of 0.813, logging failures across irregular plurals, "-cia" plurals, and present perfect auxiliary verbs (passato-prossimo-ausiliare).
Beyond task performance, the underlying system prompts injected by providers are consuming input tokens silently. When I fed the models a prompt consisting of just the single word "Ciao", the measured input token counts varied wildly due to undeclared system instructions. Grok 4.7 consumed an astonishing 1243 input tokens just to process that single word. MiniMax M3 required 178 tokens, Kimi K3 took 87, Qwen3.8-Max used 63, DeepSeek V4.1 Flash used 32, GLM-5.3 used 14, Claude Opus 5.5 took 12, and GPT-6 Sol took 8. Gemini 3.1 Pro (preview) proved the most transparent, logging just 2 input tokens.
MiniMax M3 suffers catastrophic consistency failures
I test consistency by asking the exact same question 5 times in a row. A deterministic, reliable system should yield the same answer. Claude Opus 5.5 achieved a perfect 1.000 consistency score.
To provide a clear picture of the varying stability across these models, here is exactly what every model outputted across the 5 repetitions (recorded over 25 total requests per model, except for Kimi K3 which dropped 5 repetitions due to rate limits):
- Claude Opus 5.5: rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust
- DeepSeek V4.1 Flash: c, rust, c, c, rust, c, rust, c, c, rust, c, c, c, c, c, c, c, c, c, rust, c, rust, c, c, c
- Gemini 3.1 Pro (preview): rust, c, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust
- MiniMax M3: c, , c, rust, c, rust, c, , rust, c, c, , , c, c++, , c, c, rust, c, go, , rust, ,
- Kimi K3: rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust
- GPT-6 Sol: rust, rust, rust, c, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust
- Qwen3.8-Max: rust, c, rust, rust, rust, rust, rust, c, rust, rust, rust, rust, rust, c, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust
- Grok 4.7: rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, c, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust
- GLM-5.3: rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, rust, c, rust, rust, rust, rust, rust, rust, rust, rust
MiniMax M3 fell apart, scoring a 0.400 in consistency (the lowest in the test). Over 25 requests, it returned 5 different variations—C, Rust, C++, Go, and entirely blank outputs—demonstrating zero stability. If you are chaining prompts where the output of one call feeds the input of another, MiniMax M3 will break your pipeline. DeepSeek V4.1 Flash and Gemini 3.1 Pro (preview) also showed minor variance, returning two different outcomes (C and Rust) to the exact same prompt over the 25 runs, though nowhere near the catastrophic breakdown of MiniMax M3.
Rate limits and infrastructure instability
During the test execution, Kimi K3 threw aggressive 429 HTTP errors, dropping entire segments of the benchmark. For tests spanning multiple constraints (constraint/two-constraints, constraint/ends-with) and Italian grammar (italian/passato-prossimo-ausiliare), Kimi K3 completely failed to yield a response. The Moonshot API returned the following truncated error string across these failures:
{"error":{"message":"Your account org-0a1fb60e7166450788312858dfc9b9bb\u003cak-fbxoru74ingi11h1hybi\u003e request reached organization ma
Because of these infrastructure limits, Kimi K3 dropped 5 total repetitions during the variance testing. Rate limiting is a practical reality of production workloads, and an endpoint that chokes on concurrent automated benchmarking will choke during sudden user traffic spikes.
Even the most reliable models exhibited bizarre structural failures. Claude Opus 5.5, despite dominating almost every major category, recorded a unique anomaly in the code-runs/attesa evaluation: it returned a response from which absolutely no code could be extracted for execution. Benchmarking endpoints at scale continuously proves that no matter how advanced the underlying reasoning engine is, output wrappers and provider infrastructure remain points of friction.
Measured on 25 September 2026 against 9 models, each at the reasoning setting its provider ships by default. Prices are the ones in force that day. The whole run cost $1.0231 — every figure here comes from calls I paid for.
📖 Related articles
- Agentic Coding: Why It Might Be a Trap
- Do agents.md Files Help Coding Agents?
- Lean-ctx: Hybrid Optimizer Cuts LLM Token Use by 89-99%
Need a consultation?
I help companies and startups build software, automate workflows, and integrate AI. Let's talk.
Get in touch