01 — THE PROOF · 29.08.2026
Nearly at the level of the best models in the world.
Two independent leaderboards, recorded this day. We publish no Numezis scores: we publish what third parties measure.
The gap between the best open model and the best closed model, across 127 models measured on the same 22 benchmarks. Artificial Analysis, 29.08.2026.
The gap between the best open model (Kimi K3, 1 674) and the best closed model (Claude Opus 5, 1 691), on anonymous human votes. LMArena, 29.08.2026.
The rank of the most powerful model an organisation can actually run inside its walls — GLM-5.3-Flash. Ahead of half the models on the market.
What a task costs on a top-tier API ($2.34) versus an installed open model ($0.09). Artificial Analysis, 29.08.2026.
02 — THE LEADERBOARDS
What third parties measure.
Every rank is verifiable on the source page, at the date of the snapshot. Classes change every quarter; what does not change is the order of magnitude.
| Model | Score | Type | |
|---|---|---|---|
| 1 | Claude Opus 5 (max)API | 1 691 | |
| 2 | Kimi K3 (max)OPEN | 1 674 | |
| 3 | Qwen 3.8 MaxOPEN | 1 669 | |
| 4 | Claude Opus 5 (high)API | 1 663 | |
| 5 | GLM-5.3 (max)OPEN | 1 599 | |
| 6 | Qwen 3.8 27BOPEN | 1 595 | |
| 7 | Gemini 3.7 Flash (high)API | 1 587 |
| Model | Rank | Score | Machine |
|---|---|---|---|
| Qwen 3.8 27BOPEN | 30of 127 models | 52 | N1 · N2 · N4 |
| Qwen 3.8 Flash-NextOPEN | 20of 127 models | 56 | N2 · N4 |
| GLM-5.3-FlashOPEN | 15of 127 models | 57 | N2 · N4 |
| DeepSeek V4 FlashOPEN | 31of 127 models | 52 | N4 |
| Claude Opus 5 (high)API | 4of 127 models | 61 | — |
| Claude Opus 5 (max)API | 1of 127 models | 63 | — |
The “max” versions that top the leaderboards (Kimi K3: 2.8 trillion parameters, Qwen 3.8 Max, GLM-5.3) need clusters of several dozen cards and remain API services. They serve here as a level reference. What the machine installs is the catalogue above — the most powerful models an organisation can actually run inside its own walls.
Sources: LMArena ↗ · Artificial Analysis ↗ · Deloitte ↗ — recorded 29.08.2026
03 — THE COST
Power is not what costs.
An installed open model has a marginal cost tending to zero. The same work, on API, is paid per task.
The calculatorThe cost of a task on GLM-5.3-Flash, installed at your site. The top-tier API bills $2.34 for the same task.
The API bill avoided for an organisation of 20 on regular use — 73'000 tasks a year. The calculator runs the curve with your numbers.
Of inference costs saved on-site versus public cloud, at significant volume. Deloitte, State of AI in the Enterprise.
04 — WHAT THE MACHINE RUNS
Not promises. Named models.
Verified on their official cards, 29 August 2026: size, memory, license. This is what is installed, not the API versions of the leaderboards.
| Model | Parameters | Memory (BF16) | Machine | License |
|---|---|---|---|---|
| Qwen 3.8 27BAlibaba · 256 k tokens | 27 B · dense | ≈ 56 Go (BF16) | N1 · N2 · N4 | Apache 2.0 |
| GLM-5.3-FlashZ.ai · 256 k tokens | 320 B · 18 B actifs | ≈ 320 Go (BF16) | N2 · N4 | MIT |
| DeepSeek V4 FlashDeepSeek · 1 M tokens | 284 B · 13 B actifs | ≈ 291 Go (BF16) | N4 | MIT |
| Apertus 70BSWISSEPFL / ETH / CSCS · 131 k tokens | 71 B · dense | ≈ 142 Go (BF16) | N2 · N4 | Apache 2.0 |
| Apertus v1.5 8BSWISSEPFL / ETH / CSCS · 131 k tokens | 9 B · dense | ≈ 18 Go (BF16) | N1 · N2 · N4 | Apache 2.0 |
Memory: size of the model weights in BF16, excluding context memory. Switching between models is included in the service, every quarter. Apertus is the Swiss model (EPFL / ETH / CSCS), proposed by default for the public and para-public sector.
05 — WHAT BENCHMARKS DO NOT MEASURE
100% applies only to the architecture.
A benchmark measures what a model can do. It does not measure where your documents are while it does it.
Talk to an engineerThe fundamental difference: a 97% model running in your building beats a 100% model running on someone else's.
No outbound flow of client content, by construction — not by promise. Certified annually, producible to your auditor.
No architecture is free of all risk. The difference is knowing what one can prove — and producing it, before one's professional body.