paper §7.3: complete 5-model leaderboard (Gemini key restricted, rejoined)
ci / gates (push) Has been cancelled
ci / rust (push) Has been cancelled
ci / sandbox-k8s (push) Has been cancelled
ci / frontend (push) Has been cancelled
ci / e2e (push) Has been cancelled

The GEMINI_API_KEY was restricted (fixing the prior 403 enforcement), so
Gemini-2.5-flash rejoins the per-role leaderboard. Full ranking:
GLM-4.7 > Gemini-2.5-flash > Claude > Kimi > GLM-5.2. Flagship GLM-5.2 still
last on the concision-weighted brief.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
This commit is contained in:
Omar Sobh
2026-06-17 14:57:04 -07:00
co-authored by Claude Opus 4.8
parent 33224f7088
commit b1991cee3b
+6 -8
View File
@@ -215,20 +215,18 @@ outputs with an independent Claude judge on clarity, impact, and concision:
| Rank | Model (backend) | Note | | Rank | Model (backend) | Note |
|:----:|-----------------|------| |:----:|-----------------|------|
| 1 | **GLM‑4.7** (Zhipu, Sonnet‑class) | tightest, most on‑brief | | 1 | **GLM‑4.7** (Zhipu, Sonnet‑class) | tight, vivid, zero wasted words |
| 2 | Claude Sonnet‑4.6 (Anthropic) | strong but slightly longer | | 2 | Gemini‑2.5‑flash (Google) | clear; "foster" slightly dilutes |
| 3 | Kimi K2 (Moonshot) | clear, a touch generic | | 3 | Claude Sonnet‑4.6 (Anthropic) | strong impact, one clause too many |
| 4 | GLM‑5.2 (Zhipu, Opus‑class) | most capable but *over‑wrote* the 25‑word brief | | 4 | Kimi K2 (Moonshot) | clear, ending a touch redundant |
| 5 | GLM‑5.2 (Zhipu, Opus‑class) | jargon‑heavy, longest — *over‑wrote* the 25‑word brief |
A notable inversion: the **flagship GLM‑5.2 ranked last** on this *concision‑weighted* A notable inversion: the **flagship GLM‑5.2 ranked last** on this *concision‑weighted*
task — its longer, richer output is an asset for complex reasoning but a liability task — its longer, richer output is an asset for complex reasoning but a liability
when the rubric rewards brevity. This is the core lesson the platform is built to when the rubric rewards brevity. This is the core lesson the platform is built to
surface: there is no globally "best" model — the winner is role‑, task‑, and surface: there is no globally "best" model — the winner is role‑, task‑, and
rubric‑dependent, so model choice belongs to the same tunable layer as topology, rubric‑dependent, so model choice belongs to the same tunable layer as topology,
measured rather than assumed. (Gemini‑2.5‑flash was excluded from this run: the measured rather than assumed.
Google key hit a 403 from Google's June‑19 API‑key‑restriction enforcement preview,
an infrastructure constraint unrelated to the model — it succeeds when the key is
restricted.)
### 7.4 Topology × model grid (the interaction effect) ### 7.4 Topology × model grid (the interaction effect)