paper §7.3: complete 5-model leaderboard (Gemini key restricted, rejoined)
The GEMINI_API_KEY was restricted (fixing the prior 403 enforcement), so Gemini-2.5-flash rejoins the per-role leaderboard. Full ranking: GLM-4.7 > Gemini-2.5-flash > Claude > Kimi > GLM-5.2. Flagship GLM-5.2 still last on the concision-weighted brief. Co-Authored-By: Claude Opus 4.8 <[email protected]>
This commit is contained in:
co-authored by
Claude Opus 4.8
parent
33224f7088
commit
b1991cee3b
@@ -215,20 +215,18 @@ outputs with an independent Claude judge on clarity, impact, and concision:
|
|||||||
|
|
||||||
| Rank | Model (backend) | Note |
|
| Rank | Model (backend) | Note |
|
||||||
|:----:|-----------------|------|
|
|:----:|-----------------|------|
|
||||||
| 1 | **GLM‑4.7** (Zhipu, Sonnet‑class) | tightest, most on‑brief |
|
| 1 | **GLM‑4.7** (Zhipu, Sonnet‑class) | tight, vivid, zero wasted words |
|
||||||
| 2 | Claude Sonnet‑4.6 (Anthropic) | strong but slightly longer |
|
| 2 | Gemini‑2.5‑flash (Google) | clear; "foster" slightly dilutes |
|
||||||
| 3 | Kimi K2 (Moonshot) | clear, a touch generic |
|
| 3 | Claude Sonnet‑4.6 (Anthropic) | strong impact, one clause too many |
|
||||||
| 4 | GLM‑5.2 (Zhipu, Opus‑class) | most capable but *over‑wrote* the 25‑word brief |
|
| 4 | Kimi K2 (Moonshot) | clear, ending a touch redundant |
|
||||||
|
| 5 | GLM‑5.2 (Zhipu, Opus‑class) | jargon‑heavy, longest — *over‑wrote* the 25‑word brief |
|
||||||
|
|
||||||
A notable inversion: the **flagship GLM‑5.2 ranked last** on this *concision‑weighted*
|
A notable inversion: the **flagship GLM‑5.2 ranked last** on this *concision‑weighted*
|
||||||
task — its longer, richer output is an asset for complex reasoning but a liability
|
task — its longer, richer output is an asset for complex reasoning but a liability
|
||||||
when the rubric rewards brevity. This is the core lesson the platform is built to
|
when the rubric rewards brevity. This is the core lesson the platform is built to
|
||||||
surface: there is no globally "best" model — the winner is role‑, task‑, and
|
surface: there is no globally "best" model — the winner is role‑, task‑, and
|
||||||
rubric‑dependent, so model choice belongs to the same tunable layer as topology,
|
rubric‑dependent, so model choice belongs to the same tunable layer as topology,
|
||||||
measured rather than assumed. (Gemini‑2.5‑flash was excluded from this run: the
|
measured rather than assumed.
|
||||||
Google key hit a 403 from Google's June‑19 API‑key‑restriction enforcement preview,
|
|
||||||
an infrastructure constraint unrelated to the model — it succeeds when the key is
|
|
||||||
restricted.)
|
|
||||||
|
|
||||||
### 7.4 Topology × model grid (the interaction effect)
|
### 7.4 Topology × model grid (the interaction effect)
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user