paper §7.4: topology x model grid — the interaction effect
Crosses pipeline x debate with 3 model casts (Claude, GLM-4.7, heterogeneous)
on a fixed decision task, judge-ranked 1-6. Central finding: the topology x
model interaction is real and non-monotone — debate amplifies the strongest
model (GLM-4.7: 2nd->1st) but degrades others (Claude: 3rd->last). The optimal
configuration is a joint choice over {topology x per-role model}, which neither
a topology-only nor model-only study can surface. 21 turns, sequential, within
subscription quota caps. Safety invariant holds across all cells.
Co-Authored-By: Claude Opus 4.8 <[email protected]>
This commit is contained in:
co-authored by
Claude Opus 4.8
parent
a14be693db
commit
33224f7088
@@ -230,6 +230,37 @@ Google key hit a 403 from Google's June‑19 API‑key‑restriction enforcement
|
|||||||
an infrastructure constraint unrelated to the model — it succeeds when the key is
|
an infrastructure constraint unrelated to the model — it succeeds when the key is
|
||||||
restricted.)
|
restricted.)
|
||||||
|
|
||||||
|
### 7.4 Topology × model grid (the interaction effect)
|
||||||
|
|
||||||
|
The platform's defining question is not "best topology" *or* "best model" in
|
||||||
|
isolation but their **interaction**. We crossed two topologies (**pipeline**, 3
|
||||||
|
stages; **debate**, propose→critique→revise→judge) with three model casts
|
||||||
|
(homogeneous Claude, homogeneous GLM‑4.7, and a heterogeneous GLM→Kimi→Claude
|
||||||
|
cast) on a fixed decision task — *"Rust or Go for a 3‑person startup backend?"* — and
|
||||||
|
ranked all six final outputs 1–6 with an independent judge on reasoning quality,
|
||||||
|
clarity, and concision:
|
||||||
|
|
||||||
|
| Topology \ cast | Claude | GLM‑4.7 | Heterogeneous |
|
||||||
|
|------------------|:------:|:-------:|:-------------:|
|
||||||
|
| **Pipeline** | 3 | **2** | 4 |
|
||||||
|
| **Debate** | **6** | **1** | 5 |
|
||||||
|
|
||||||
|
Two results stand out. First, **all six cells converged on the same decision** ("Go")
|
||||||
|
— so on a task with a clear prior, topology and model move *justification quality*,
|
||||||
|
not the answer. Second, and central: **the topology × model interaction is real and
|
||||||
|
non‑monotone.** Debate *amplified* the strongest model (GLM‑4.7: pipeline 2nd →
|
||||||
|
debate 1st, its extra adversarial round adding a hiring‑cost nuance) but *degraded*
|
||||||
|
the others (Claude: 3rd → dead last, its judge wasting words on an "upstream agents
|
||||||
|
agree" meta‑citation; heterogeneous: 4th → 5th). In other words, **the best topology
|
||||||
|
depends on the model, and vice‑versa** — debate is not a universal upgrade; it pays
|
||||||
|
off only with a model strong enough to use the extra rounds. This is precisely the
|
||||||
|
result the platform is built to produce and that neither a topology‑only nor a
|
||||||
|
model‑only study can see: the optimal *configuration* is a joint choice over
|
||||||
|
{topology × per‑role model}, discovered empirically, with the §5 safety invariant
|
||||||
|
holding across every cell. The sweep was run sequentially under the subscription
|
||||||
|
plans' quota/concurrency caps (21 turns total), the practical envelope for this
|
||||||
|
class of experiment.
|
||||||
|
|
||||||
## 8. Limitations and future work
|
## 8. Limitations and future work
|
||||||
|
|
||||||
- Results are offline; the production path (real models + tool‑using turns with
|
- Results are offline; the production path (real models + tool‑using turns with
|
||||||
|
|||||||
Reference in New Issue
Block a user