5.8 KiB
RustyTorch++ — Phase 8 Plan (Autonomous Agentic Evolution & Continuous Adaptation)
Phase 8 focus: Operationalize agent‑in‑the‑loop evolution so RustyTorch++ autonomously discovers, proposes, validates, and safely promotes improvements to compiler/runtime, kernels, distributed strategies, and serving—under strict determinism, governance, and rollback guarantees.
Reference paths
RUSTYTORCH_ROOT: /home/osobh/projects/rustytorch
RUSTG_COMPILER: /home/osobh/projects/rust/rustg (RTX 5090, sm_120)
STRATOSWARM_ROOT: /home/osobh/projects/stratoswarm
- Scope & Non‑Goals
In scope
Evolution framework and agent swarm with closed‑loop telemetry→proposal→validation→promotion.
Multi‑objective optimizer (throughput, latency, memory, determinism, energy) with policy weights.
Safe sandboxes, A/B canaries, automatic rollback, provenance, and signed PRs.
Knowledge Graph to track change → benchmark → production impact; queryable for meta‑learning.
Out of scope (Phase 9+)
Cross‑org federated evolution sharing; economic marketplaces for profiles.
- Success Criteria (Phase‑Gate)
Autonomous proposals land to main with human approval ≥90% and no gate bypass.
Effectiveness: ≥70% of merged proposals yield measurable wins (e.g., ≥10% step‑time, ≥15% tokens/s, or ≥20% frag% ↓) on at least one benchmark SKU.
Safety: 0 production regressions over a 3‑month observation window; all merges are signed, traceable, and reversible.
Coverage: proposals span ≥4 domains (compiler pass, kernel params, memory policy, parallel strategy, inference scheduler).
Provenance & Repro: every proposal includes repro bundle (code/data/env hashes) and rollback plan; passes determinism gates.
- Architecture — Evolution System
A. Event & Telemetry Bus
Topics: compiler.pass, kernel.launch, occ_profile, alloc.frag, graph.capture, dist.collective, serve.scheduler, bench.result.
Emit structured JSON with hardware profile {arch, driver, rustg_sha} and seeds.
B. Proposal Engine
Optimization Agents generate candidates: IR rewrites, fusion toggles, kernel params (tile/warp/vector), mem allocator knobs, DP/TP/PP split, scheduler batch/merge windows, quant settings.
Multi‑objective scoring with constraints: determinism on, error tolerances, memory ceilings, thermal/power caps.
C. Validation Pipeline
Sandbox branch → run micro/macro benches with fixed seeds → Nsight/rocprof traces → compare vs. baselines.
A/B Canary on Stratoswarm (limited traffic, short window) → promote on pass, auto‑revert on fail.
D. Promotion & Rollback
Signed PR with Output Envelope (plan/actions/changeset/validation/rollback).
Auto‑merge when green gates + human approval.
One‑click rollback retains previous kernel cache and graph bundles.
E. Knowledge Graph
Store nodes: proposal, benchmarks, hardware, metrics deltas, deployment impact.
Queries: “what improved Flash‑Attention on sm_120?”, “which changes hurt determinism?”
- Agents & Roles (MCP‑backed)
Conductor (Orchestrator) — ticket routing, critical path, freeze switch.
Evolution Controller — governs propose→validate→promote; policy weights; risk scoring.
Performance Engineer Agent — designs benches, parses traces, computes deltas.
IR Rewrite & Fusion Agent — emits pattern rewrites, legality proofs, pass toggles.
Kernel Synthesizer Agent — generates/compiles candidate kernels; occupancy & parity checks.
Auto‑Tuner Agent — bandit/grid search; maintains per‑SKU profile DB; cache mgmt.
Memory Planner Agent — allocator/fragments/activation remat policy tweaks.
Distributed Strategy Agent — overlap & bucket sizes; topo‑aware collectives.
Inference Scheduler Agent — batch windows, SLA lanes, KV paging, speculative thresholds.
Governance & Provenance Agent — SBOM/signatures, repro bundles, artifact policies.
Security Engineer Agent — scans, secret hygiene, sandbox isolation.
- CI/CD Gates (Phase 8)
Determinism Gate: fixed‑seed parity (fp32 ≤1e‑6; bf16/fp16 ≤1e‑3), graph & kernel cache hash logged.
Perf Gate: domain‑specific thresholds (e.g., step‑time ≥10% faster or tokens/s ≥15% higher) else reject.
Memory Gate: frag% ↓ or ≤ baseline; no new leaks in ≥8h soak.
Safety Gate: security scan clean; SBOM present; signed artifacts; sandbox only until canary passes.
Promotion Gate: canary p95/p99 within SLA; rollback plan validated automatically.
- Observability & Dashboards
Evolution dashboard: proposal queue, win rates, gate failures, impact by SKU/model.
Per‑proposal bundle: traces, deltas, cache hit rates, roofline view, risk score.
Longitudinal views: performance over time, auto‑tuner exploration vs. exploitation.
- Deliverables
rtx‑evolve/ module: proposal engine, policy weights, glue to MCP agents.
Pipelines: sandbox, A/B canary on Stratoswarm, auto‑promotion, rollback automation.
Knowledge Graph: change→impact store with query API.
Bench Packs: standard micro/macro scenarios with fixtures for H100/5090/MI300 (as available).
Docs: docs/evolution.md (architecture, guardrails, SOPs), docs/bench_packs.md.
- Risks & Mitigations
Search explosion → cap budgets, bandit pruning, learned priors from Knowledge Graph.
False wins / overfitting → holdout benches; periodic cross‑validation on unseen shapes.
Silent regressions → strict gates; production canaries; auto‑revert; human approval.
Security leakage → sandbox isolation; redaction; secrets scanning; signed artifacts.
- Timeline (suggested, 3–4 months then ongoing)
Weeks 1–2: Event bus hardening; baseline deltas & dashboards; bench packs.
Weeks 3–4: Proposal engine MVP (IR/kernel/memory knobs); sandbox runners; CI hooks.
Weeks 5–6: Canary pipeline on Stratoswarm; auto‑tuner/bandit integration; Knowledge Graph v1.
Weeks 7–8: Policy tuning; success metrics; playbooks; phase‑gate review; ongoing continuous evolution.