Files
rustytorch/docs/archive/legacy/phase11.md
T
2026-03-04 00:08:42 +00:00

10 KiB
Raw Blame History

RustyTorch++ — PyTorch Parity & Superset Master Plan (Phases 010)

Purpose: A precise, phasebyphase specification of everything we must recreate in Rust (PyTorch parity) and everything we will surpass (superset) — explicitly tied to our GPUnative compiler rustg and the artifacts we ship. This consolidates the roadmap with unambiguous scope, acceptance tests, and crate ownership.

Context

Progress snapshot (per your tracker, 20250811): Phases 19 marked COMPLETE; Phase 10 in progress; Phase 0 foundations largely complete with a few docs/Contrib items pending.

Roots/paths:

RUSTYTORCH_ROOT: /home/osobh/projects/rustytorch

RUSTG_COMPILER: /home/osobh/projects/rust/rustg (primary target: RTX 5090 / sm_120)

STRATOSWARM_ROOT: /home/osobh/projects/stratoswarm

  1. Foundation (Repo, Tooling, Contracts)

Recreate (parity): N/A (setup); establish reproducible builds; CI; docs baseline akin to PyTorch contributor experience.

Rust + rustg deliverables:

Workspace layout, build graph, lint/format gates; CUDA/ROCm/Metal toolchain detection; Nsight/rocprof wiring.

rtx-governance (SBOM, signing), rtx-bench (criterion + traces), rtx-profiler shims.

Superset:

Determinism & perf CI gates from day zero; signed artifacts.

Exit tests: CI green on GPU runner; perf/determinism smoke suites publish artifacts; contributor guide complete.

Status: ~90% per tracker → Action: finish env docs + CONTRIBUTING.

  1. Core Compiler & Runtime (CUDA/ROCm/Metal)

Recreate (parity): PyTorch CUDA/ROCm backends & runtime semantics (streams, events, graphs, allocators).

Rust + rustg deliverables:

Crates: rtx-compiler, rtx-runtime, rtx-kernel, rtx-bench, rtx-profiler.

Passes: canonicalize → type/shape → layout normalize → early fusion; lower to rustg → fatbin/hsaco.

Runtime: Device/Stream/Event/Graph/Module/TensorStorage; pooled allocator; CUDA/HIP Graph capture.

Superset:

Zerocost abstractions; allocator frag telemetry; deterministic mode.

Exit tests: ≥20% steptime vs eager; 6h soak w/ no leaks; fixedseed tolerances met.

Status: COMPLETE (20250811)

  1. Tensor API, Autograd & Graph IR

Recreate (parity): torch.Tensor (dense first), views/strides/broadcasting, autograd (reversemode), graph capture/export.

Rust + rustg deliverables:

Crates: rtx-tensor, rtx-autograd, rtx-ir (midIR + serializer), rtx-compiler wiring to rustg.

Backward registry for core ops; AMPaware gradients.

Superset:

Typelevel safety (const generics for shapes/dtypes); IR roundtrip goldens.

Exit tests: Correct grads vs analytical/FD; IR roundtrip; ≥20% steptime drop vs P1.

Status: COMPLETE (per tracker; continue IR serialization polish)

  1. Distributed Training & Parallelism

Recreate (parity): DDP, FSDP (ZeROstyle), pipeline & tensor parallelism, elastic training, distributed ckpt.

Rust + rustg deliverables:

Crates: rtx-dist (process groups, NCCL/RCCL/MPI), rtx-fsdp (or integrated), distributed checkpointing.

Topology discovery (NVLink/IB), gradient buckets, overlap.

Superset:

Stratoswarmnative scheduling; auto hybrid DP/TP/PP planner; WAL + elastic recovery.

Exit tests: 1→8 GPU ≥0.8× efficiency; multinode ≥0.7×; FSDP memory ↓ ≥40%.

Status: COMPLETE

  1. Advanced Compilation & AutoKernel Synthesis

Recreate (parity): TorchInductorclass fusions/compilation paths.

Rust + rustg deliverables:

Crates: rtx-synth (pattern lib + template emitters), autotuner w/ persistent cache, hardware profile DB.

AOT graph compiler/loader; synthesized kernels (attention/MLP/norm/conv pilot).

Superset:

2040% steptime wins; 1.5× tokens/s; cache hit ≥80%; AOT load <100ms.

Exit tests: Macro wins ≥20%; determinism parity; cache hit threshold met.

Status: COMPLETE

  1. Inference Runtime & Serving

Recreate (parity): vLLMclass scheduler, paged KV, streaming, quantization; multinode serving.

Rust + rustg deliverables:

Crates: rtx-infer (continuous batching, SLA lanes, speculative decoding), quant (INT8/INT4/FP8), KV paging.

gRPC/HTTP/WebSocket servers; client SDKs (Rust/Python initial).

Superset:

Hotreload models; Stratoswarm autoscaling; blue/green + canary; p99 < 150ms targets.

Exit tests: ≥1.5× throughput vs eager; cache hit ≥85%; brownout under overload.

Status: COMPLETE

  1. SelfOptimizing Platform (Unified Data+Compute, Governance)

Recreate (parity): profiler, bottleneck tools, reproducibility docs; dataset → model pipelines.

Rust + rustg deliverables:

Crates: rtx-graph (ETL dialect + ML ops), zerocopy IO (GDS/RDMA), provenance & SBOM pipeline.

Autotuning agents with sandbox → canary → promotion; dashboards.

Superset:

Telemetrydriven optimization with guardrails; signed reproducible builds; unified ETL+Model graph.

Exit tests: E2E walltime ↓ ≥15%; autotune wins ≥1025% without regressions.

Status: COMPLETE

  1. Ecosystem, SDKs & Interop

Recreate (parity): LibTorch ABI, PyTorch Python API surface (core), ONNX, DLPack, utils.

Rust + rustg deliverables:

Crates: rtx-bindings (C ABI + PyO3 Python), ONNX import/export, DLPack zerocopy.

Docs site, examples, release channels; plugin registry scaffolding.

Superset:

Featuregated bindings; WASM preview; agentgenerated examples; typed FFI boundaries.

Exit tests: SDK install matrices; interop parity ≥95% on suites; signed artifacts.

Status: COMPLETE

  1. Autonomous Agentic Evolution

Recreate (parity): N/A — beyond PyTorch.

Rust + rustg deliverables:

Crates: rtx-evolve (proposal engine, multiobjective optimizer, sandbox), knowledge graph of changes.

MCP allowlists per agent; CI/CD hooks for propose→validate→promote.

Superset:

≥70% merged proposals produce measurable wins; 0 production regressions; full provenance.

Exit tests: Canary p95/p99 within SLA; rollback rehearsed; determinism gates pass.

Status: COMPLETE

  1. Global MultiTenant Platform & Federation

Recreate (parity): Largescale ops playbooks, quotas, audit, compliance (PyTorch ecosystem + infra tools).

Rust + rustg deliverables:

Crates: rtx-platform (regions, routing, quotas, billing, residency, audit), artifact CDN for AOT/kernels.

Federation (DP budgets, secure aggregation), SRE automation & DR.

Superset:

Activeactive regions; signed usage records; policy linter; 99.95% regional SLO.

Exit tests: Global p95 ≤ 1.3× singleregion; accurate metering (<1% error); DR drills pass.

Status: COMPLETE

  1. 1.0 Release, Governance & LongTerm Sustainability

Recreate (parity): Stable/unstable API taxonomy; release discipline; docs at operator coverage.

Rust + rustg deliverables (to finish):

API freeze + semver; LTS vs Fast channels; full docs with stable/unstable markers.

TSC governance charter; community plugin registry; sponsorship & partner program.

Superset:

Signed, reproducible releases; crash telemetry optin; model hub with perf leaderboard.

Exit tests: 100% build/test across matrices; docs complete; governance operational; partner SKUs validated.

Status: IN PROGRESS

PyTorch Feature Family → RustyTorch++ Mapping (Parity Matrix)

"Recreate in Rust" checklist mapping major torch.* surfaces to crates/modules. Status reflects your tracker (most complete), and remaining items for 1.0.

PyTorch family

RustyTorch++ crate/module

Status

Notes / Remaining for 1.0

torch, torch.Tensor

rtx-tensor, rtx-autograd

Add complex dtype ops coverage list in docs.

torch.nn, torch.nn.functional

rtx-nn (module set within rtx-tensor/rtx-autograd)

Publish parity table per module.

torch.cuda, torch.cuda.memory

rtx-runtime (CUDA), allocator

Document CUDA graph semantics.

torch.mps, torch.xpu

rtx-runtime backends (Metal/LevelZero)

/♻︎

MPS parity subset; publish support matrix.

torch.amp

AMP scaler in rtx-runtime/rtx-autograd

Tolerance tables in docs.

torch.distributed.*

rtx-dist, rtx-fsdp

Userlevel API sugar (oneliner init) in docs.

torch.compile/inductor

rtx-compiler, rtx-synth, rtx-ir

Expose AOT format docs.

torch.export, torch.fx, torch.jit

rtx-ir + AOT (rtx-compiler)

FXstyle transforms doc.

torch.onnx

rtx-bindings + ONNX I/O

Operator mapping table published.

torch.profiler

rtx-profiler + Nsight/rocprof

Trace schema docs.

torch.utils.data

rtx-data (in rtx-graph ETL dialect)

Examples: sharded iterators.

torch.utils.dlpack

rtx-bindings DLPack

Interop examples.

torch.optim

rtx-optim

Fused optimizers doc.

torch.sparse, torch.masked, torch.nested

rtx-sparse

/♻︎

Publish operator coverage matrix.

torch.linalg, torch.fft, torch.signal, torch.special

rtx-linalg, rtx-fft, rtx-signal, specials in core

/♻︎

Ensure grad coverage; doc edge cases.

torch.random

RNG in rtx-runtime

Repro seeds across distributed.

torch.package, torch.hub

rtx-package, rtx-hub

Model cards + provenance.

torch.utils.mobile_optimizer

Mobile targets

iOS/Android/WASM targets doc.

Legend: complete • ♻︎ polishing/extend docs

Acceptance Contracts (per category)

For each mapped family above, 1.0 requires: (a) API reference with stable/unstable badges; (b) examples; (c) determinism/perf baselines; (d) interop tests (where applicable).

Remaining 1.0 Checklist (Phase 10)

Publish stable/unstable labels across docs; generate API diffs in CI.

Finalize Metal/LevelZero support matrix (feature table + fallbacks).

Release AOT graph & kernel bundle format spec.

Ship model hub minimal portal + CLI; signed submissions.

Governance: TSC charter, RFC repo, release cadence (LTS vs Fast).

"Oneliner distributed init" helper for newcomers.

Windows build & test gates.

Ownership & RACI (abbrev.)

Rust Engineer — runtime, tensor, autograd, backends.

Frontend Compiler Agent — IR/passes, export/AOT.

Distributed Strategy Agent — rtxdist/FSDP, planner.

Kernel Synthesizer & AutoTuner — rtxsynth, cache, profile DB.

Performance Engineer — benches, profiler, gates.

Inference Scheduler Agent — rtxinfer (batching, KV, decoding).

Governance/Provenance — SBOM/sign, model hub, artifacts.

Docs & SDK Agent — bindings, docs site, examples, API badges.

Executive Summary (why this surpasses PyTorch)

By Phase 10, RustyTorch++ provides a GPUnative, Rustsafe, selfoptimizing platform that unifies training, compilation, inference, orchestration, and governance. We match PyTorchs breadth and exceed it with agentic optimization, AOT portability, signed reproducibility, and planetscale multitenancy — all without Python overhead.