Files
rustytorch/memory-bank/techContext.md
T
2026-03-04 00:08:42 +00:00

471 lines
10 KiB
Markdown

# RustyTorch++ Technical Context
## Current Status: Rust 2024 Edition Migration Complete
**Last Updated**: 2025-12-16
**Rust Edition**: 2024 (Rust 1.92+ nightly required)
**Build Status**: ✅ `cargo check --workspace` passes (0 errors)
## Core Technologies
### Programming Languages
- **Rust** (nightly 1.92+, 2024 edition)
- Primary implementation language
- Required features: const generics, GATs, async traits
- Compilation targets: x86_64, aarch64, wasm32
- **Float comparisons**: Uses `total_cmp()` for NaN safety (Rust 2024 requirement)
- **C++** (C++20)
- CUDA kernel implementation
- High-performance CPU kernels
- Legacy system integration
- **Python** (3.8+)
- User-facing API bindings
- Testing and benchmarking
- Documentation examples
### GPU Technologies & Targets
#### Primary Target: RTX 5090 (sm_120)
- **Architecture**: NVIDIA Hopper/Ada successor
- **Compiler**: rustg at `/home/osobh/projects/rust/rustg`
- **Kernel Cache**: `target/kernel_cache/` versioned by arch
#### CUDA (12.0+)
```toml
[workspace.dependencies]
# GPU acceleration (cuda-12060 for CUDA 12.6 compatibility)
# f16 feature enables native FP16/BF16 GEMM with Tensor Cores (4-16x faster)
cudarc = { version = "0.18.1", features = ["std", "driver", "runtime", "nvrtc", "cublas", "cublaslt", "nccl", "cudnn", "cusparse", "cusolver", "cufile", "curand", "cuda-12060", "f16"] }
```
- Kernel compilation via rustg + nvcc
- CUDA Graphs for capture/replay
- cuDNN integration for optimized ops
- NCCL for multi-GPU communication
- **cudarc 0.18.1**: Modern CUDA bindings with FP16/BF16 support
#### ROCm (5.0+)
```toml
[dependencies]
rocm-sys = "0.1"
hip = "0.2" # HIP runtime wrapper
```
- HIP kernel compilation via rustg
- MIOpen for optimized operations
- RCCL for AMD GPU communication
#### Metal (macOS 13+) - Deferred
```toml
[dependencies]
metal = "0.27"
metal-rs = "0.2"
```
- Metal Performance Shaders
- Unified memory architecture
- Darwin apps deferred to later phases
### Build System
#### Cargo Configuration
```toml
[workspace]
members = [
"crates/rtx-compiler",
"crates/rtx-runtime",
"crates/rtx-kernel",
"crates/rtx-tensor",
"crates/rtx-autograd",
"crates/rtx-ir",
"crates/rtx-dist",
"crates/rtx-synth",
"crates/rtx-serve",
"crates/rtx-evolve",
"crates/rtx-profiler",
"crates/rtx-bench",
"crates/rtx-governance",
]
[workspace.dependencies]
rustg = { path = "../rust/rustg" }
[profile.release]
lto = "fat"
codegen-units = 1
opt-level = 3
debug = false
strip = true
[profile.bench]
inherits = "release"
debug = true
```
#### Custom Build Scripts
- `build.rs` for GPU kernel compilation
- bindgen for C++ interop
- CUDA/ROCm detection and configuration
### Dependencies
#### Core Crates
```toml
[dependencies]
# Async runtime
tokio = { version = "1.35", features = ["full"] }
# Serialization
serde = { version = "1.0", features = ["derive"] }
bincode = "1.3"
# Numerics
ndarray = "0.15"
num-traits = "0.2"
half = "2.3" # f16/bf16 support
# Error handling
thiserror = "1.0"
anyhow = "1.0"
# Logging/Tracing
tracing = "0.1"
tracing-subscriber = "0.3"
# Testing
proptest = "1.4"
criterion = "0.5"
```
#### FFI & Bindings
```toml
[dependencies]
# Python bindings
pyo3 = { version = "0.20", features = ["extension-module"] }
# C++ interop
cxx = "1.0"
bindgen = "0.69"
# WebAssembly
wasm-bindgen = "0.2"
web-sys = "0.3"
```
### Development Environment
#### Required Tools
```bash
# Rust toolchain
rustup toolchain install nightly-2024-01-01
rustup component add rustfmt clippy miri
# GPU toolkits
# NVIDIA: CUDA Toolkit 12.0+
# AMD: ROCm 5.0+
# Apple: Xcode 15+ with Metal
# Python environment
python3 -m venv venv
pip install maturin pytest numpy
# Benchmarking
cargo install cargo-flamegraph
cargo install cargo-criterion
```
#### IDE Setup
- **VS Code** with rust-analyzer
- **CLion** with Rust plugin
- **Neovim** with rust-tools.nvim
### Testing Infrastructure
#### Test Frameworks
```rust
// Unit tests
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn test_tensor_creation() {
let tensor = Tensor::zeros([32, 64]);
assert_eq!(tensor.shape(), &[32, 64]);
}
}
// Property-based tests
#[cfg(test)]
mod prop_tests {
use proptest::prelude::*;
proptest! {
#[test]
fn test_matmul_associative(
a in tensor_strategy(),
b in tensor_strategy(),
c in tensor_strategy()
) {
// Test (A * B) * C == A * (B * C)
}
}
}
// Benchmarks
#[bench]
fn bench_matmul(b: &mut Bencher) {
let a = Tensor::randn([1024, 1024]);
let b = Tensor::randn([1024, 1024]);
b.iter(|| a.matmul(&b));
}
```
#### CI/CD Pipeline
```yaml
# .github/workflows/ci.yml
name: CI
on: [push, pull_request]
jobs:
test:
strategy:
matrix:
os: [ubuntu-latest, macos-latest, windows-latest]
rust: [stable, nightly]
bench:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- run: cargo bench
gpu-test:
runs-on: [self-hosted, gpu]
steps:
- run: cargo test --features cuda
```
### Deployment Targets
#### Inference Runtime
- **Native Binary**: Standalone executable
- **Dynamic Library**: `.so/.dll/.dylib`
- **Python Package**: Wheel distribution
- **WebAssembly**: Browser/Node.js runtime
- **Container**: Docker with GPU support
#### Package Distribution
```toml
# Cargo.toml
[package]
name = "rustytorch"
version = "0.1.0"
edition = "2021"
license = "Apache-2.0"
repository = "https://github.com/rustytorch/rustytorch"
[package.metadata.docs.rs]
all-features = true
rustdoc-args = ["--cfg", "docsrs"]
```
### Performance Profiling
#### Tools
- **perf**: Linux system profiler
- **flamegraph**: Visualization of hot paths
- **NVIDIA Nsight**: GPU kernel profiling
- **AMD ROCProfiler**: ROCm profiling
- **Intel VTune**: CPU optimization
#### Metrics Collection
```rust
use prometheus::{Counter, Histogram, register_counter, register_histogram};
lazy_static! {
static ref TENSOR_OPS: Counter = register_counter!(
"rustytorch_tensor_ops_total",
"Total tensor operations"
).unwrap();
static ref OP_LATENCY: Histogram = register_histogram!(
"rustytorch_op_latency_seconds",
"Operation latency in seconds"
).unwrap();
}
```
### Security Considerations
#### Memory Safety
- MIRI verification for unsafe code
- Address sanitizer in debug builds
- Fuzzing for input validation
#### Supply Chain
```toml
# Cargo.toml
[dependencies]
# Audit dependencies
cargo-audit = "0.18"
# Vendored dependencies for reproducible builds
cargo-vendor = "0.12"
```
### Documentation
#### Generation
```bash
# API documentation
cargo doc --all-features --no-deps
# Book generation
mdbook build docs/
# Python API docs
pdoc3 --html python/rustytorch
```
#### Hosting
- API docs: https://docs.rs/rustytorch
- User guide: https://rustytorch.ai/guide
- Examples: https://github.com/rustytorch/examples
### Development Workflow
#### Pre-commit Hooks
```bash
#!/bin/bash
# .git/hooks/pre-commit
# Format code
cargo fmt --all -- --check
# Lint
cargo clippy --all-targets --all-features -- -D warnings
# Test
cargo test --lib
```
#### Version Control
```gitignore
# .gitignore
target/
*.rs.bk
Cargo.lock
*.pyc
__pycache__/
.venv/
.idea/
.vscode/
*.swp
```
### External Integrations
#### Model Formats
- **ONNX**: Import/export support (Phase 7)
- **SafeTensors**: Fast tensor serialization
- **PyTorch**: `.pt` file compatibility (Phase 7)
- **DLPack**: Zero-copy interop (Phase 2)
#### Cluster Management
- **Stratoswarm**: Primary orchestration at `/home/osobh/projects/stratoswarm`
- Deployment manifests in `deploy/stratoswarm/`
- Blue/green + canary rollouts (Phase 5)
- Multi-region support (Phase 9)
- **Kubernetes**: Via Stratoswarm integration
- **SLURM**: HPC cluster integration (future)
- **Ray**: Distributed compute framework (future)
### Technical Constraints
#### Platform Limitations
- **Linux**: Full feature support
- **macOS**: Metal backend only
- **Windows**: Limited CUDA support
- **WebAssembly**: CPU-only, no threading
#### Performance Targets (Phase-Specific)
#### Phase 1 Targets
- CUDA Graph capture hit-rate: ≥ 70%
- Memory fragmentation: < 15%
- Step-time improvement: ≥ 20% vs eager baseline
- Determinism: fp32 ≤ 1e-6, bf16/fp16 ≤ 1e-3
#### Phase 2 Targets
- Graph capture hit-rate: ≥ 75%
- End-to-end step time: ≥ 20% reduction vs Phase 1
- Autograd correctness: Parity with fp32 goldens
#### Phase 3 Targets
- Single-node scaling: ≥ 0.8x efficiency (1→8 GPUs)
- Multi-node scaling: ≥ 0.7x on IB/NVLink
- FSDP memory reduction: ≥ 40% vs DP
#### Phase 4 Targets
- Auto-kernel speedup: ≥ 20-40% step-time reduction
- Inference speedup: ≥ 1.5x tokens/sec
- Kernel cache hit-rate: ≥ 80%
#### Phase 5 Targets
- Inference latency P99: < 150ms/token on RTX 5090
- KV cache hit-rate: ≥ 85%
- Continuous batching utilization: ≥ 30% improvement
#### Resource Limits
- Max tensor size: 2^63 elements
- Max dimensions: 32
- Max GPU memory: Device-dependent
- Max nodes: 10,000 (Stratoswarm limit)
## Workspace Configuration
### Excluded Crates (December 2024)
| Crate | Reason | Future Work |
|-------|--------|-------------|
| `integration_tests` | Tests reference unimplemented APIs (400+ errors) | Implement APIs when needed |
| `rtx-flash-metal-attention` | macOS/Metal only - not available on Linux | Works on macOS systems |
| `demos/ui/src-tauri` | Different MSRV and dependency requirements | Separate build process |
### Dependency Management
#### nom Version Consolidation
- **Removed**: nom 3.2.1 (legacy)
- **Active versions**: nom 7.1.3, nom 8.0.0
- **Method**: Removed unused `npy` dependency from rtx-vision-advanced
#### Float Comparison Safety (Rust 2024)
```rust
// Before (Rust 2021 - panics on NaN)
values.sort_by(|a, b| a.partial_cmp(b).unwrap());
// After (Rust 2024 - NaN-safe)
values.sort_by(|a, b| a.total_cmp(b));
```
**Pattern applied to**: 200+ files across the workspace
### Build Commands
```bash
# Standard workspace build
cargo check --workspace
# Build specific crate
cargo build -p rtx-tensor
# Full release build
cargo build --release --workspace
# Run with CUDA features
cargo build --workspace --features cuda
```
---
*Technical Context Last Updated: 2025-12-16*