Initial commit
This commit is contained in:
@@ -0,0 +1,202 @@
|
||||
# Architecture Overview
|
||||
|
||||
RustyTorch++ is a modular ML framework with 56 specialized crates organized for separation of concerns.
|
||||
|
||||
## Design Philosophy
|
||||
|
||||
### 1. GPU-Native First
|
||||
|
||||
Unlike frameworks that wrap CPU-first libraries, RustyTorch++ is designed from the ground up for GPU computation:
|
||||
|
||||
- Memory allocated directly on GPU
|
||||
- Operations dispatch to GPU kernels by default
|
||||
- Zero-copy data transfers where possible
|
||||
- Automatic kernel fusion for common patterns
|
||||
|
||||
### 2. Memory Safety
|
||||
|
||||
Rust's ownership model eliminates common GPU programming errors:
|
||||
|
||||
- No double-free of GPU memory
|
||||
- No use-after-free tensor operations
|
||||
- Compile-time shape validation where possible
|
||||
- Thread-safe concurrent GPU access
|
||||
|
||||
### 3. Production Ready
|
||||
|
||||
Built-in infrastructure for deployment:
|
||||
|
||||
- HTTP/gRPC serving APIs
|
||||
- Monitoring and telemetry
|
||||
- Distributed training support
|
||||
- Model versioning and A/B testing
|
||||
|
||||
## Crate Architecture
|
||||
|
||||
```
|
||||
┌─────────────────────────────────────────────────────────────────┐
|
||||
│ Meta Crates │
|
||||
│ rtx │ rtx-core │ rtx-training │ rtx-inference-stack │
|
||||
└───────────────────────────────────────────────────────────────────┘
|
||||
│ │ │
|
||||
▼ ▼ ▼
|
||||
┌─────────────────┐ ┌─────────────────┐ ┌─────────────────────────┐
|
||||
│ Core (11) │ │ Training (11) │ │ Production (6) │
|
||||
│ ─────────────── │ │ ─────────────── │ │ ───────────────────────│
|
||||
│ rtx-tensor │ │ rtx-transformers│ │ rtx-serving-api │
|
||||
│ rtx-runtime │ │ rtx-distributed │ │ rtx-inference │
|
||||
│ rtx-autograd │ │ rtx-flash-attn │ │ rtx-streaming │
|
||||
│ rtx-memory │ │ rtx-rl │ │ rtx-hub │
|
||||
│ rtx-kernel │ │ rtx-compress │ │ rtx-config │
|
||||
│ rtx-nn │ │ rtx-federated │ │ rtx-monitoring │
|
||||
│ rtx-graph │ │ ... │ │ │
|
||||
│ ... │ │ │ │ │
|
||||
└─────────────────┘ └─────────────────┘ └─────────────────────────┘
|
||||
│ │ │
|
||||
└────────────────────┼─────────────────────┘
|
||||
▼
|
||||
┌─────────────────────────────────────────────────────────────────┐
|
||||
│ Models (7) │
|
||||
│ rtx-vision │ rtx-nlg │ rtx-multimodal │ rtx-diffuse │
|
||||
└─────────────────────────────────────────────────────────────────┘
|
||||
│
|
||||
▼
|
||||
┌─────────────────────────────────────────────────────────────────┐
|
||||
│ Specialized & Data (14) │
|
||||
│ rtx-science │ rtx-geom │ rtx-cfd │ rtx-etl │ rtx-feature-store │
|
||||
└─────────────────────────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
## Core Infrastructure
|
||||
|
||||
### Tensor System
|
||||
|
||||
The `rtx-tensor` crate provides the foundational tensor type:
|
||||
|
||||
```rust
|
||||
// GPU-backed tensor with automatic memory management
|
||||
let tensor = Tensor::randn([1024, 1024], &Device::cuda(0)?)?;
|
||||
|
||||
// Zero-copy views
|
||||
let slice = tensor.narrow(0, 0, 512)?; // No memory copy
|
||||
|
||||
// Automatic differentiation ready
|
||||
let tracked = tensor.requires_grad(true)?;
|
||||
```
|
||||
|
||||
### Runtime System
|
||||
|
||||
The `rtx-runtime` crate manages GPU resources:
|
||||
|
||||
- Device discovery and selection
|
||||
- Memory allocation with pooling
|
||||
- Stream scheduling for concurrency
|
||||
- Kernel compilation and caching
|
||||
|
||||
### Autograd Engine
|
||||
|
||||
The `rtx-autograd` crate implements automatic differentiation:
|
||||
|
||||
- Tape-based gradient recording
|
||||
- Topological backward pass
|
||||
- Gradient checkpointing for memory efficiency
|
||||
- Higher-order derivatives (Hessian, Jacobian)
|
||||
|
||||
## Data Flow
|
||||
|
||||
```
|
||||
Input Data
|
||||
│
|
||||
▼
|
||||
┌─────────────────┐
|
||||
│ Preprocessing │ ← rtx-preprocessing
|
||||
│ (Transforms) │
|
||||
└────────┬────────┘
|
||||
│
|
||||
▼
|
||||
┌─────────────────┐
|
||||
│ Tensor │ ← rtx-tensor
|
||||
│ (GPU Memory) │
|
||||
└────────┬────────┘
|
||||
│
|
||||
▼
|
||||
┌─────────────────┐
|
||||
│ Model │ ← rtx-nn, rtx-transformers
|
||||
│ (Forward) │
|
||||
└────────┬────────┘
|
||||
│
|
||||
▼
|
||||
┌─────────────────┐
|
||||
│ Loss │ ← rtx-losses
|
||||
│ (Objective) │
|
||||
└────────┬────────┘
|
||||
│
|
||||
▼
|
||||
┌─────────────────┐
|
||||
│ Backward │ ← rtx-autograd
|
||||
│ (Gradients) │
|
||||
└────────┬────────┘
|
||||
│
|
||||
▼
|
||||
┌─────────────────┐
|
||||
│ Optimizer │ ← rtx-transformers (optimizers)
|
||||
│ (Update) │
|
||||
└─────────────────┘
|
||||
```
|
||||
|
||||
## Memory Model
|
||||
|
||||
### GPU Memory Hierarchy
|
||||
|
||||
```
|
||||
┌─────────────────────────────────────┐
|
||||
│ Global Memory (HBM) │ 48GB+ on high-end GPUs
|
||||
│ ┌─────────────────────────────┐ │
|
||||
│ │ Tensor Pool (Cached) │ │ Pre-allocated blocks
|
||||
│ │ ┌───────────────────────┐ │ │
|
||||
│ │ │ Active Tensors │ │ │ Currently in use
|
||||
│ │ └───────────────────────┘ │ │
|
||||
│ │ ┌───────────────────────┐ │ │
|
||||
│ │ │ Gradient Storage │ │ │ Accumulated gradients
|
||||
│ │ └───────────────────────┘ │ │
|
||||
│ └─────────────────────────────┘ │
|
||||
└─────────────────────────────────────┘
|
||||
│
|
||||
▼
|
||||
┌─────────────────────────────────────┐
|
||||
│ Shared Memory (SRAM) │ Per-block cache
|
||||
│ L2 Cache │ GPU-wide cache
|
||||
└─────────────────────────────────────┘
|
||||
```
|
||||
|
||||
### Memory Management Strategy
|
||||
|
||||
1. **Tensor Pool**: Pre-allocate and reuse GPU memory
|
||||
2. **Lazy Allocation**: Allocate on first use
|
||||
3. **Memory Mapping**: Use pinned host memory for efficient transfers
|
||||
4. **Gradient Checkpointing**: Trade compute for memory in training
|
||||
|
||||
## Thread Safety
|
||||
|
||||
All public types are `Send + Sync` where appropriate:
|
||||
|
||||
```rust
|
||||
// Share device handle across threads
|
||||
let device = Arc::new(Device::cuda(0)?);
|
||||
|
||||
// Parallel data loading
|
||||
let handles: Vec<_> = (0..4)
|
||||
.map(|i| {
|
||||
let device = Arc::clone(&device);
|
||||
std::thread::spawn(move || {
|
||||
Tensor::randn([1000, 1000], &device)
|
||||
})
|
||||
})
|
||||
.collect();
|
||||
```
|
||||
|
||||
## Next Steps
|
||||
|
||||
- [Crate Organization](./crate-organization.md) - Detailed crate breakdown
|
||||
- [Tensor System](./tensor-system.md) - Deep dive into tensors
|
||||
- [Autograd Engine](./autograd-engine.md) - How gradients work
|
||||
Reference in New Issue
Block a user