933 B
933 B
Ring Attention
RustyTorch++ supports Ring Attention for 16M+ token contexts across multiple GPUs.
Overview
Ring Attention enables arbitrarily long sequences by:
- Partitioning sequences across GPUs
- Ring-communicating KV blocks between devices
- Computing local attention with accumulated context
Usage
use rtx_transformers::RingAttention;
let ring = RingAttention::new(config)
.with_num_devices(8)
.with_chunk_size(16384); // 16K tokens per device
// Handles 8 x 16K = 128K tokens seamlessly
let output = ring.forward(&query, &key, &value).await?;
Scaling Efficiency
| Total Tokens | Devices | Efficiency | Communication Overhead |
|---|---|---|---|
| 128K | 2 | 95% | 5% |
| 512K | 4 | 91% | 9% |
| 2M | 8 | 86% | 14% |