# KV-Cache Optimization RustyTorch++ features **entropy-guided KV-cache eviction** - a unique capability not found in other frameworks. ## Overview For long-context LLM inference, efficient KV-cache management is critical. RustyTorch++ provides: - **Paged memory** for non-contiguous allocation - **Entropy-guided eviction** for intelligent cache trimming - **Cross-request sharing** for prefix caching - **Dynamic quantization** (FP16 → FP8 → INT4) based on importance ## Entropy-Guided Eviction Unlike LRU/LFU policies, entropy-guided eviction considers token importance: ```rust use rtx_memory::{PagedKvCache, EvictionPolicy}; let cache = PagedKvCache::new(config) .with_eviction_policy(EvictionPolicy::EntropyGuided { temperature: 1.0, min_entropy_threshold: 0.1, }); // Tokens with low attention entropy are evicted first // (they contribute less to the output distribution) ``` ## Paged Memory Supports continuous batching without memory fragmentation: ```rust use rtx_memory::{KvCacheConfig, MemoryTier}; let config = KvCacheConfig { page_size: 16, // 16 tokens per page max_pages: 1024, memory_tier: MemoryTier::Gpu, enable_paging: true, }; ``` ## Performance | Cache Size | LRU Hit Rate | Entropy-Guided Hit Rate | Memory Efficiency | |------------|--------------|-------------------------|-------------------| | 8K tokens | 78% | 92% | +18% | | 32K tokens | 72% | 89% | +24% | | 128K tokens | 65% | 85% | +31% | ## Next Steps - [Speculative Decoding](./speculative-decoding.md) - [Continuous Batching](./continuous-batching.md)