208 lines
7.1 KiB
Markdown
208 lines
7.1 KiB
Markdown
RustyTorch++ — Phase 5 Plan (Inference Runtime & Serving Stack)
|
||
|
||
Objective (3–4 months): Ship a production‑grade inference runtime with vLLM‑class scheduling (continuous batching + paged KV cache), low‑latency decoding (speculative/assisted), quantization (int8/int4/fp8), and multi‑tenant serving. Integrate Stratoswarm manifests for deployment (staged), keep determinism/observability strong, and provide simple APIs & CLIs.
|
||
|
||
Reference paths
|
||
|
||
RUSTYTORCH_ROOT: /home/osobh/projects/rustytorch
|
||
|
||
RUSTG_COMPILER: /home/osobh/projects/rust/rustg (RTX 5090 sm_120 primary)
|
||
|
||
STRATOSWARM_ROOT: /home/osobh/projects/stratoswarm
|
||
|
||
Training/distributed work continues on the Phase 3/4 track. Phase 5 focuses on serving.
|
||
|
||
1. Scope & Non‑Goals
|
||
|
||
In scope
|
||
|
||
Inference scheduler with continuous batching, priority/SLA lanes, and admission control
|
||
|
||
Paged KV cache (GPU first; CPU/NVMe offload hooks) with eviction policies
|
||
|
||
Speculative & Assisted Decoding (draft+verify; optional smaller draft model) and early‑exit
|
||
|
||
Quantization: PTQ (int8/int4/fp8) with per‑layer error tracking; QAT hooks for later
|
||
|
||
Parallel inference: tensor parallel + pipeline parallel presets; single‑node focus (multi‑node staged)
|
||
|
||
APIs/SDKs: gRPC/HTTP streaming; simple Python & Rust client SDKs
|
||
|
||
Observability: latency histograms (p50/p95/p99), queue depth, batch merge stats, cache hit rates
|
||
|
||
Security & governance: model provenance, signed artifacts, per‑tenant isolation knobs
|
||
|
||
Stratoswarm: deployment manifests, autoscaling policies (queue‑depth & latency), staged rollouts
|
||
|
||
Out of scope (Phase 6+)
|
||
|
||
Multi‑tenant billing; full multi‑region global routing
|
||
|
||
Online learning/fine‑tuning; advanced A/B traffic shaping
|
||
|
||
2. Success Criteria (Exit / Phase‑Gate)
|
||
|
||
Latency & throughput: Achieve ≥1.5× tokens/sec vs. eager baseline, while holding p99 under target (e.g., <150 ms/token on RTX 5090 for target models).
|
||
|
||
Scheduler efficacy: Continuous batching improves GPU utilization ≥30%; queueing theory model predicts and matches live behavior within ±10%.
|
||
|
||
Cache effectiveness: KV cache hit‑rate ≥85% for steady workloads; offload path keeps p99 within 1.2× of GPU‑only at 1.5× load.
|
||
|
||
Quantization: int8/int4 reduce latency ≥25% with quality drop ≤ specified metric (e.g., Δ<0.5 ppl or task metric within budget).
|
||
|
||
Reliability: graceful degradation (brownout) under overload; zero‑dataloss restarts with warm cache rebuild.
|
||
|
||
Deployability: Stratoswarm charts support blue/green and canary; autoscaling based on queue depth + latency.
|
||
|
||
3. Architecture Work
|
||
|
||
A. Inference Scheduler
|
||
|
||
Request lanes: realtime, standard, bulk; per‑lane queue + SLAs; preemption support
|
||
|
||
Batch builder: dynamic merge window, shape‑aware; max batch size/latency budget; backpressure
|
||
|
||
Tokenizer + pre/post integrated in the same graph for zero‑copy handoff
|
||
|
||
B. KV Cache
|
||
|
||
Paged design: segment per head/layer; LRU/LFU eviction; residency (GPU ↔ CPU ↔ NVMe)
|
||
|
||
Memory accounting: bytes per token/head; pressure signals
|
||
|
||
Prefetch & write‑back policies; GPUDirect Storage hooks (future)
|
||
|
||
C. Decoding
|
||
|
||
Speculative (draft+verify) with dynamic acceptance; assisted (guided tokens) hooks
|
||
|
||
Early‑exit strategies; nucleus/top‑k sampling; temperature per request
|
||
|
||
D. Quantization
|
||
|
||
PTQ pipelines: per‑layer calibration, error metrics; formats for int8/int4/fp8
|
||
|
||
Runtime dequant kernels; mixed‑precision scheduler
|
||
|
||
E. Parallel Inference
|
||
|
||
Tensor parallel for large layers; pipeline stages for deep models; activation shipping policies
|
||
|
||
Single‑node first; multi‑node design doc and stubs
|
||
|
||
F. APIs & SDKs
|
||
|
||
gRPC/HTTP streaming (/generate, /embed); cancellation; timeouts; tracing headers
|
||
|
||
Rust/Python clients; auth hooks; rate limiting per tenant
|
||
|
||
G. Deployment (Stratoswarm)
|
||
|
||
Charts: replicas, resource requests/limits, node selectors (GPU), anti‑affinity
|
||
|
||
HPA policies: queue_depth + p95_latency
|
||
|
||
Blue/green + canary; rollback automation; secrets via external store
|
||
|
||
4. Benchmarks & Test Matrix
|
||
|
||
Workloads
|
||
|
||
LLMs: GPT‑class 7B–13B; sequence 1k–8k; batch 1–N
|
||
|
||
Embedding models (optional) for throughput testing
|
||
|
||
Metrics
|
||
|
||
tokens/sec, p50/p95/p99 latency; queue times; batch size distribution; cache hit‑rate; GPU util; mem pressure; OOM count
|
||
|
||
Scenarios
|
||
|
||
Steady state, spiky arrivals, long prompts, short prompts, multi‑tenant mixed SLAs, overload/brownout
|
||
|
||
5. CI Gates (Phase 5)
|
||
|
||
Latency Gate: p99 under SLA for standard lane at target QPS; realtime lane honored
|
||
|
||
Throughput Gate: ≥1.5× tokens/sec vs. eager baseline on 5090
|
||
|
||
Cache Gate: hit‑rate ≥85% on steady workloads
|
||
|
||
Reliability Gate: restart with warm cache within budget; overload degrades gracefully (no errors)
|
||
|
||
Security Gate: signed models; SBOM; no secrets in logs
|
||
|
||
6. Observability & Telemetry
|
||
|
||
Traces: per‑request timeline (enqueue → batch → decode → stream)
|
||
|
||
Metrics: queue_depth, merge_wait_ms, batch_size, tokens_per_s, lat_p50/p95/p99, kv_hit_pct, gpu_util, alloc_frag_pct
|
||
|
||
Dashboards: SLA lanes, cache tiers, autoscaling behavior, brownout triggers
|
||
|
||
7. Deliverables
|
||
|
||
rtx-serve crate: scheduler, KV cache, decoding, quantization hooks
|
||
|
||
Parallel inference presets (TP/PP) and configs
|
||
|
||
gRPC/HTTP server; Rust/Python client SDKs; examples
|
||
|
||
Benchmarks & dashboards; CI gates
|
||
|
||
Stratoswarm Helm charts + runbooks (blue/green, canary, autoscaling)
|
||
|
||
Documentation: docs/serving.md (APIs, tuning, SLOs, deployment)
|
||
|
||
8. RACI — Phase 5 Agent Ownership
|
||
|
||
Conductor (Orchestrator) — Accountable: delivery, cross‑agent planning
|
||
|
||
Inference Scheduler Agent — Responsible: lanes, batch builder, brownout strategies
|
||
|
||
Memory Planner Agent — Responsible: KV paging, residency policies, memory pressure signals
|
||
|
||
Kernel Synthesizer & Auto‑Tuner — Consulted: dequant kernels, fused epilogues
|
||
|
||
Performance Engineer — Accountable: latency/throughput benches; cache hit‑rate analysis; gates
|
||
|
||
Security Engineer & Auditor — Accountable: provenance, signing, secrets hygiene
|
||
|
||
DevOps Engineer — Responsible: Stratoswarm charts, autoscaling, rollouts
|
||
|
||
SRE Engineer — Responsible: SLOs, alerting, runbooks, chaos drills
|
||
|
||
Agent Organizer — Accountable: tool allowlists, routing, freeze on gate failure
|
||
|
||
9. Risks & Mitigations
|
||
|
||
Queue collapse / tail latency → lane isolation; max queue depth; dynamic merge windows; preemption
|
||
|
||
Cache thrash → admission control; long‑prompt handling; tiered cache with hysteresis
|
||
|
||
Quantization quality loss → per‑layer error budgets; auto fallback to higher precision
|
||
|
||
Speculative instability → acceptance threshold tuning; abort on drift; fallback to baseline
|
||
|
||
Autoscaling oscillations → damped controllers; combined signals (queue + p95)
|
||
|
||
10. Timeline (suggested)
|
||
|
||
Weeks 1–2: Scheduler lanes + batch builder; gRPC/HTTP skeleton; metrics schema
|
||
|
||
Weeks 3–4: KV paging + cache hit‑rate instrumentation; quantization hooks; dequant kernels
|
||
|
||
Weeks 5–6: Speculative decoding; TP/PP presets; benchmarks; dashboards
|
||
|
||
Weeks 7–8: Stratoswarm charts; chaos/brownout drills; CI gates; docs; phase‑gate review
|
||
|
||
11. Phase‑6 Handover Seeds
|
||
|
||
Multi‑tenant quota/billing; global routing and edge POPs
|
||
|
||
Online quantization & precision adaptation under SLAs
|
||
|
||
Cross‑node inference (tensor+pipeline) and shared KV cache via RDMA/GDS
|
||
|
||
— End of Phase 5 Plan —
|