feat(inference): concrete EAGLE draft model + real tokenizer at the serving boundary
GPU Tests / Check GPU Availability (push) Successful in 1s
CI / Build (ubuntu-latest) (push) Failing after 6s
CI / Build (macos-latest) (push) Failing after 11s
Documentation / Build User Guide (push) Successful in 7s
CI / Clippy Check (push) Failing after 21s
CI / Format Check (push) Failing after 6s
GPU Tests / CUDA Tests (11.8) (push) Has been skipped
GPU Tests / CUDA Tests (12.1) (push) Has been skipped
CI / Build CPU-Only (Explicit) (push) Failing after 6s
GPU Tests / Metal Tests (push) Has been skipped
CI / Test (macos-latest) (push) Has been skipped
CI / Test (ubuntu-latest) (push) Has been skipped
CI / Python Bindings (maturin) (macos-latest) (push) Has been skipped
CI / Python Bindings (maturin) (ubuntu-latest) (push) Has been skipped
CI / WASM Build + Size Check (push) Has been skipped
CI / Distributed Training Tests (push) Has been skipped
CI / CI Success (push) Failing after 0s
Performance Benchmarks / Run Benchmarks (push) Successful in 28s
Documentation / Build API Documentation (push) Failing after 25s

EAGLE (rtx-inference/src/eagle.rs, ~610 lines, mirrors medusa.rs
conventions): EagleDraftHead autoregressive FFN with Concat/Add/
Attention feature fusion, EagleHeads draft model with draft/
draft_steps (per-step top-k for candidate trees) and teacher-forced
training_loss; implements the speculative::EagleDraftModel trait so it
plugs into the orchestration layer. 38 unit tests.

Tokenizer (rtx-inference/src/tokenizer.rs): ServingTokenizer enum —
Vocab (HuggingFace tokenizers, loadable from tokenizer.json) or
ByteLevel fallback preserving previous behavior. rtx-serving-api's
AppState and rtx-streaming's token generator now encode/decode through
it (with_engine_and_tokenizer / set_tokenizer added; existing
signatures unchanged). Also fixes two pre-existing compile errors in
rtx-streaming (missing import, stray .await) that blocked its lib
tests entirely.

Tests: rtx-inference 328 pass, rtx-serving-api 193 pass, rtx-streaming
53 pass (2 pre-existing mock-server connection failures unrelated to
these changes).

Co-Authored-By: Claude Fable 5 <[email protected]>
This commit is contained in:
osobh
2026-07-09 21:54:05 -07:00
co-authored by Claude Fable 5
parent 0cbfc1a739
commit 733b02cd8b
9 changed files with 1124 additions and 21 deletions
@@ -47,16 +47,33 @@ impl ServingServer {
}
}
/// Create a serving server backed by a live inference engine
/// Create a serving server backed by a live inference engine.
///
/// Uses byte-level tokenization by default; use
/// [`ServingServer::with_engine_and_tokenizer`] to attach a real
/// vocabulary tokenizer.
#[must_use]
pub fn with_engine(
config: ServerConfig,
engine: std::sync::Arc<tokio::sync::RwLock<rtx_inference::InferenceEngine>>,
) -> Self {
Self::with_engine_and_tokenizer(config, engine, rtx_inference::ServingTokenizer::default())
}
/// Create a serving server backed by a live inference engine and an
/// explicit tokenizer (e.g. loaded via
/// `rtx_inference::ServingTokenizer::from_file`).
#[must_use]
pub fn with_engine_and_tokenizer(
config: ServerConfig,
engine: std::sync::Arc<tokio::sync::RwLock<rtx_inference::InferenceEngine>>,
tokenizer: rtx_inference::ServingTokenizer,
) -> Self {
Self {
config,
state: inference::AppState {
engine: Some(engine),
tokenizer,
},
}
}