feat(demos,inference): wire simulation demos to real compute; fix embedding lookup and weight-name aliases
GPU Tests / Check GPU Availability (push) Successful in 0s
GPU Tests / Metal Tests (push) Has been skipped
CI / Format Check (push) Failing after 6s
CI / Clippy Check (push) Failing after 7s
Performance Benchmarks / Run Benchmarks (push) Failing after 7s
GPU Tests / CUDA Tests (11.8) (push) Has been skipped
GPU Tests / CUDA Tests (12.1) (push) Has been skipped
CI / Build (ubuntu-latest) (push) Failing after 7s
Documentation / Build User Guide (push) Successful in 8s
CI / Build (macos-latest) (push) Failing after 9s
CI / Test (macos-latest) (push) Has been skipped
CI / Test (ubuntu-latest) (push) Has been skipped
CI / Python Bindings (maturin) (macos-latest) (push) Has been skipped
CI / Python Bindings (maturin) (ubuntu-latest) (push) Has been skipped
CI / WASM Build + Size Check (push) Has been skipped
CI / Distributed Training Tests (push) Has been skipped
CI / Build CPU-Only (Explicit) (push) Failing after 43s
Documentation / Build API Documentation (push) Failing after 48s
CI / CI Success (push) Failing after 0s
GPU Tests / Check GPU Availability (push) Successful in 0s
GPU Tests / Metal Tests (push) Has been skipped
CI / Format Check (push) Failing after 6s
CI / Clippy Check (push) Failing after 7s
Performance Benchmarks / Run Benchmarks (push) Failing after 7s
GPU Tests / CUDA Tests (11.8) (push) Has been skipped
GPU Tests / CUDA Tests (12.1) (push) Has been skipped
CI / Build (ubuntu-latest) (push) Failing after 7s
Documentation / Build User Guide (push) Successful in 8s
CI / Build (macos-latest) (push) Failing after 9s
CI / Test (macos-latest) (push) Has been skipped
CI / Test (ubuntu-latest) (push) Has been skipped
CI / Python Bindings (maturin) (macos-latest) (push) Has been skipped
CI / Python Bindings (maturin) (ubuntu-latest) (push) Has been skipped
CI / WASM Build + Size Check (push) Has been skipped
CI / Distributed Training Tests (push) Has been skipped
CI / Build CPU-Only (Explicit) (push) Failing after 43s
Documentation / Build API Documentation (push) Failing after 48s
CI / CI Success (push) Failing after 0s
Demos: - rtx-distllm-demo: real rtx-tensor weights per shard, real scaled-dot-product attention forward, metrics measured (Instant) instead of hardcoded constants; network topology remains a documented simulation fed by real tensor byte sizes. - rtx-model-zoo: MockInferenceEngine deleted; RealInferenceEngine loads a tiny real transformer into rtx_inference::InferenceEngine and runs genuine engine.infer per request; domain outputs are explicitly- labeled toy proxies derived from real output tokens. - rtx-inference-profiler: mock models deleted; profiles real matmul/softmax pipelines on rtx-tensor with measured latency/memory. Inference-path bugs the demos surfaced (fixed here): - ForwardPass::apply_embedding misused Tensor::gather for the embedding lookup — gather returns the indices' shape, silently dropping the hidden dim and breaking every downstream broadcast. Now uses the existing Tensor::embedding_lookup ([vocab,hidden] x [batch,seq] -> [batch,seq,hidden]). - Attention weight lookup accepts both self_attn. (HF-LLaMA) and attention. prefixes; final layer norm accepts norm.weight / model.norm.weight / ln_f.weight aliases. - Integration fixture gains the final norm weight; the previously always-failing engine tests now pass (8/8 model_loading_test). End-to-end inference through the real engine now works for the first time — verified via model_zoo_demo producing real forward-pass outputs across all categories. Co-Authored-By: Claude Fable 5 <[email protected]>
This commit is contained in:
@@ -12,6 +12,36 @@ use distllm_shared::{
|
||||
// Model Configurations
|
||||
// ============================================================================
|
||||
|
||||
/// A small model configuration used for **real tensor compute** demo runs.
|
||||
///
|
||||
/// The larger configs below (`llama_70b_config`, `llama_405b_config`, ...)
|
||||
/// describe genuinely trillion-parameter-scale models: their weight tensors
|
||||
/// are far too large to actually allocate (would require hundreds of GB of
|
||||
/// RAM per shard). They remain useful for cluster-planning /
|
||||
/// memory-estimation code paths that only read `ModelConfig` fields and
|
||||
/// never allocate real weights. Anywhere this demo performs *real*
|
||||
/// `Tensor::randn` weight allocation and real attention compute
|
||||
/// (`ModelShard::load`, `DistributedLLM::generate`, `run_demo`), this small
|
||||
/// config (or something similarly sized) should be used instead, so the
|
||||
/// real compute stays representative without exhausting host memory.
|
||||
#[must_use]
|
||||
pub fn tiny_realcompute_config() -> ModelConfig {
|
||||
ModelConfig {
|
||||
name: "tiny-realcompute-demo".to_string(),
|
||||
num_params: 0.05,
|
||||
num_layers: 8,
|
||||
hidden_dim: 256,
|
||||
num_heads: 8,
|
||||
num_kv_heads: 8,
|
||||
intermediate_dim: 512,
|
||||
vocab_size: 32000,
|
||||
max_seq_len: 4096,
|
||||
head_dim: 32,
|
||||
rope_theta: 10000.0,
|
||||
dtype: DataType::Float32,
|
||||
}
|
||||
}
|
||||
|
||||
/// LLaMA 7B model configuration.
|
||||
#[must_use]
|
||||
pub fn llama_7b_config() -> ModelConfig {
|
||||
|
||||
Reference in New Issue
Block a user