3.3 KiB
Phase 13: Classical ML Superset — Making scikit‑learn Obsolete
Executive Summary
While RustyTorch++ already excels at deep learning, distributed training, inference, and agentic optimization, scikit-learn still dominates classical ML workflows with its breadth, usability, and ecosystem integration. Phase 13 bridges that gap—adding full coverage of classical algorithms, preprocessing, validation, and Python-native APIs. The goal: make scikit-learn obsolete by offering GPU-native, Rust-safe, high-performance, agent-augmented classical ML within the RustyTorch++ stack.
What We Already Offer (Aligned with scikit-learn)
Classical algorithms like linear/logistic regression, k-means, SVMs, and clustering through existing modules—already GPU-capable and faster than Python implementations.
Pipeline construction via rtx-graph.
Distributed AutoML support through rtx-evolve.
Model serialization, governance, and deployment pipelines.
What’s Missing (scikit-learn still wins on these)
Full algorithm coverage:
Implement: Decision Trees, Random Forests, Gradient Boosted Trees, Naive Bayes, KNN, ElasticNet, Gaussian Processes, HMMs, Association Rule Mining.
Complete preprocessing suite:
Add: Scalers (Standard, MinMax, Robust), Encoders (OneHot, Label), Transformers (PolynomialFeatures, Imputation, Feature Selection).
Model selection utilities:
Add: Grid/Random Search, Bayesian Optimization, Cross-validation strategies (KFold, StratifiedKFold, Time Series), standard metrics (precision, recall, F1, R²).
Python ecosystem interop:
Provide drop-in APIs and seamless NumPy/pandas compatibility.
Documentation API parity:
Replicate scikit-learn’s style and usability for easy migration.
Phase 13 Deliverables
Crates & Modules
rtx-ml-classic: Implements full algorithm suite (trees, ensembles, Bayesian methods, KNN, etc.), GPU or CPU enabled.
rtx-preprocessing: Parity with scikit-learn’s transformers.
rtx-validation: CV strategies, hyperparameter search, standard metrics.
rtx-sklearn-py: Python wrappers with scikit-learn-like API and DLPack/NumPy bridging.
rtx-docs-sklearn: Documentation modeled after scikit-learn’s structure and clarity.
rtx-automeasure: Agentic model selection and AutoML on classical side.
Exit Tests / KPIs
Algorithm coverage ≥95% of scikit-learn v1.x APIs.
Preprocessing transformers match scikit-learn behavior.
GridSearchCV parity with vectorization and GPU speedup (10–50× faster).
API tests: Python migration example works unchanged (e.g., same syntax).
Agent AutoML yields models within 5% of hand-tuned scikit-learn baselines using less compute.
Phase 13 Roadmap
Build rtx-ml-classic with core to advanced algorithms.
Add rtx-preprocessing and validation modules.
Create Python binding crate rtx-sklearn-py.
Draft docs with usage examples & scikit-like style.
Integrate AutoML agent support for classical ML.
Match coverage, then measure performance improvements & developer feedback.
Migration Goal
Enable drop-in replacement for scikit-learn code:
Before (scikit-learn):
from sklearn.ensemble import RandomForestClassifier clf = RandomForestClassifier(n_estimators=100).fit(X, y)
After (RustyTorch++ Phase 13):
from rustytorch_ml import RandomForestClassifier clf = RandomForestClassifier(n_estimators=100).fit(X, y)
Outcome:
Same syntax.
Instant GPU acceleration.
Memory safety and Rust-native performance.