Distill a repeated Jev classification task into a local model, on the fly — same answers, your hardware.
-
Updated
Oct 1, 2026 - Python
Distill a repeated Jev classification task into a local model, on the fly — same answers, your hardware.
Distil an expensive LLM API call on a narrow task into a small local model: capture traffic, curate, LoRA fine-tune with MLX on Apple Silicon, evaluate against the teacher, and serve an OpenAI-compatible cascade that escalates low-confidence requests.
Cost-aware routing across agentic pipelines. Empirical benchmarks of accuracy, cost, and dangerous rate.
Studying Minimum Sufficient Inference: when objective execution evidence can stop LLM inference without sacrificing reliability.
Adaptive Cascade Tuning via Importance Sampling
LLM serving router that prices every request in dollars and milliseconds. Semantic cache plus a confidence-gated cheap-to-GPT-4 cascade: 96.9% lower cost than always calling the strong model, within 1.6 accuracy points, median latency 450ms to sub-ms on repeated traffic. FastAPI, SQLite cost ledger, reproducible benchmark.
Ask the cheap model first. Pay for the strong one only when the cheap answer fails a check, under a budget, with a cost record for every call.
NSGA-II search framework for CIFAR-10 big/little dynamic inference cascades under embedded memory constraints.
Route each LLM call to the cheapest model that will get it right. Verifier-gated cascades with an offline threshold optimizer — 94% of the strongest model's accuracy at 34% lower cost.
Divergence-aware multi-agent routing that cuts LLM inference cost 10-100x. One signal routes queries, keys a multilingual cache, and detects hallucinations at AUC 0.90. CIKM 2026.
To associate your repository with the model-cascade topic, visit your repo's landing page and select "manage topics."