ML Engineer — LLM quantization & inference on consumer hardware
Dallas–Fort Worth, TX · Portfolio · Hugging Face · LinkedIn · ttimmsinternational@gmail.com
I quantize and serve large models on hardware that isn't supposed to run them — 16 GB Blackwell GPUs, Jetson edge boards — and I reproduce every published number against its confidence interval before I call it done.
15 models on Hugging Face · ~7,800 downloads/month (Sept 2026).
Open to ML Engineer roles (inference optimization, model compression, edge deployment) — DFW or remote.
Bible AI Assistant — Local Scripture Q&A on a 16 GB card: hybrid RAG over 31k verses feeding an SFT fine-tune, with a verifiable-citation reward (cited verse must exist in the index, quote must match) scaffolded for a GRPO stage. sha256-pinned benchmark protocol, 476 tests, full CI/CD. Primary project — building toward local SOTA on a 5070 Ti.
MoE Pruning + NVFP4 — A 50%-expert-pruned MoE coder model, quantized to fit 16 GB VRAM. SWE-bench Verified 52.0% (26/50, officially graded), HumanEval+/MBPP+ reproduced inside published confidence intervals, CI-checked reproduction pipeline. Model on Hugging Face →
Ornith REAP-50 + NVFP4 — Same pipeline on a different 35B-A3B MoE base (MIT): 256→128 experts, MTP head and vision tower stripped, GPTQ-NVFP4A16 — 12.47 GiB. SWE-bench Verified 44.0% (22/50, official harness, same 50-instance slice as the KAT run), HumanEval+ 84.2% / MBPP+ 89.2%. Three upstream vLLM gaps for this architecture patched locally. Model on Hugging Face →
ZAYA1 NVFP4 W4A4 — 4-bit weights and activations on native Blackwell tensor cores: 9.5 tok/s single-stream from a 6.02 GB checkpoint, 2,400+ combined downloads on Hugging Face (Sept 2026). Includes a benchmark I retracted and corrected in public once I found the CUDA-graph path corrupting output. Model on Hugging Face →
Godspeed Coding Agent — A coding agent built from scratch: deny-first permission engine, SHA-256 hash-chained audit trail. SWE-bench Lite 34.8% single-shot / 52.2% oracle best-of-5, $0 API spend.
Sovereign Edge — Five-agent personal AI system running entirely on a Jetson Orin Nano — zero cloud dependencies.
Also: Manna Trading (multi-agent trading pipeline) · an open llama.cpp PR fixing an NVFP4 quantizer crash
Python · PyTorch · vLLM / CUTLASS · TRL / Unsloth · NVFP4 · GGUF · GPTQ / AWQ · CUDA · Docker



