AI systems engineer focused on verifiable LLM inference, computer-vision deployment, model compression, and agent reliability. I work as an AI Engineer at Multiverse Computing, turning research ideas into measured, reproducible systems for real hardware.
- ProvenanceGuard — source-aware factuality verification for MCP-based LLM agents, published on arXiv.
- LibreYOLO — open-source object-detection library; I contributed the merged Core ML export implementation.
- Agent Knowledge Vault — local-first, redacted knowledge capture for preserving Codex CLI and Claude Code session knowledge in Obsidian.
- TurboQuant CPU — CPU-focused TurboQuant implementation work with modified
llama.cppbuilds and reproducible benchmark artifacts for x86-64 Linux and Apple silicon.
- Vision BuildKit — provider-neutral tooling for building, validating, and packaging computer-vision models. Its evidence records distinguish verified deployments from failed or incomplete backend gates across ONNX Runtime, OpenVINO, TensorRT, and Axelera.
- Fused Memory — scoped multi-signal retrieval for long-term agents, evaluated through 61,600 paired retrieval cases with reproducible benchmark artifacts.
- iOS calorie and exercise tracker (in development) — currently designing a native iOS app for calorie intake and exercise tracking.
- DSpark Speed Lab (most recent project) — currently improving DSpark inference speed through exactness-first speculative-decoding research with frozen baselines, offline replay, matched runtime gates, and reproducible experiment protocols.
- ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents, 2026.
- Scaling Laws for Energy Efficiency of Local LLMs, 2025.
- LLM inference, speculative decoding, and agent systems
- Computer-vision deployment and hardware-aware validation
- Model compression, quantization, and on-device AI
- Reproducible benchmarks and evidence-gated engineering
Python, PyTorch, Swift, C++, ONNX Runtime, OpenVINO, TensorRT, Core ML, llama.cpp, and ML experiment tooling.

