i build the systems behind real-time AI products - inference stacks, backend services, and the infrastructure that keeps them fast and reliable under load.
at june labs, i built a voice AI platform end to end and took its STT + LLM + TTS pipeline from 1.5s to 366ms e2e latency on a single H100 by replacing network hops with shared-memory IPC.
before that, i worked on multi-cloud Kubernetes and GPU deployment infrastructure at DynamoAI, including deployments for Lenovo and PayPal. earlier, i worked on Kubernetes, Terraform, and GitOps systems at Bizongo.
currently building June, an open-source AI workspace for legal teams, and Lisn, local voice dictation for Linux.
i tend to work on problems where latency, reliability, deployment, and product constraints all meet.
stack: Python · Go · TypeScript · Kubernetes · PostgreSQL · Redis · vLLM · TensorRT-LLM



