Focused on ML Systems and Inference Infrastructure — the layer between research and production.
Currently building foundational depth in:
- LLM serving engines (vLLM, Triton, TGI)
- Kubernetes and container orchestration
- Systems programming and Linux internals
Actively working toward:
- First open source contribution to vLLM
- ML Systems Architect trajectory (entry via Inference Infra Engineer)
Nothing shipped yet — but the foundation is being laid deliberately. Check back in 60 days.
- PagedAttention paper (Kwon et al., 2023)
- Designing Data-Intensive Applications — Kleppmann
- Linux internals + CUDA memory hierarchy
I care about systems that serve models at scale — not just models themselves.