Relevant Experience & Education Highlights
Director-level AI software engineering leader with extensive experience owning GPU AI stacks including PyTorch, vLLM, and inference engines at a major semiconductor company, strongly positioned to lead Modular's GenAI Enterprise team.
1. Serve as the primary engineering point of contact for enterprise prospects and customers
Acted as technical owner and support lead for strategic AI clients and key server GPU customers in deep learning applications.
- Technical owner of strategic AI client accounts globally, collaborating on GPU AI software solutions.
- Provided technical support for key customers of server GPUs, interfacing on DL applications and system configurations.
- Coordinated issues and feature planning for frameworks like TensorFlow-ROCm with external partners.
2. 3+ years of management experience with a proven track record of building high-performing collaborative teams
Directed and managed AI software engineering teams focused on GPU AI frameworks, inference, and ecosystems at a leading semiconductor company.
- Directed AI Software Engineering team owning GPU AI software including PyTorch, TensorFlow, JAX, ONNX-Runtime, and MIGraphX.
- Led Sr. Director role overseeing ROCm GPU AI layer with LLM inference engines like vLLM, SGLang, MIGraphX, and TGI.
- Managed deep learning frameworks team to enable and optimize TensorFlow, PyTorch, and ONNX-RT on AMD GPUs using ROCm.
- Built teams contributing to AI ecosystems including OpenAI Triton, DeepSpeed, CuPy, and rocMLIR graph compiler.
3. Hands-on experience with production AI workloads including scaling and serving large models
Owned development and optimization of production GPU inference engines, frameworks, and operators for LLMs and GenAI at scale.
- Owned LLM and Generative AI GPU inference engines including vLLM, SGLang, MIGraphX, TGI, and HuggingFace open-source models.
- Developed high-performance AI operators for LLMs in Triton, CK, and ASM, supporting production AITER project kernels.
- Led AI models team focusing on HuggingFace and open-source LLMs with public MAD automation tooling.
- Optimized key DL models on data center GPUs and contributed to TensorFlow-ROCm for production releases.
4. Proficiency in Mojo, Python, C++, or CUDA programming languages, plus familiarity with frameworks such as MAX, Pytorch, or vLLM
Demonstrated deep expertise in C++, Python, PyTorch, vLLM, and CUDA-aligned GPU programming for AI workloads.
- Expertise in C++ and Python for GPU AI software development across ROCm stack.
- Led PyTorch, JAX, TensorFlow, and ONNX-Runtime support with ROCm/MIGraphX execution providers.
- Integrated vLLM and other SOTA inference solutions for GenAI GPU serving.
- Developed HIP-dependent workloads supporting TensorFlow, Caffe, and PaddlePaddle ecosystems.
Requirements & Candidate Alignment
| Modular Requirement | Candidate Qualification |
|---|---|
| 7+ Years of experience in a customer-facing technical roles where you solved hard problems alongside clients | 16.6 years total experience including technical ownership of global strategic AI clients and support for key DL GPU customers |
| 3+ Years of management experience with a proven track record of building high-performing collaborative teams | Director and Sr. Director roles leading GPU AI software teams on frameworks, inference, and ecosystems |
| In-depth understanding of the GenAI execution lifecycle, deployment methodologies, along with testing and quality evaluation methodologies | Owned full GenAI stack including rocMLIR graph compiler, MIGraphX inference, vLLM serving, and PyTorch foundation contributions |
| Hands-on experience with production AI workloads including scaling and serving large models | Led production LLM inference with vLLM, SGLang, TGI, HuggingFace models, and high-performance Triton operators |
| Proficiency in Mojo, Python, C++, or CUDA programming languages, plus familiarity with frameworks such as MAX, Pytorch, or vLLM | Proficient in C++, Python, PyTorch, vLLM, and CUDA workloads across ROCm GPU AI software |
| Proven ability to translate complex technical concepts into clear, concise, and understandable language that resonates with for both clients and internal teams | Served as technical representative for PyTorch and OpenXLA foundations, coordinating client issues and feature planning |
| Demonstrated ability to rapidly prototype and deliver Proofs of Concept (POCs) that win customer confidence | Optimized and enabled DL frameworks and models on GPUs, delivering production-ready ROCm features and client demos |
If you would like to discuss this candidate or other critical roles, here is a link to my calendar to schedule a call:
Schedule a CallContact:
Finding the signal in the noise since 2005