Candidate Profile — Candidate #3
Position: Modular — Engineering Director, GenAI Enterprise (Remote)
Candidate Location: Austin, TX Metro
Experience: 16.6 years experience

Relevant Experience & Education Highlights

Director-level AI software engineering leader with extensive experience owning GPU AI stacks including PyTorch, vLLM, and inference engines at a major semiconductor company, strongly positioned to lead Modular's GenAI Enterprise team.

1. Serve as the primary engineering point of contact for enterprise prospects and customers

Acted as technical owner and support lead for strategic AI clients and key server GPU customers in deep learning applications.

  • Technical owner of strategic AI client accounts globally, collaborating on GPU AI software solutions.
  • Provided technical support for key customers of server GPUs, interfacing on DL applications and system configurations.
  • Coordinated issues and feature planning for frameworks like TensorFlow-ROCm with external partners.

2. 3+ years of management experience with a proven track record of building high-performing collaborative teams

Directed and managed AI software engineering teams focused on GPU AI frameworks, inference, and ecosystems at a leading semiconductor company.

  • Directed AI Software Engineering team owning GPU AI software including PyTorch, TensorFlow, JAX, ONNX-Runtime, and MIGraphX.
  • Led Sr. Director role overseeing ROCm GPU AI layer with LLM inference engines like vLLM, SGLang, MIGraphX, and TGI.
  • Managed deep learning frameworks team to enable and optimize TensorFlow, PyTorch, and ONNX-RT on AMD GPUs using ROCm.
  • Built teams contributing to AI ecosystems including OpenAI Triton, DeepSpeed, CuPy, and rocMLIR graph compiler.

3. Hands-on experience with production AI workloads including scaling and serving large models

Owned development and optimization of production GPU inference engines, frameworks, and operators for LLMs and GenAI at scale.

  • Owned LLM and Generative AI GPU inference engines including vLLM, SGLang, MIGraphX, TGI, and HuggingFace open-source models.
  • Developed high-performance AI operators for LLMs in Triton, CK, and ASM, supporting production AITER project kernels.
  • Led AI models team focusing on HuggingFace and open-source LLMs with public MAD automation tooling.
  • Optimized key DL models on data center GPUs and contributed to TensorFlow-ROCm for production releases.

4. Proficiency in Mojo, Python, C++, or CUDA programming languages, plus familiarity with frameworks such as MAX, Pytorch, or vLLM

Demonstrated deep expertise in C++, Python, PyTorch, vLLM, and CUDA-aligned GPU programming for AI workloads.

  • Expertise in C++ and Python for GPU AI software development across ROCm stack.
  • Led PyTorch, JAX, TensorFlow, and ONNX-Runtime support with ROCm/MIGraphX execution providers.
  • Integrated vLLM and other SOTA inference solutions for GenAI GPU serving.
  • Developed HIP-dependent workloads supporting TensorFlow, Caffe, and PaddlePaddle ecosystems.

Requirements & Candidate Alignment

Modular RequirementCandidate Qualification
7+ Years of experience in a customer-facing technical roles where you solved hard problems alongside clients16.6 years total experience including technical ownership of global strategic AI clients and support for key DL GPU customers
3+ Years of management experience with a proven track record of building high-performing collaborative teamsDirector and Sr. Director roles leading GPU AI software teams on frameworks, inference, and ecosystems
In-depth understanding of the GenAI execution lifecycle, deployment methodologies, along with testing and quality evaluation methodologiesOwned full GenAI stack including rocMLIR graph compiler, MIGraphX inference, vLLM serving, and PyTorch foundation contributions
Hands-on experience with production AI workloads including scaling and serving large modelsLed production LLM inference with vLLM, SGLang, TGI, HuggingFace models, and high-performance Triton operators
Proficiency in Mojo, Python, C++, or CUDA programming languages, plus familiarity with frameworks such as MAX, Pytorch, or vLLMProficient in C++, Python, PyTorch, vLLM, and CUDA workloads across ROCm GPU AI software
Proven ability to translate complex technical concepts into clear, concise, and understandable language that resonates with for both clients and internal teamsServed as technical representative for PyTorch and OpenXLA foundations, coordinating client issues and feature planning
Demonstrated ability to rapidly prototype and deliver Proofs of Concept (POCs) that win customer confidenceOptimized and enabled DL frameworks and models on GPUs, delivering production-ready ROCm features and client demos

If you would like to discuss this candidate or other critical roles, here is a link to my calendar to schedule a call:

Schedule a Call

Contact:

Jason Rath

TalentPros.AI

Finding the signal in the noise since 2005

512-993-8228

Jason@TalentPros.AI

Top