Relevant Experience & Education Highlights
Brings extensive hands-on experience optimizing AI kernels and GPU libraries for inference workloads, aligning strongly with designing high-performance kernels for Modular's AI infrastructure.
1. Design, implement, and optimize performance-critical kernels for AI inference workloads (e.g., GEMM, attention, communication, fusion)
Demonstrated expertise in developing and tuning kernels for machine learning operators directly matching GEMM and transformer-based models.
- Developed GPU libraries for linear algebra (GEMM) used in TensorFlow, PyTorch, BERT, and Transformer models at a major GPU manufacturer.
- Optimized algorithms for machine learning operators and frameworks on a custom AI stack at a leading semiconductor company.
- Tuned ImageNet models to improve inference and initialization time across computational graphs on mobile processors.
2. Lead kernel-level optimization efforts across single-GPU, multi-GPU, and heterogeneous hardware environments
Proven track record in multi-GPU systems and heterogeneous accelerators, including performance tuning and CI infrastructure.
- Coded in HIP, C++, and Python for GPU optimization using data structures at a major GPU manufacturer.
- Created Dockerized multi-GPU CI system using Jenkins for testing GPU libraries.
- Implemented support for concurrent graph execution using tightly coupled memory on tensor processors.
- Developed tools for profiling and debugging neural network deployment on Snapdragon processors.
3. Work closely with compiler and runtime teams to influence code generation, scheduling, and kernel fusion strategies
Hands-on compiler development and build optimization experience supports collaboration on code generation and fusion.
- Coded in C++ and Python for front-end compiler team, working on lexical analysis and build engineering at a leading EDA company.
- Refactored CMake build code using design patterns for a neural processing engine project.
- Produced merge request tool building and testing static analysis in Docker across platforms.
4. Deep understanding of GPU architecture, including memory hierarchies, synchronization, and execution models
Background in low-level programming on GPUs and custom accelerators provides foundation for architectural optimizations.
- Optimized for Hexagon Tensor Processor, focusing on mobile drivers, SDK, and AI software stack at a major semiconductor company.
- Produced server-side C++ code using object-oriented programming, data structures, sockets, and threads for imaging software.
- Coded in C++, C, and Python emphasizing low-level optimization for vector operations and memory management.
Requirements & Candidate Alignment
| Modular Requirement | Candidate Qualification |
|---|---|
| 5+ Years of experience in performance-critical systems or kernel development (or equivalent depth of expertise) | 9.8 years of software engineering experience focused on AI optimization, GPU libraries, and kernel-level performance |
| Strong proficiency in C/C++ and low-level programming | Extensive C/C++ expertise across GPU libraries, compiler front-ends, tensor processors, and server-side imaging software |
| Extensive hands-on experience with GPU kernel programming (CUDA, HIP, or equivalent) | HIP and equivalent GPU programming for linear algebra libraries supporting TensorFlow, PyTorch, and Transformers |
| Deep understanding of GPU architecture, including memory hierarchies, synchronization, and execution models | Optimized memory and synchronization in concurrent graph execution and tightly coupled memory on tensor processors |
| Proven track record of delivering measurable performance improvements in production systems | Delivered inference speedups by tuning ImageNet models and optimizing ML operators on production AI stacks |
| Experience with PTX, assembly-level tuning, or code generation frameworks (e.g., Triton) | Compiler and code generation experience including front-end lexical analysis and build optimization tools |
| Experience optimizing distributed or multi-GPU inference pipelines | Built multi-GPU CI systems with Docker and Jenkins for distributed testing of GPU libraries |
| Understanding of modern AI models (e.g., transformers, LLMs, diffusion) from a systems and performance perspective | Optimized for Transformers and BERT in GPU libraries integrated with PyTorch and TensorFlow inference |
If you would like to discuss this candidate or other critical roles, here is a link to my calendar to schedule a call:
Schedule a CallContact:
Finding the signal in the noise since 2005