Candidate Profile — Candidate #2
Position: Modular — Senior AI Kernel Engineer (Remote)
Candidate Location: San Francisco Bay Area
Experience: 9.8 years experience

Relevant Experience & Education Highlights

Brings extensive hands-on experience optimizing AI kernels and GPU libraries for inference workloads, aligning strongly with designing high-performance kernels for Modular's AI infrastructure.

1. Design, implement, and optimize performance-critical kernels for AI inference workloads (e.g., GEMM, attention, communication, fusion)

Demonstrated expertise in developing and tuning kernels for machine learning operators directly matching GEMM and transformer-based models.

  • Developed GPU libraries for linear algebra (GEMM) used in TensorFlow, PyTorch, BERT, and Transformer models at a major GPU manufacturer.
  • Optimized algorithms for machine learning operators and frameworks on a custom AI stack at a leading semiconductor company.
  • Tuned ImageNet models to improve inference and initialization time across computational graphs on mobile processors.

2. Lead kernel-level optimization efforts across single-GPU, multi-GPU, and heterogeneous hardware environments

Proven track record in multi-GPU systems and heterogeneous accelerators, including performance tuning and CI infrastructure.

  • Coded in HIP, C++, and Python for GPU optimization using data structures at a major GPU manufacturer.
  • Created Dockerized multi-GPU CI system using Jenkins for testing GPU libraries.
  • Implemented support for concurrent graph execution using tightly coupled memory on tensor processors.
  • Developed tools for profiling and debugging neural network deployment on Snapdragon processors.

3. Work closely with compiler and runtime teams to influence code generation, scheduling, and kernel fusion strategies

Hands-on compiler development and build optimization experience supports collaboration on code generation and fusion.

  • Coded in C++ and Python for front-end compiler team, working on lexical analysis and build engineering at a leading EDA company.
  • Refactored CMake build code using design patterns for a neural processing engine project.
  • Produced merge request tool building and testing static analysis in Docker across platforms.

4. Deep understanding of GPU architecture, including memory hierarchies, synchronization, and execution models

Background in low-level programming on GPUs and custom accelerators provides foundation for architectural optimizations.

  • Optimized for Hexagon Tensor Processor, focusing on mobile drivers, SDK, and AI software stack at a major semiconductor company.
  • Produced server-side C++ code using object-oriented programming, data structures, sockets, and threads for imaging software.
  • Coded in C++, C, and Python emphasizing low-level optimization for vector operations and memory management.

Requirements & Candidate Alignment

Modular RequirementCandidate Qualification
5+ Years of experience in performance-critical systems or kernel development (or equivalent depth of expertise)9.8 years of software engineering experience focused on AI optimization, GPU libraries, and kernel-level performance
Strong proficiency in C/C++ and low-level programmingExtensive C/C++ expertise across GPU libraries, compiler front-ends, tensor processors, and server-side imaging software
Extensive hands-on experience with GPU kernel programming (CUDA, HIP, or equivalent)HIP and equivalent GPU programming for linear algebra libraries supporting TensorFlow, PyTorch, and Transformers
Deep understanding of GPU architecture, including memory hierarchies, synchronization, and execution modelsOptimized memory and synchronization in concurrent graph execution and tightly coupled memory on tensor processors
Proven track record of delivering measurable performance improvements in production systemsDelivered inference speedups by tuning ImageNet models and optimizing ML operators on production AI stacks
Experience with PTX, assembly-level tuning, or code generation frameworks (e.g., Triton)Compiler and code generation experience including front-end lexical analysis and build optimization tools
Experience optimizing distributed or multi-GPU inference pipelinesBuilt multi-GPU CI systems with Docker and Jenkins for distributed testing of GPU libraries
Understanding of modern AI models (e.g., transformers, LLMs, diffusion) from a systems and performance perspectiveOptimized for Transformers and BERT in GPU libraries integrated with PyTorch and TensorFlow inference

If you would like to discuss this candidate or other critical roles, here is a link to my calendar to schedule a call:

Schedule a Call

Contact:

Jason Rath

TalentPros.AI

Finding the signal in the noise since 2005

512-993-8228

Jason@TalentPros.AI

Top