Candidate Profile — Candidate #5
Position: Modular — Senior AI Kernel Engineer (Remote)
Candidate Location: Los Altos, California
Experience: 21.9 years experience

Relevant Experience & Education Highlights

Exceptional expertise in GPU kernel optimization, Triton compiler leadership, CUDA performance tuning for AI inference, and hardware heterogeneity makes this candidate a strong fit for leading high-performance kernel development at Modular.

1. Design, implement, and optimize performance-critical kernels for AI inference workloads (e.g., GEMM, attention, communication, fusion)

Optimized kernels and runtimes for data-center AI training and inference at scale.

  • Optimized CUDA kernels for widely deployed training workloads, delivering multiple megawatts of efficiency improvements.
  • Improved GPU utilization for LLM inference through autoscaling, dynamic throughput tracking, and engineering support for analytics.
  • Enhanced PyTorch ecosystem with build caching and debugging for torch.compile flow resilience.
  • Developed safe, efficient programmability for data-center AI training and inference at a leading autonomous driving and AI company.

2. Lead kernel-level optimization efforts across single-GPU, multi-GPU, and heterogeneous hardware environments

Led teams to address hardware heterogeneity and continuous performance measurement across NVIDIA and AMD platforms.

  • Built team in programming languages organization to deploy GPU runtimes frequently and support Triton compiler improvements for AMD and NVIDIA targeting.
  • Created test infrastructure for continuous validation of PyTorch models across backends with targeted improvements.
  • Drove $100s of millions in efficiency savings through improvements in native server compilers, high-performance libraries, and infrastructure utilization.

3. Deep understanding of GPU architecture, including memory hierarchies, synchronization, and execution models

Contributed to GPU compute architecture modeling, standards, and programming models.

  • Proposed, modeled, and evaluated GPU compute hardware features with focus on memory consistency in cache hierarchy.
  • Represented in Khronos OpenCL and Heterogeneous System Architecture committees specializing in memory consistency.
  • Maintained official OpenCL C++ header (cl.hpp) and contributed to high-level programming models and runtime interfaces for heterogeneous compute platforms.
  • Addressed C++ concurrency, parallelism, and fundamental primitives in high-scale systems handling billions of queries per second.

4. Work closely with compiler and runtime teams to influence code generation, scheduling, and kernel fusion strategies

Influenced compiler ownership, hardware support, and developer experience in production AI systems.

  • Fostered internal ownership of Triton compiler and planned enhancements for multi-vendor hardware support.
  • Led engineering efforts in native server performance, delivering broad codebase improvements for developers.
  • Contributed to C++ standards committee and INCITS steering for programming languages and systems software.
  • Collaborated on runtime systems and programming models for future fusion architectures at a major semiconductor company.

Requirements & Candidate Alignment

Modular RequirementCandidate Qualification
5+ Years of experience in performance-critical systems or kernel development (or equivalent depth of expertise)21.9 years across GPU programming, AI efficiency leadership, and compiler/runtime optimization roles
Strong proficiency in C/C++ and low-level programmingExpert C/C++ proficiency, including concurrency primitives, folly libraries, OpenCL C++ APIs, and standards committee representation
Extensive hands-on experience with GPU kernel programming (CUDA, HIP, or equivalent)Hands-on CUDA kernel programming, optimizing for LLM inference, training workloads, and data-center AI
Deep understanding of GPU architecture, including memory hierarchies, synchronization, and execution modelsDeep GPU architecture knowledge, from modeling memory consistency, cache hierarchies, CUDA, and heterogeneous compute platforms
Proven track record of delivering measurable performance improvements in production systemsDelivered $100s millions in efficiency savings and multiple MWs of GPU improvements in production infrastructure
Experience with PTX, assembly-level tuning, or code generation frameworks (e.g., Triton)Led Triton compiler efforts, including ownership, AMD/NVIDIA targeting, and PyTorch integration
Experience optimizing distributed or multi-GPU inference pipelinesOptimized multi-GPU inference, focusing on LLM scale-up, hardware utilization, CUDA, and heterogeneous backends
Understanding of modern AI models (e.g., transformers, LLMs, diffusion) from a systems and performance perspectiveDeep systems expertise with LLMs, driving inference throughput, autoscaling, and kernel optimizations

If you would like to discuss this candidate or other critical roles, here is a link to my calendar to schedule a call:

Schedule a Call

Contact:

Jason Rath

TalentPros.AI

Finding the signal in the noise since 2005

512-993-8228

Jason@TalentPros.AI

Top