Candidate Profile — Candidate #3
Position: Modular — Senior AI Kernel Engineer (Remote)
Candidate Location: San Francisco Bay Area
Experience: 30.2 years experience

Relevant Experience & Education Highlights

Brings decades of hands-on expertise in designing and optimizing high-performance CUDA kernels for AI inference, delivering dramatic speedups and innovative performance models, perfectly aligning with Modular's mission to revolutionize AI infrastructure.

1. Design, implement, and optimize performance-critical kernels for AI inference workloads (e.g., GEMM, attention, communication, fusion)

Excelled in crafting fused and memory-efficient kernels that outperform industry standards for neural network operations.

  • Developed memory-efficient CUDA kernels for MBConv+SE blocks and ConvFirst blocks, achieving up to 14x and 5x speedups over PyTorch Inductor as an independent researcher.
  • Implemented block-fusion kernels in CUDA and PTX for an efficient inference engine on NVIDIA Xavier SoC at a leading autonomous driving technology company, delivering ~4x speedup over NVIDIA TensorRT on object detection and semantic segmentation tasks.
  • Created the Spio kernel library for PyTorch featuring named tensors, runtime compilation, kernel performance models, and torch.compile integration, with initial kernels outperforming built-in PyTorch grouped convolution kernels.
  • Published original research introducing ConvFirst building block for convolutional neural networks and custom kernels enabling ~4x speedup over ConvNeXt with equal accuracy.

2. Lead kernel-level optimization efforts across single-GPU, multi-GPU, and heterogeneous hardware environments

Drove kernel optimizations yielding substantial performance gains in production AI systems on GPU hardware.

  • Improved efficiency of neural network software on Intel Gen graphics processors using OpenCL, achieving approximately 4x speedup over existing codebase at a major semiconductor company.
  • Developed high-performance NVIDIA GPU kernels for an autopilot system at a leading electric vehicle company.
  • Led development of graph rewriting compiler for neural networks and enhanced core system-level code for production ADAS system at a leading autonomous driving technology company.

3. Analyze performance using profilers, hardware counters, and microbenchmarks; translate insights into concrete improvements

Invented advanced performance modeling techniques and applied them to achieve breakthrough efficiencies in AI workloads.

  • Introduced Waterline performance model for sequences of parallel kernels, correcting errors in widely used Roofline analysis, as part of five years of original research published in 'On the Efficiency of Convolutional Neural Networks.'
  • Developed theory of neural network efficiency unifying model efficiency, computational efficiency, and latency, demonstrating co-optimization of model and program for superior performance.
  • Performed research on high-performance pre-processing pipelines and devised multi-scale convnet models at a leading autonomous driving technology company.

4. Work closely with compiler and runtime teams to influence code generation, scheduling, and kernel fusion strategies

Collaborated on compiler innovations and kernel frameworks to enhance AI system performance.

  • Led development of graph rewriting compiler for neural networks at a leading autonomous driving technology company.
  • Integrated Spio kernel library with torch.compile for runtime compilation and performance modeling as an independent researcher.
  • Developed client-server framework for testing and benchmarking inference engines.

Requirements & Candidate Alignment

Modular RequirementCandidate Qualification
5+ Years of experience in performance-critical systems or kernel development (or equivalent depth of expertise)30.2 years total experience, including deep kernel optimization for AI inference across multiple roles
Strong proficiency in C/C++ and low-level programmingStrong proficiency in C/C++ demonstrated through extensive GPU kernel development
Extensive hands-on experience with GPU kernel programming (CUDA, HIP, or equivalent)Extensive hands-on experience with CUDA kernel programming, including fused kernels, TensorRT, and PTX-level tuning
Deep understanding of GPU architecture, including memory hierarchies, synchronization, and execution modelsDeep understanding evidenced by Waterline performance model and memory-efficient kernel designs including CUDA and TensorRT.
Proven track record of delivering measurable performance improvements in production systemsProven track record with ~4x speedups over NVIDIA TensorRT and up to 14x over PyTorch Inductor in production AI inference
Experience with PTX, assembly-level tuning, or code generation frameworks (e.g., Triton)Experience with PTX in block-fusion kernels and graph rewriting compilers for neural networks
Contributions to open-source kernel libraries, compilers, or performance toolsContributions include developing Spio CUDA kernel library for PyTorch with torch.compile integration
Experience optimizing distributed or multi-GPU inference pipelinesExperience optimizing inference on NVIDIA Xavier SoC and multi-scale convnet models including CUDA and TensorRT.

If you would like to discuss this candidate or other critical roles, here is a link to my calendar to schedule a call:

Schedule a Call

Contact:

Jason Rath

TalentPros.AI

Finding the signal in the noise since 2005

512-993-8228

Jason@TalentPros.AI

Top