Candidate Profile — Candidate #4
Position: Modular — Senior AI Kernel Engineer (Remote)
Candidate Location: New York Metro
Experience: 23.6 years experience

Relevant Experience & Education Highlights

Extensive 23.6 years of compiler and low-level systems expertise, including GPU kernel programming with CUDA, CUTLASS, NVSHMEM, and OpenAI Triton Compiler for AI inference, aligns strongly with designing and optimizing high-performance AI kernels at Modular.

1. Design, implement, and optimize performance-critical kernels for AI inference workloads (e.g., GEMM, attention, communication, fusion)

Led development of AI compilers and inference pipelines optimized for GPUs and distributed systems.

  • Developed AI compiler for PyTorch using OpenAI Triton Compiler and MLIR at an AI research firm, enabling distributed inference and training.
  • Implemented distributed AI inference/training pipelines using OpenMPI, CUDA, CUTLASS, NVSHMEM, and PyTorch at an AI research firm.
  • Optimized kernels for NVIDIA Triton Inference Server and LLaMA-CPP, alongside Apple METAL backend for Triton compiler at an AI research firm.
  • Engineered NVIDIA CUDA kernels and PyTorch C++ optimizations at a quantum computing startup.

2. Extensive hands-on experience with GPU kernel programming (CUDA, HIP, or equivalent)

Delivered low-level GPU programming across CUDA and related frameworks for high-performance computing.

  • Developed NVIDIA CUDA solutions for AI and quantum workloads at a quantum computing startup.
  • Built assemblers and codegen for IBM Quantum ISAs using LLVM backends at a leading quantum research organization.
  • Optimized ARM64 codegen and atomics in LLVM/Clang for ThunderX architecture at a semiconductor company.
  • Contributed to GPU-accelerated supercomputing projects with Sandia, LANL, and Cray at a semiconductor company.

3. Work closely with compiler and runtime teams to influence code generation, scheduling, and kernel fusion strategies

Drove compiler toolchain advancements and collaborations across hardware and software stacks.

  • Led Clang/LLVM backends for Solaris Intel/SPARC, including JIT and OpenMP ports at a large enterprise software company.
  • Maintained GCC4/GCC5 for Intel/AMD/SPARC and developed kernel drivers at a pioneering systems company.
  • Developed quantum compilers using LLVM/MLIR, OpenQASM 3.0 front-end, and RISC-V on FPGA at quantum startups and a leading research organization.
  • Retargeted TVM and specialized compilers for TensorFlow models with modulo scheduling in LLVM at a specialized compiler firm.

Requirements & Candidate Alignment

Modular RequirementCandidate Qualification
5+ Years of experience in performance-critical systems or kernel development (or equivalent depth of expertise)23.6 years across compiler engineering, kernel development, GPU programming, and low-level systems at AI firms, quantum organizations, semiconductor companies, and enterprise software leaders
Strong proficiency in C/C++ and low-level programmingExpertise in C++ (C++98-17), C, Assembler (SPARC, x86_64, ARM64, PPC64LE), including contributions to Apache C++ Standard Library PMC and BOOST maintainer
Extensive hands-on experience with GPU kernel programming (CUDA, HIP, or equivalent)Hands-on with CUDA, CUTLASS, NVSHMEM, Apple METAL for AI inference/training and Triton Inference Server at AI research and quantum firms
Deep understanding of GPU architecture, including memory hierarchies, synchronization, and execution modelsDemonstrated through ARM64 ThunderX optimizations, supercomputing codegen, quantum assemblers, and distributed MPI/CUDA pipelines
Proven track record of delivering measurable performance improvements in production systemsDelivered optimizations in AI compilers for PyTorch/Triton, LLVM retargeting for AArch64/flang, and high-performance kernel drivers
Experience with PTX, assembly-level tuning, or code generation frameworks (e.g., Triton)Extensive experience with OpenAI Triton Compiler backends, PTX via CUDA, assemblers for quantum silicon, and LLVM/MLIR codegen
Experience optimizing distributed or multi-GPU inference pipelinesOptimized pipelines using OpenMPI, NVSHMEM, CUDA for distributed AI inference/training at an AI research firm including NVIDIA Triton and Triton Inference Server.
Contributions to open-source kernel libraries, compilers, or performance toolsContributed to GCC/Clang/LLVM, GNU Binutils, GDB/lldb, Apache C++ Std Library, OpenSolaris Project, and KDE Solaris

If you would like to discuss this candidate or other critical roles, here is a link to my calendar to schedule a call:

Schedule a Call

Contact:

Jason Rath

TalentPros.AI

Finding the signal in the noise since 2005

512-993-8228

Jason@TalentPros.AI

Top