Candidate Profile — Candidate #4
Position: Modular — Senior AI Kernel Engineer (Remote)
Candidate Location: New York Metro
Experience: 23.6 years experience
Relevant Experience & Education Highlights
Extensive 23.6 years of compiler and low-level systems expertise, including GPU kernel programming with CUDA, CUTLASS, NVSHMEM, and OpenAI Triton Compiler for AI inference, aligns strongly with designing and optimizing high-performance AI kernels at Modular.
1. Design, implement, and optimize performance-critical kernels for AI inference workloads (e.g., GEMM, attention, communication, fusion)
Led development of AI compilers and inference pipelines optimized for GPUs and distributed systems.
- Developed AI compiler for PyTorch using OpenAI Triton Compiler and MLIR at an AI research firm, enabling distributed inference and training.
- Implemented distributed AI inference/training pipelines using OpenMPI, CUDA, CUTLASS, NVSHMEM, and PyTorch at an AI research firm.
- Optimized kernels for NVIDIA Triton Inference Server and LLaMA-CPP, alongside Apple METAL backend for Triton compiler at an AI research firm.
- Engineered NVIDIA CUDA kernels and PyTorch C++ optimizations at a quantum computing startup.
2. Extensive hands-on experience with GPU kernel programming (CUDA, HIP, or equivalent)
Delivered low-level GPU programming across CUDA and related frameworks for high-performance computing.
- Developed NVIDIA CUDA solutions for AI and quantum workloads at a quantum computing startup.
- Built assemblers and codegen for IBM Quantum ISAs using LLVM backends at a leading quantum research organization.
- Optimized ARM64 codegen and atomics in LLVM/Clang for ThunderX architecture at a semiconductor company.
- Contributed to GPU-accelerated supercomputing projects with Sandia, LANL, and Cray at a semiconductor company.
3. Work closely with compiler and runtime teams to influence code generation, scheduling, and kernel fusion strategies
Drove compiler toolchain advancements and collaborations across hardware and software stacks.
- Led Clang/LLVM backends for Solaris Intel/SPARC, including JIT and OpenMP ports at a large enterprise software company.
- Maintained GCC4/GCC5 for Intel/AMD/SPARC and developed kernel drivers at a pioneering systems company.
- Developed quantum compilers using LLVM/MLIR, OpenQASM 3.0 front-end, and RISC-V on FPGA at quantum startups and a leading research organization.
- Retargeted TVM and specialized compilers for TensorFlow models with modulo scheduling in LLVM at a specialized compiler firm.
Requirements & Candidate Alignment
| Modular Requirement | Candidate Qualification |
|---|---|
| 5+ Years of experience in performance-critical systems or kernel development (or equivalent depth of expertise) | 23.6 years across compiler engineering, kernel development, GPU programming, and low-level systems at AI firms, quantum organizations, semiconductor companies, and enterprise software leaders |
| Strong proficiency in C/C++ and low-level programming | Expertise in C++ (C++98-17), C, Assembler (SPARC, x86_64, ARM64, PPC64LE), including contributions to Apache C++ Standard Library PMC and BOOST maintainer |
| Extensive hands-on experience with GPU kernel programming (CUDA, HIP, or equivalent) | Hands-on with CUDA, CUTLASS, NVSHMEM, Apple METAL for AI inference/training and Triton Inference Server at AI research and quantum firms |
| Deep understanding of GPU architecture, including memory hierarchies, synchronization, and execution models | Demonstrated through ARM64 ThunderX optimizations, supercomputing codegen, quantum assemblers, and distributed MPI/CUDA pipelines |
| Proven track record of delivering measurable performance improvements in production systems | Delivered optimizations in AI compilers for PyTorch/Triton, LLVM retargeting for AArch64/flang, and high-performance kernel drivers |
| Experience with PTX, assembly-level tuning, or code generation frameworks (e.g., Triton) | Extensive experience with OpenAI Triton Compiler backends, PTX via CUDA, assemblers for quantum silicon, and LLVM/MLIR codegen |
| Experience optimizing distributed or multi-GPU inference pipelines | Optimized pipelines using OpenMPI, NVSHMEM, CUDA for distributed AI inference/training at an AI research firm including NVIDIA Triton and Triton Inference Server. |
| Contributions to open-source kernel libraries, compilers, or performance tools | Contributed to GCC/Clang/LLVM, GNU Binutils, GDB/lldb, Apache C++ Std Library, OpenSolaris Project, and KDE Solaris |
If you would like to discuss this candidate or other critical roles, here is a link to my calendar to schedule a call:
Schedule a CallContact:
Finding the signal in the noise since 2005