Candidate Profile — Candidate #6
Position: Modular — Senior AI Kernel Engineer (Remote)
Candidate Location: San Francisco Bay Area, CA
Experience: 9.3 years experience
Relevant Experience & Education Highlights
Expertise in GPU kernel development using PTX, quantization for large language models, and proven performance optimizations positions this candidate strongly for leading AI kernel engineering at Modular.
1. Design, implement, and optimize performance-critical kernels for AI inference workloads (e.g., GEMM, attention, communication, fusion)
Developed and tuned low-level kernels for deep learning inference and transformer models.
- Developed GPU kernels in PTX alongside quantization for serving large language models at a cutting-edge AI company.
- Optimized infrastructures including GPU kernels for transformer-based natural language models at a major technology company.
- Built end-to-end deep learning pipelines for image classification, object detection, and segmentation at a prominent enterprise software company.
- Implemented forward and back propagation with performance improvements in Apache SystemML at an enterprise software company.
2. Lead kernel-level optimization efforts across single-GPU, multi-GPU, and heterogeneous hardware environments
Delivered measurable speedups and efficiency gains in distributed GPU-accelerated systems.
- Improved distributed neural network training performance by 5 times using PySpark in Apache SystemML.
- Designed hierarchical indexing with chunking for Apache Spark, accelerating multidimensional data queries 2-4 times.
- Enabled Spark with HDFS to support HDF4, HDF5, and NetCDF4 formats, saving 10 times storage space.
- Integrated U-Net models for satellite image segmentation with parallel prediction in ClimateSpark pipelines.
3. Deep understanding of GPU architecture, including memory hierarchies, synchronization, and execution models
Applied low-level expertise through open-source contributions and hardware-aware optimizations.
- Contributed over 59K lines as TensorFlow committer, fixing issues and adding deep learning features.
- Performed end-to-end performance optimization for large language models and quantization techniques.
- Utilized TensorFlow, PyTorch, and related frameworks for systems-level AI model tuning.
- Worked on PTX and assembly-level tuning for GPU kernels in production inference systems.
Requirements & Candidate Alignment
| Modular Requirement | Candidate Qualification |
|---|---|
| Experience: 5+ years of experience in performance-critical systems or kernel development (or equivalent depth of expertise) | 9.3 years across GPU kernel development, AI optimization, and distributed systems roles |
| Programming: Strong proficiency in C/C++ and low-level programming | Low-level programming via PTX kernels, SystemML implementations, and TensorFlow contributions |
| GPU Kernels: Extensive hands-on experience with GPU kernel programming (CUDA, HIP, or equivalent) | GPU kernel programming including PTX for large language model inference and quantization |
| GPU Architecture: Deep understanding of GPU architecture, including memory hierarchies, synchronization, and execution models | GPU architecture expertise from end-to-end optimizations and performance profiling in production |
| Performance Track Record: Proven track record of delivering measurable performance improvements in production systems | Delivered gains like 5x training speedup, 2-4x query acceleration, and 10x storage efficiency |
| PTX/Assembly: Experience with PTX, assembly-level tuning, or code generation frameworks (e.g., Triton) | PTX and Triton experience in kernel development for AI inference workloads |
| AI Models: Understanding of modern AI models (e.g., transformers, LLMs, diffusion) from a systems and performance perspective | Transformers and LLMs expertise through serving optimizations and quantization |
| Open-Source: Contributions to open-source kernel libraries, compilers, or performance tools | TensorFlow committer with extensive code contributions to deep learning frameworks |
If you would like to discuss this candidate or other critical roles, here is a link to my calendar to schedule a call:
Schedule a CallContact:
Finding the signal in the noise since 2005