Candidate Profile — Candidate #1
Position: Modular — Senior AI Kernel Engineer (Remote)
Candidate Location: Toronto, Ontario
Experience: 4.8 years experience
Relevant Experience & Education Highlights
Demonstrated deep expertise in compiler optimizations, GPU kernel programming via Triton, and AI inference workloads positions this candidate strongly for leading kernel design and performance tuning at Modular.
1. Design, implement, and optimize performance-critical kernels for AI inference workloads
Applied advanced kernel and runtime optimizations to accelerate AI/ML inference and training across diverse models.
- Enhanced rematerialization optimizer by improving memory heuristics accuracy and postprocessing coverage at a leading AI systems company.
- Adapted compiler for KV-caching, kernel, and runtime optimizations achieving GPU-impossible inference speeds for LLMs including GPT family, Llama, Mixtral MoE, and BERT.
- Designed dynamic recompile loop for stable training compilation across multiple LLM architectures.
2. Extensive hands-on experience with GPU kernel programming (CUDA, HIP, or equivalent)
Integrated and modified GPU kernel languages like Triton for custom accelerators and tensor parallelism.
- Designed integration contract between Triton language and compiler with an AI accelerator startup, defining translation semantics from GPU kernel parallelism to tensor parallelism.
- Modified triton-to-linalg pass and built cleanup pass to convert Triton into memref representation for custom dialect operations.
- Resolved rank broadcasting and shape mismatch issues in tensor operations within custom compiler dialects.
3. Proven track record of delivering measurable performance improvements in production systems
Improved compile times, execution performance, and optimization techniques in heterogeneous compiler environments.
- Improved compile and execution time of optimizing compiler for HPC on specialized platforms at a major technology company's heterogeneous compiler lab.
- Implemented new compiler analysis and transformation passes within LLVM mid-end architecture.
- Researched innovative optimization opportunities for AI/ML workloads in heterogeneous environments.
Requirements & Candidate Alignment
| Modular Requirement | Candidate Qualification |
|---|---|
| 5+ Years of experience in performance-critical systems or kernel development (or equivalent depth of expertise) | 4.8 years across machine learning stack engineering roles focused on compilers, kernels, and AI optimizations |
| Strong proficiency in C/C++ and low-level programming | Strong C++ proficiency demonstrated in compiler passes, low-level optimizations, and kernel integrations |
| Extensive hands-on experience with GPU kernel programming (CUDA, HIP, or equivalent) | Extensive Triton experience for GPU kernel programming, including parallelism translation to tensor ops for accelerators |
| Deep understanding of GPU architecture, including memory hierarchies, synchronization, and execution models | Deep GPU knowledge via memory heuristics, KV-caching adaptations, and Triton-to-custom dialect conversions |
| Proven track record of delivering measurable performance improvements in production systems | Delivered improvements in rematerialization, compile/execution times, and inference speeds for LLMs |
| Experience with PTX, assembly-level tuning, or code generation frameworks (e.g., Triton) | Triton expertise including pass modifications, type converters, and cleanup for linalg and custom dialects |
| Experience optimizing distributed or multi-GPU inference pipelines | Optimized multi-GPU inference through kernel adaptations and tensor parallelism semantics |
| Understanding of modern AI models (e.g., transformers, LLMs, diffusion) from a systems and performance perspective | Strong LLM systems knowledge with optimizations for GPT, Llama, Mixtral MoE, BERT, and multimodal models |
If you would like to discuss this candidate or other critical roles, here is a link to my calendar to schedule a call:
Schedule a CallContact:
Finding the signal in the noise since 2005