Relevant Experience & Education Highlights
Extensive expertise in compiler infrastructure, hardware-software co-design, and optimization for AI accelerators positions this candidate as an exceptional fit for developing and scaling software kernels at d-Matrix.
1. Development, enhancement, and maintenance of software kernels for next-generation AI hardware
Proven track record in implementing and optimizing kernels across diverse AI hardware platforms.
- Implemented LLVM optimizations and code generation on wafer-scale tensor processing unit at an AI accelerator company.
- Drove optimizations within Halide language to leverage GPU texture hardware and DSP scatter/gather at a specialized imaging processing firm.
- Conducted performance engineering on ARM NEON for touch sensor processing, improving performance by 200% through C representations and machine code at a mobile computing company.
- Optimized OpenCL performance for Snapdragon GPUs and provided LLVM compiler feedback on hardware specs at a major semiconductor provider.
2. Experience building software kernels for HW architectures and mapping algorithms to architecture
Deep understanding of mapping computational graphs and algorithms to specialized hardware like DSPs, GPUs, and AI accelerators.
- Built MLIR pass generating Risc-V Vector intrinsics from MLIR Affine dialect, improving vector recognition and lowering to LLVM for pre-tapeout AI chip at a cutting-edge AI computing startup.
- Coordinated kernel scheduler for data cluster serving petabytes to over 64K processors, establishing metrics for GEMM components and data transfer at an innovative AI hardware company.
- Implemented optimizations for ARM CPUs, Snapdragon Hexagon DSP, Adreno GPUs, and CEVA XM4 DSPs at a digital signal processing firm.
- Designed Python algorithms identifying patterns in data-flow graphs from TensorFlow for ML algorithms on Cerebras CS-1 at a wafer-scale engine company.
3. Compiler infrastructure and experience with ML compilers such as MLIR, LLVM
Hands-on work with MLIR, LLVM, and related tools aligns directly with building compiler infrastructure for AI workloads.
- Implemented control flow and math optimizations for P4/P4+/P416 language using LLVM, coordinating ISA specification with hardware designers at a high-performance networking company.
- Performed performance analysis of BERT models and generated CI/CD perf tests for optimization changelogs at an AI research lab.
- Implemented instructions and optimizations for Float, BFloat16, FP8 data types using state-of-the-art LLVM and MLIR at an AI computing startup.
- Broke down MNIST, BERT-Large, DLRM, GPT-2 tasks to schedule massively parallel asynchronous tasks on AI hardware.
4. Experience implementing operators commonly used in ML workloads—GEMMs, Convolutions, BLAS, SIMD operators
Strong foundation in ML operators and SIMD processing demonstrated through practical implementations.
- Established metrics for GEMM components, Risc-V hosts, and data transfer subsystems tailored for AI card performance.
- Collaborated on ISA simplification and optimization for scalar and vector instruction lanes supporting ML workloads.
- Optimized code base for BLAS-like operations and SIMD on DSPs and GPUs including Adreno and Hexagon.
- Added ARM NEON machine code to critical sections for SIMD-accelerated sensor processing.
Requirements & Candidate Alignment
| d-Matrix Requirement | Candidate Qualification |
|---|---|
| Minimum: MS in computer engineering, math, physics, or a related degree with 5+ years of industry experience or a PhD... with 1+ years of industry experience | 34.2 years of industry experience in compiler engineering and AI hardware optimization, complemented by B.S. in Computer Science from a top-tier public university |
| Strong grasp of computer architecture, data structures, system software, and machine learning fundamentals | Extensive expertise in computer architecture and ML fundamentals, evidenced by Risc-V vectorization, data-flow graph analysis, and BERT performance tuning |
| Proficient in C/C++ and Python development in Linux environments and using standard development tools | Proficient in C/C++, Python, LLVM, and MLIR, applied in Linux-based optimizations for AI chips and DSPs |
| Experience implementing algorithms in high-level languages such as C/C++ and Python | Implemented algorithms in C/C++ and Python for data encapsulation patterns in TensorFlow-generated ML graphs |
| Experience implementing algorithms for specialized hardware such as FPGAs, DSPs, GPUs, and AI accelerators using libraries such as CUDA, etc. | Implemented for GPUs (Adreno, Snapdragon), DSPs (Hexagon, CEVA XM4), and AI accelerators (wafer-scale TPU, Risc-V AI chips) using LLVM, OpenCL, and Halide |
| Experience in implementing operators commonly used in ML workloads—GEMMs, Convolutions, BLAS, SIMD operators for operations like softmax, layer normalization, pooling, etc. | Implemented GEMM metrics, SIMD vector intrinsics, and optimizations for BFloat16/FP8 in ML tasks like BERT and MNIST |
| Experience with development for embedded SIMD vector processors such as Tensilica | Developed for embedded SIMD including ARM NEON, Risc-V Vector, Hexagon DSP, aligning with vector processor workflows |
| Self-motivated team player with a strong sense of ownership and leadership | Led cross-team coordination on ISA design, kernel scheduling, and hardware-software optimizations across multiple AI startups |
If you would like to discuss this candidate or other critical roles, here is a link to my calendar to schedule a call:
Schedule a CallContact:
Finding the signal in the noise since 2005