Candidate Profile — Candidate #4
Position: d-Matrix — Software Engineer, Staff - Kernels (Santa Clara, CA)
Candidate Location: San Jose, CA
Experience: 7.3 years experience
Relevant Experience & Education Highlights
Demonstrates deep expertise in developing compilers and kernels for AI hardware architectures, including GEMM, convolutions, and attention mechanisms, with strong alignment to d-Matrix's requirements for software kernels on next-generation AI compute engines.
1. Development, enhancement, and maintenance of software kernels for next-generation AI hardware
Built and optimized kernels for AI accelerators, GPUs, and specialized processors.
- Performed HW/SW co-design for AI architecture at a major semiconductor company, creating data-movement compiler maximizing memory-hierarchy efficiency for high-throughput GEMM, convolution, and multi-head attention workloads.
- Developed flexible attention compiler supporting arbitrary tensor shapes, masks, padding, reordering, layer fusion, and software pipelining for VLIW-SIMD cores.
- Optimized GPU kernels using CUDA and HIP for finite volume simulations, implementing operator fusion, loop flattening, tiling, achieving 10x performance on single GPU and 20x on multi-GPU setups.
- Implemented wrappers for sparse linear solvers across Linux, Mac, Windows using Clang, MPITrampoline, and MPI at a national research laboratory.
2. Strong understanding of hardware architectures and mapping algorithms to architecture
Mapped computational graphs and algorithms from AI frameworks to diverse hardware targets.
- Optimized deep learning models for neutronics simulations using Torch Compiler and TVM at a national laboratory, reducing memory usage and inference time, benchmarking across 200+ isotopes.
- Developed software pipelines for reactor simulations launching 4000+ instances, implementing MPI and NCCL-based parallel algorithms for generative AI data processing.
- Built ML library for time series classification with time series forest and KNN algorithms during Google Summer of Code, achieving 6x speedup via data randomization and unified APIs.
3. Full-stack toolchain, compiler infrastructure, and hardware-software co-design
Collaborated across teams to optimize compilers, validate silicon, and deploy production AI models.
- Collaborated with architecture, emulation, quantization teams at a major semiconductor company to validate silicon, implement inference flows, and deploy production-grade AI models.
- Developed full research proposal as principal investigator for large language model optimizing training with ZeRO-Infinity and DeepSpeed at a national laboratory.
- Created RangeEnclosures.jl Julia package and range over-approximation algorithms improving bounds by 70%, implementing branch and bound for polynomial bounding at NumFOCUS.
- Implemented Julia packages for range enclosures and branch-and-bound algorithms minimizing relative errors in mathematical modeling.
Requirements & Candidate Alignment
| d-Matrix Requirement | Candidate Qualification |
|---|---|
| Minimum: MS in computer engineering, math, physics, or a related degree with 5+ years of industry experience or a PhD in computer engineering, math, physics, or a related degree with 1+ years of industry experience. | MS in Computer Science from a top research university with 7.3 years of industry experience in deep learning compilers and kernel engineering. |
| Strong grasp of: computer architecture, data structures, system software, and machine learning fundamentals. | Expertise in computer architecture, data structures, algorithms, system software including Linux, MPI, OpenMP, and machine learning fundamentals demonstrated through kernel optimizations and compiler development. |
| Proficient in: C/C++ and Python development in Linux environments and using standard development tools. | Proficient in C/C++, Python, Linux, Git, with experience in CUDA, HIP, Clang, MPITrampoline, and standard tools for algorithm implementation. |
| Experience implementing algorithms in high-level languages such as C/C++ and Python. | Implemented algorithms in C/C++, Python, Julia for ML workloads including sentiment analysis with scikit-learn, NLTK, and time series classification. |
| Experience implementing algorithms for specialized hardware such as FPGAs, DSPs, GPUs, and AI accelerators using libraries such as CUDA, etc. | Developed kernels for GPUs using CUDA, HIP, NCCL, DeepSpeed, and AI accelerators like XDNA architecture with VLIW-SIMD cores at a major semiconductor company. |
| Experience in implementing operators commonly used in ML workloads—GEMMs, Convolutions, BLAS, SIMD operators for operations like softmax, layer normalization, pooling, etc. | Optimized operators including GEMMs, Convolutions, multi-head attention, layer fusion, software pipelining, and SIMD for memory-bound workloads. |
| Experience with development for embedded SIMD vector processors such as Tensilica. | Optimized for VLIW-SIMD cores with automated op-shape mapping, minimizing memory roundtrips, aligning with embedded SIMD vector processors. |
| Self-motivated team player with a strong sense of ownership and leadership. | Led HW/SW co-design projects, served as principal investigator for research proposals, and collaborated across architecture, emulation, and software teams. |
If you would like to discuss this candidate or other critical roles, here is a link to my calendar to schedule a call:
Schedule a CallContact:
Finding the signal in the noise since 2005