Candidate Profile — Candidate #4
Position: d-Matrix — Software Engineer, Staff - Kernels (Santa Clara, CA)
Candidate Location: San Jose, CA
Experience: 7.3 years experience

Relevant Experience & Education Highlights

Demonstrates deep expertise in developing compilers and kernels for AI hardware architectures, including GEMM, convolutions, and attention mechanisms, with strong alignment to d-Matrix's requirements for software kernels on next-generation AI compute engines.

1. Development, enhancement, and maintenance of software kernels for next-generation AI hardware

Built and optimized kernels for AI accelerators, GPUs, and specialized processors.

  • Performed HW/SW co-design for AI architecture at a major semiconductor company, creating data-movement compiler maximizing memory-hierarchy efficiency for high-throughput GEMM, convolution, and multi-head attention workloads.
  • Developed flexible attention compiler supporting arbitrary tensor shapes, masks, padding, reordering, layer fusion, and software pipelining for VLIW-SIMD cores.
  • Optimized GPU kernels using CUDA and HIP for finite volume simulations, implementing operator fusion, loop flattening, tiling, achieving 10x performance on single GPU and 20x on multi-GPU setups.
  • Implemented wrappers for sparse linear solvers across Linux, Mac, Windows using Clang, MPITrampoline, and MPI at a national research laboratory.

2. Strong understanding of hardware architectures and mapping algorithms to architecture

Mapped computational graphs and algorithms from AI frameworks to diverse hardware targets.

  • Optimized deep learning models for neutronics simulations using Torch Compiler and TVM at a national laboratory, reducing memory usage and inference time, benchmarking across 200+ isotopes.
  • Developed software pipelines for reactor simulations launching 4000+ instances, implementing MPI and NCCL-based parallel algorithms for generative AI data processing.
  • Built ML library for time series classification with time series forest and KNN algorithms during Google Summer of Code, achieving 6x speedup via data randomization and unified APIs.

3. Full-stack toolchain, compiler infrastructure, and hardware-software co-design

Collaborated across teams to optimize compilers, validate silicon, and deploy production AI models.

  • Collaborated with architecture, emulation, quantization teams at a major semiconductor company to validate silicon, implement inference flows, and deploy production-grade AI models.
  • Developed full research proposal as principal investigator for large language model optimizing training with ZeRO-Infinity and DeepSpeed at a national laboratory.
  • Created RangeEnclosures.jl Julia package and range over-approximation algorithms improving bounds by 70%, implementing branch and bound for polynomial bounding at NumFOCUS.
  • Implemented Julia packages for range enclosures and branch-and-bound algorithms minimizing relative errors in mathematical modeling.

Requirements & Candidate Alignment

d-Matrix RequirementCandidate Qualification
Minimum: MS in computer engineering, math, physics, or a related degree with 5+ years of industry experience or a PhD in computer engineering, math, physics, or a related degree with 1+ years of industry experience.MS in Computer Science from a top research university with 7.3 years of industry experience in deep learning compilers and kernel engineering.
Strong grasp of: computer architecture, data structures, system software, and machine learning fundamentals.Expertise in computer architecture, data structures, algorithms, system software including Linux, MPI, OpenMP, and machine learning fundamentals demonstrated through kernel optimizations and compiler development.
Proficient in: C/C++ and Python development in Linux environments and using standard development tools.Proficient in C/C++, Python, Linux, Git, with experience in CUDA, HIP, Clang, MPITrampoline, and standard tools for algorithm implementation.
Experience implementing algorithms in high-level languages such as C/C++ and Python.Implemented algorithms in C/C++, Python, Julia for ML workloads including sentiment analysis with scikit-learn, NLTK, and time series classification.
Experience implementing algorithms for specialized hardware such as FPGAs, DSPs, GPUs, and AI accelerators using libraries such as CUDA, etc.Developed kernels for GPUs using CUDA, HIP, NCCL, DeepSpeed, and AI accelerators like XDNA architecture with VLIW-SIMD cores at a major semiconductor company.
Experience in implementing operators commonly used in ML workloads—GEMMs, Convolutions, BLAS, SIMD operators for operations like softmax, layer normalization, pooling, etc.Optimized operators including GEMMs, Convolutions, multi-head attention, layer fusion, software pipelining, and SIMD for memory-bound workloads.
Experience with development for embedded SIMD vector processors such as Tensilica.Optimized for VLIW-SIMD cores with automated op-shape mapping, minimizing memory roundtrips, aligning with embedded SIMD vector processors.
Self-motivated team player with a strong sense of ownership and leadership.Led HW/SW co-design projects, served as principal investigator for research proposals, and collaborated across architecture, emulation, and software teams.

If you would like to discuss this candidate or other critical roles, here is a link to my calendar to schedule a call:

Schedule a Call

Contact:

Jason Rath

TalentPros.AI

Finding the signal in the noise since 2005

512-993-8228

Jason@TalentPros.AI

Top