Candidate Profile — Candidate #2
Position: d-Matrix — Software Engineer, Staff - Kernels (Santa Clara, CA)
Candidate Location: San Francisco Bay Area
Experience: 19.9 years experience

Relevant Experience & Education Highlights

Expertise in developing high-performance GPU kernels, optimizing ML operators like GEMMs and convolutions using CUDA and MLIR/LLVM, and experience in small-team environments positions this candidate as an exceptional fit for productizing software kernels on d-Matrix's next-generation AI hardware.

1. Development, enhancement, and maintenance of software kernels for next-generation AI hardware

Extensive track record building and optimizing kernels for GPU architectures aligns directly with kernel development for AI compute engines.

  • Developed CUDA kernels in CUTLASS library at a leading GPU company to enable deep learning primitives on NVIDIA Ampere, Turing, and Volta architectures targeting CUDA and Tensor Cores.
  • Implemented forward and backward convolutions for NVIDIA GPUs, including Implicit GEMM convolution presented at major industry conference.
  • Optimized FP8 GEMMs with blockwise scaling on NVIDIA H100 GPUs at a major tech company, extending to groupwise scaling and adopted in large-scale models.
  • Improved GPU performance for LLMs including Llama 70B and 405B on NVIDIA H100 through scaling techniques ensuring numerical accuracy.

2. Experience implementing algorithms for specialized hardware such as FPGAs, DSPs, GPUs, and AI accelerators using libraries such as CUDA

Proven ability to map algorithms to GPU hardware using CUDA and related tools matches requirements for AI accelerator optimization.

  • Engineered codegen for GEMMs on NVIDIA A100 Tensor Cores using OpenXLA/IREE MLIR compiler at a major tech company, matching CUTLASS and cuBLAS performance.
  • Boosted half-precision performance from 144 to 238 TFLOPs and single-precision from 77 to 118 TFLOPs on NVIDIA A100, adding support for batch matmul, split-k, bfloat16, and mixed-precision.
  • Developed compiler passes using LLVM for GPU kernels in C++AMP at a leading GPU company.
  • Scoped and evaluated codegen efforts for NVIDIA Hopper architecture to enhance GPU programmability for LLMs.

3. Experience in implementing operators commonly used in ML workloads—GEMMs, Convolutions, BLAS, SIMD operators

Deep hands-on work with core ML operators on GPUs directly supports mapping computational graphs to d-Matrix hardware.

  • Implemented GEMMs, convolutions, and related BLAS operations via CUTLASS and OpenXLA at multiple major companies, targeting Tensor Cores.
  • Added support for advanced features like batch matmul, split-k, and mixed-precision datatypes in MLIR/LLVM codegen.
  • Optimized performance for large models including Llama 70B/405B and Deepseek V3/R1 using FP8 scaling techniques on H100 GPUs.

4. Experience with ML compilers and algorithms, such as MLIR, LLVM, TVM, Glow

Strong compiler infrastructure background, including collaboration with experts, fits building d-Matrix's toolchain.

  • Developed MLIR/LLVM codegen in OpenXLA compiler for NVIDIA GPUs at a major tech company.
  • Contributed to IREE MLIR compiler for high-performance GEMM codegen matching vendor libraries.
  • Worked in small-team startup environment scaling test-time compute, long context, and reinforcement learning at current role.

Requirements & Candidate Alignment

d-Matrix RequirementCandidate Qualification
Minimum: MS in computer engineering, math, physics, or a related degree with 5+ years of industry experience or a PhD in computer engineering, math, physics, or a related degree with 1+ years of industry experiencePhD in Computer Science from a top research university with 19.9 years of industry experience
Strong grasp of computer architecture, data structures, system software, and machine learning fundamentalsExpertise in computer architecture, algorithms, data structures, system software, embedded systems, and machine learning
Proficient in C/C++ and Python development in Linux environments and using standard development toolsProficient in C, C++, Python, compilers, debugging, and programming in Linux-compatible environments including CUDA and OpenCL
Experience implementing algorithms in high-level languages such as C/C++ and PythonImplemented algorithms in C/C++ for GPU kernels, MLIR/LLVM codegen, and deep learning primitives
Experience implementing algorithms for specialized hardware such as FPGAs, DSPs, GPUs, and AI accelerators using libraries such as CUDA, etc.Developed for GPUs and DSPs using CUDA, CUTLASS, C++AMP, OpenCL, including NVIDIA A100, H100, Hopper, Ampere, Turing, Volta
Experience with development for embedded SIMD vector processors such as TensilicaExperience with digital signal processors, embedded systems, firmware, and SIMD operations
Experience with ML compilers and algorithms, such as MLIR, LLVM, TVM, Glow, etc.Hands-on with MLIR, LLVM, OpenXLA, IREE for GPU codegen and optimization
Prior startup, small team, or incubation experienceCurrent role in small-team startup focused on test-time compute, long context, and reinforcement learning

If you would like to discuss this candidate or other critical roles, here is a link to my calendar to schedule a call:

Schedule a Call

Contact:

Jason Rath

TalentPros.AI

Finding the signal in the noise since 2005

512-993-8228

Jason@TalentPros.AI

Top