Candidate Profile — Candidate #3
Position: d-Matrix — Software Engineer, Staff - Kernels (Santa Clara, CA)
Candidate Location: San Francisco Bay Area
Experience: 34.2 years experience

Relevant Experience & Education Highlights

Extensive expertise in compiler infrastructure, hardware-software co-design, and optimization for AI accelerators positions this candidate as an exceptional fit for developing and scaling software kernels at d-Matrix.

1. Development, enhancement, and maintenance of software kernels for next-generation AI hardware

Proven track record in implementing and optimizing kernels across diverse AI hardware platforms.

  • Implemented LLVM optimizations and code generation on wafer-scale tensor processing unit at an AI accelerator company.
  • Drove optimizations within Halide language to leverage GPU texture hardware and DSP scatter/gather at a specialized imaging processing firm.
  • Conducted performance engineering on ARM NEON for touch sensor processing, improving performance by 200% through C representations and machine code at a mobile computing company.
  • Optimized OpenCL performance for Snapdragon GPUs and provided LLVM compiler feedback on hardware specs at a major semiconductor provider.

2. Experience building software kernels for HW architectures and mapping algorithms to architecture

Deep understanding of mapping computational graphs and algorithms to specialized hardware like DSPs, GPUs, and AI accelerators.

  • Built MLIR pass generating Risc-V Vector intrinsics from MLIR Affine dialect, improving vector recognition and lowering to LLVM for pre-tapeout AI chip at a cutting-edge AI computing startup.
  • Coordinated kernel scheduler for data cluster serving petabytes to over 64K processors, establishing metrics for GEMM components and data transfer at an innovative AI hardware company.
  • Implemented optimizations for ARM CPUs, Snapdragon Hexagon DSP, Adreno GPUs, and CEVA XM4 DSPs at a digital signal processing firm.
  • Designed Python algorithms identifying patterns in data-flow graphs from TensorFlow for ML algorithms on Cerebras CS-1 at a wafer-scale engine company.

3. Compiler infrastructure and experience with ML compilers such as MLIR, LLVM

Hands-on work with MLIR, LLVM, and related tools aligns directly with building compiler infrastructure for AI workloads.

  • Implemented control flow and math optimizations for P4/P4+/P416 language using LLVM, coordinating ISA specification with hardware designers at a high-performance networking company.
  • Performed performance analysis of BERT models and generated CI/CD perf tests for optimization changelogs at an AI research lab.
  • Implemented instructions and optimizations for Float, BFloat16, FP8 data types using state-of-the-art LLVM and MLIR at an AI computing startup.
  • Broke down MNIST, BERT-Large, DLRM, GPT-2 tasks to schedule massively parallel asynchronous tasks on AI hardware.

4. Experience implementing operators commonly used in ML workloads—GEMMs, Convolutions, BLAS, SIMD operators

Strong foundation in ML operators and SIMD processing demonstrated through practical implementations.

  • Established metrics for GEMM components, Risc-V hosts, and data transfer subsystems tailored for AI card performance.
  • Collaborated on ISA simplification and optimization for scalar and vector instruction lanes supporting ML workloads.
  • Optimized code base for BLAS-like operations and SIMD on DSPs and GPUs including Adreno and Hexagon.
  • Added ARM NEON machine code to critical sections for SIMD-accelerated sensor processing.

Requirements & Candidate Alignment

d-Matrix RequirementCandidate Qualification
Minimum: MS in computer engineering, math, physics, or a related degree with 5+ years of industry experience or a PhD... with 1+ years of industry experience34.2 years of industry experience in compiler engineering and AI hardware optimization, complemented by B.S. in Computer Science from a top-tier public university
Strong grasp of computer architecture, data structures, system software, and machine learning fundamentalsExtensive expertise in computer architecture and ML fundamentals, evidenced by Risc-V vectorization, data-flow graph analysis, and BERT performance tuning
Proficient in C/C++ and Python development in Linux environments and using standard development toolsProficient in C/C++, Python, LLVM, and MLIR, applied in Linux-based optimizations for AI chips and DSPs
Experience implementing algorithms in high-level languages such as C/C++ and PythonImplemented algorithms in C/C++ and Python for data encapsulation patterns in TensorFlow-generated ML graphs
Experience implementing algorithms for specialized hardware such as FPGAs, DSPs, GPUs, and AI accelerators using libraries such as CUDA, etc.Implemented for GPUs (Adreno, Snapdragon), DSPs (Hexagon, CEVA XM4), and AI accelerators (wafer-scale TPU, Risc-V AI chips) using LLVM, OpenCL, and Halide
Experience in implementing operators commonly used in ML workloads—GEMMs, Convolutions, BLAS, SIMD operators for operations like softmax, layer normalization, pooling, etc.Implemented GEMM metrics, SIMD vector intrinsics, and optimizations for BFloat16/FP8 in ML tasks like BERT and MNIST
Experience with development for embedded SIMD vector processors such as TensilicaDeveloped for embedded SIMD including ARM NEON, Risc-V Vector, Hexagon DSP, aligning with vector processor workflows
Self-motivated team player with a strong sense of ownership and leadershipLed cross-team coordination on ISA design, kernel scheduling, and hardware-software optimizations across multiple AI startups

If you would like to discuss this candidate or other critical roles, here is a link to my calendar to schedule a call:

Schedule a Call

Contact:

Jason Rath

TalentPros.AI

Finding the signal in the noise since 2005

512-993-8228

Jason@TalentPros.AI

Top