Candidate Profile — Candidate #1
Position: d-Matrix — Software Engineer, Staff - Kernels (Santa Clara, CA)
Candidate Location: Not specified
Experience: 15.9 years experience

Relevant Experience & Education Highlights

Extensive expertise in ML compilers, hardware acceleration, and deep learning optimizations positions this candidate strongly for developing and productizing software kernels on next-generation AI hardware at d-Matrix.

1. Experience building software kernels for HW architectures

Demonstrated deep knowledge in mapping algorithms and computational graphs to specialized hardware through compiler stacks and acceleration frameworks.

  • Developed MLIR-enabled compilation stack targeting commodity and in-house ML accelerators at a major autonomous vehicle technology company.
  • Led frontend for PyTorch models to MLIR ecosystem, optimizing for heterogeneous compute platforms in latency-critical inference.
  • Co-created and maintained Tensor Compute Primitives (mlir-tcp), an open-source MLIR dialect for ML programs.
  • Co-maintained Torch-MLIR project providing PyTorch to MLIR compiler support.

2. Experience implementing algorithms for specialized hardware such as FPGAs, DSPs, GPUs, and AI accelerators

Hands-on work optimizing deep learning algorithms for programmable logic and AI engines aligns with hardware-software co-design needs.

  • Advanced ML acceleration algorithms at an FPGA and adaptive computing company, including AI Engines for cloud and edge use cases.
  • Developed TQT, a quantization technique for fixed-point inference on deep neural networks, accepted at a leading conference.
  • Created GraFX, a source-to-source compiler framework automating PyTorch model optimization and deployment.
  • Explored novel architectures for efficient ML workloads on programmable hardware.

3. Experience with ML compilers and algorithms, such as MLIR, LLVM, TVM, Glow

Proven track record building full-stack toolchains for ML model deployment bridges AI frameworks to underlying architectures.

  • Worked at intersection of compilers and algorithms for deep learning optimizations and accelerated deployments.
  • Tech led ML compiler efforts integrating state-of-the-art PyTorch models with MLIR middle-end.
  • Designed software architecture for compiler frameworks enhancing model development lifecycles.
  • Integrated components across ML training and deployment stacks.

Requirements & Candidate Alignment

d-Matrix RequirementCandidate Qualification
Education: MS in computer engineering, math, physics, or a related degree with 5+ years of industry experience or a PhD... with 1+ yearsMS in Electrical Engineering from a top-tier research university with 15.9 years of industry experience
Strong grasp: of computer architecture, data structures, system software, and machine learning fundamentalsStrong alignment through work on hardware architectures, ML compilers, and deep learning optimizations for accelerators
Proficient: in C/C++ and Python development in Linux environments and using standard development toolsProficient in Python with Jupyter Notebook; extensive compiler and framework development implying C/C++ proficiency
Experience implementing algorithms for specialized hardware such as FPGAs, DSPs, GPUs, and AI accelerators using libraries such as CUDA, etc.Direct experience with FPGAs and AI Engines at an adaptive computing company; MLIR stacks for commodity and in-house accelerators
Experience in implementing operators commonly used in ML workloads—GEMMs, Convolutions, BLAS, SIMD operators...Aligned through deep learning optimizations, quantization for neural networks, and PyTorch model deployments
Self-motivated team player with a strong sense of ownership and leadershipStaff-level Tech Lead roles demonstrating ownership in compiler frontends and framework creation
Preferred: Experience with ML frameworks such as TensorFlow and/or PyTorchHands-on with PyTorch including SOTA models, Torch-MLIR, TensorFlow, and GraFX framework
Preferred: Experience with ML compilers and algorithms, such as MLIR, LLVM, TVM, Glow, etc.Extensive with MLIR, including Torch-MLIR, mlir-tcp dialect, and compilation stacks

If you would like to discuss this candidate or other critical roles, here is a link to my calendar to schedule a call:

Schedule a Call
← Previous#1#2#3#4#5Next →

Contact:

Jason Rath

TalentPros.AI

Finding the signal in the noise since 2005

512-993-8228

Jason@TalentPros.AI

Top