Candidate Profile — Candidate #5
Position: d-Matrix — Software Engineer, Staff - Kernels (Santa Clara, CA)
Candidate Location: Morgan Hill, CA
Experience: 9.6 years experience

Relevant Experience & Education Highlights

Demonstrates strong alignment with kernel development for AI hardware through extensive GPU acceleration of CNN inference using CUDA, TensorRT, and TensorFlow/PyTorch, complemented by hardware-software co-design expertise.

1. Development, enhancement, and maintenance of software kernels for next-generation AI hardware

Contributed to productizing AI software stacks in embedded devices and high-performance GPU environments.

  • Accelerated convolutional neural networks inference using Nvidia GPU ecosystem including TensorRT, ONNX, and CUDA kernels at a leading biotech diagnostics company.
  • Developed GPU acceleration for neural network algorithms in post-primary analysis on multiple A100 and H100 GPUs.
  • Architected deep convolutional neural networks with hardware-software co-verification and silicon architecture focus.
  • Coded video signal processing algorithms in C and OpenCL for Intel CodeBuilder and Nvidia SDK environments.

2. Experience implementing algorithms for specialized hardware such as FPGAs, DSPs, GPUs, and AI accelerators using libraries such as CUDA, etc.

Applied GPGPU programming expertise to optimize ML workloads on Nvidia GPUs.

  • Implemented GPU kernels using CUDA, TensorRT, ONNX, Trtexec, and Nsys profiling for AI inference.
  • Utilized OpenCL for video signal processing on Nvidia and Intel hardware platforms.
  • Deployed vLLM for inference of large language models alongside GPU acceleration.
  • Developed algorithms for multi-node and multi-GPU settings at a leading semiconductor company.

3. Experience with ML frameworks such as TensorFlow and/or PyTorch

Leveraged major deep learning frameworks for computer vision, NLP, and inference tasks.

  • Built TensorFlow infrastructure for NLP algorithms including Google-BERT and Transformers in multi-GPU environments.
  • Employed TensorFlow and PyTorch with associated Python frameworks for neural network acceleration.
  • Applied unsupervised machine learning for texture synthesis in video codec development at a major consumer electronics company.
  • Developed computer vision techniques using OpenCV 3.0 for plant measurement and tracking at an agritech company.

4. Experience in implementing operators commonly used in ML workloads—GEMMs, Convolutions, BLAS, SIMD operators for operations like softmax, layer normalization, pooling, etc.

Focused on convolution-heavy workloads and multidimensional signal processing relevant to ML operators.

  • Accelerated convolutions in CNNs for real-time inference in nanopore sequencing applications.
  • Implemented pattern recognition, segmentation, and region growing in multidimensional signal processing for image and video.
  • Developed blur recognition, de-convolution, noise reduction, and white balance in camera signal processing.
  • Architected vehicle inspection systems capturing structural defects via image processing algorithms.

Requirements & Candidate Alignment

d-Matrix RequirementCandidate Qualification
Minimum: MS in computer engineering, math, physics, or a related degree with 5+ years of industry experience or a PhD in computer engineering, math, physics, or a related degree with 1+ years of industry experience.PhD in Electrical Engineering from a top research university with 9.6 years of industry experience.
Strong grasp of computer architecture, data structures, system software, and machine learning fundamentals.Deep expertise in computer architecture, silicon architecture, hardware-software co-design, and deep learning fundamentals.
Proficient in C/C++ and Python development in Linux environments and using standard development tools.Proficient in C/C++, Python, VS Code, and Linux-based tools including Nvidia SDK and Intel CodeBuilder.
Experience implementing algorithms in high-level languages such as C/C++ and Python.Extensive experience implementing video processing, computer vision, and ML algorithms in C/C++ and Python.
Experience implementing algorithms for specialized hardware such as FPGAs, DSPs, GPUs, and AI accelerators using libraries such as CUDA, etc.Strong hands-on experience with GPUs and AI accelerators using CUDA, OpenCL, TensorRT, and ONNX.
Experience in implementing operators commonly used in ML workloads—GEMMs, Convolutions, BLAS, SIMD operators for operations like softmax, layer normalization, pooling, etc.Proven implementation of convolutions and related operators in CNN inference and multidimensional signal processing.
Experience with ML frameworks such as TensorFlow and/or PyTorch.Hands-on work with TensorFlow infrastructure and PyTorch for NLP, CV, and inference workloads.
Self-motivated team player with a strong sense of ownership and leadership.Demonstrated leadership in roles such as GPU Principal Engineer and Director of Image Processing and Algorithms.

If you would like to discuss this candidate or other critical roles, here is a link to my calendar to schedule a call:

Schedule a Call

Contact:

Jason Rath

TalentPros.AI

Finding the signal in the noise since 2005

512-993-8228

Jason@TalentPros.AI

Top