Candidate Profile — Candidate #5
Position: Adobe — Machine Learning Infrastructure Engineer
Candidate Location: San Francisco Bay Area
Experience: 13.6 years experience

Relevant Experience & Education Highlights

Deep expertise in GPU optimization, PyTorch and TensorFlow performance engineering, custom CUDA kernels, and large-scale ML model training aligns strongly with building scalable infrastructures for Adobe's Firefly generative AI foundation models.

1. You'll build and optimize infrastructures that power large foundation model training on thousands of GPUs

Demonstrated leadership in scaling ML training and inference infrastructures with GPU acceleration matches core responsibilities for foundation model stacks.

  • Led ML GPU optimization efforts at a leading augmented reality and social media company, focusing on high-performance production environments.
  • Improved TensorFlow and PyTorch performance on GPUs at a premier technology and AI company, optimizing models including BERT, ResNet50, and NCF with official MLPerf submissions.
  • Developed custom CUDA kernels for critical bottlenecks, enhancing training and inference throughput in large-scale ML workflows.
  • Engineered features and leveraged LLMs in recommendation systems at a major ads platform within the same premier company, driving model quality and revenue.

2. You will profile GPU utilization, trace inference and training runs and help craft strategies for optimizing our ML model latency

Hands-on experience with profiling tools and optimization techniques directly supports GPU utilization analysis and latency reduction strategies.

  • Profiled performance using Nsight and optimized GPU memory bandwidth, kernel fusion, and model quantization (INT8/FP8) to maximize throughput.
  • Deployed models with TensorRT and Triton Inference Server, slashing inference latency in production environments.
  • Bridged model development and hardware by transforming computationally expensive models into efficient, low-latency applications across the ML lifecycle.

3. We'll work together to architect and optimize end-to-end ML pipelines, ensuring they're scalable, efficient, and robust

Full lifecycle experience from architecture design to deployment aligns with architecting scalable ML pipelines and systems optimizations.

  • Architected novel deep learning models in PyTorch and JAX, spanning design, training, and scalable deployment.
  • Applied distillation, transformers, and performance engineering techniques, including work on BERT and recommendation models with LLMs.
  • Contributed to GPU modeling and simulation at a prominent consumer technology company and a major semiconductor firm, enhancing hardware utilization.

Requirements & Candidate Alignment

Adobe RequirementCandidate Qualification
Education: Graduate, PhD, or postgraduate degree in Computer Science, Computer Engineering, or a related field—or equivalent experienceDual Master's degrees in Computer Science from a top-tier research university and Electrical Engineering from a premier technical institute
Experience: 2+ years ML Engineering experience, specializing in generative AI like LLMs13.6 years total experience including ML engineering leadership and specialization in LLMs, transformers, and performance optimization
Technical Skills: Strong Python and deep learning engineering skills, paired with experience in training and inferencing with PyTorch or TensorFlowStrong Python, PyTorch, TensorFlow, and JAX expertise demonstrated in training, inferencing, and optimizing models like BERT and ResNet50
Domain Knowledge: Familiarity with distillation, transformers, and diffusion models. Experience with generative image and video is a plusHands-on with distillation, transformers (BERT), and LLMs in generative AI contexts, including recommendation systems
Deployment Technologies: Knowledge of deployment technologies such as Docker, ML Ops, and ML services is valuableDeployment experience with TensorRT, Triton Inference Server, and MLOps workflows for scalable production ML systems
Cloud Platforms: Experience with cloud platforms like Azure and AWS is a plusFull ML lifecycle scaling supports cloud-adjacent infrastructures, with production deployment expertise
Problem-Solving: Excellent problem-solving abilities and capacity to analyze complex issues and drive solutions with a data-driven approachData-driven optimizations via profiling, feature engineering, and performance modeling across GPU and ML systems

If you would like to discuss this candidate or other critical roles, here is a link to my calendar to schedule a call:

Schedule a Call

Contact:

Jason Rath

TalentPros.AI

Finding the signal in the noise since 2005

512-993-8228

Jason@TalentPros.AI

Top