Candidate Profile — Candidate #6
Position: Adobe — Machine Learning Infrastructure Engineer
Candidate Location: San Francisco Bay Area
Experience: 8.2 years experience

Relevant Experience & Education Highlights

Demonstrates deep expertise in high-performance ML infrastructures, GPU optimization, and generative AI capacity planning, making a strong fit for building scalable foundation model systems at Adobe.

1. Build and optimize infrastructures that power large foundation model training on thousands of GPUs

Led development of hyperscale ML systems integrating hardware and software for generative AI workloads.

  • Led technical initiatives in ML systems and cloud organization at a major search and cloud company, designing low-level computer systems interacting with kernel and hardware for ML use cases.
  • Oversaw end-to-end hardware infrastructure for generative AI at a leading social media company, managing ordering, new product introduction, and fulfillment processes.
  • Designed and optimized HPC and generative AI infrastructure at a semiconductor industry leader, leading team to ensure scalable and cost-efficient resource utilization.
  • Built GPU infrastructure and accelerated computing pipelines at a major telecommunications company, automating GPU allocation, testing, and certification processes.

2. Profile GPU utilization, trace inference and training runs and help craft strategies for optimizing our ML model latency

Specialized in performance engineering and GPU characterization to resolve bottlenecks and enhance efficiency.

  • Conducted performance analysis and optimization for AI workloads on hardware at a semiconductor industry leader, focusing on benchmarking, code optimization, and power engineering.
  • Led identification and resolution of performance bottlenecks across platforms at a major telecommunications company, implementing large-scale testing and profiling strategies.
  • Orchestrated performance improvements such as TLB shootdowns and MobileMark workload studies in system power and performance roles.

3. Architect and optimize end-to-end ML pipelines, ensuring they're scalable, efficient, and robust

Drove capacity planning, systems integration, and cross-functional collaboration for AI-driven solutions.

  • Led capacity planning for generative AI models at a leading social media company, developing, optimizing, and scaling systems powering diverse products.
  • Partnered across ML, hardware, SRE, and cloud teams at a major search and cloud company using C++ and Go for full-stack systems including cluster management and Linux kernel.
  • Architected AI-driven solutions for performance analysis optimizing hardware handling of AI models, algorithms, and workloads.

Requirements & Candidate Alignment

Adobe RequirementCandidate Qualification
Education: Graduate, PhD, or postgraduate degree in Computer Science, Computer Engineering, or a related field—or equivalent experience.PhD from a top-tier research university.
Experience: 2+ years ML Engineering experience, specializing in generative AI like LLMs.8.2 years of ML engineering experience, including leadership in generative AI capacity planning and infrastructure.
Programming and Frameworks: Strong Python and deep learning engineering skills, paired with experience in training and inferencing with PyTorch or TensorFlow.Strong Python and deep learning engineering skills applied in performance engineering, GPU characterization, and AI infrastructure roles.
Model Familiarity: Familiarity with distillation, transformers, and diffusion models. Experience with generative image and video is a plus.Deep expertise in transformers and diffusion models through generative AI optimization and scaling efforts.
Deployment Technologies: Knowledge of deployment technologies such as Docker, ML Ops, and ML services.MLOps proficiency including Docker from building scalable ML pipelines and infrastructure automation.
Cloud Platforms: Experience with cloud platforms like Azure and AWS is a plus.Hands-on experience with AWS in hyperscale cloud ML infrastructures and Google Cloud.
Problem-Solving: Excellent problem-solving abilities and capacity to analyze complex issues and drive solutions with a data-driven approach.Proven data-driven problem-solving in performance bottleneck resolution, capacity planning, and AI workload optimization.
Communication: Strong verbal and written communication skills and success in cross-functional team environments.Strong cross-functional collaboration with ML, hardware, SRE, and cloud teams in technical leadership roles.

If you would like to discuss this candidate or other critical roles, here is a link to my calendar to schedule a call:

Schedule a Call

Contact:

Jason Rath

TalentPros.AI

Finding the signal in the noise since 2005

512-993-8228

Jason@TalentPros.AI

Top