Candidate Profile — Candidate #1
Position: Adobe — Machine Learning Infrastructure Engineer
Candidate Location: Santa Clara County, California
Experience: 21.6 years experience

Relevant Experience & Education Highlights

Extensive expertise in large-scale LLM inference and GPU-accelerated ML infrastructure at leading technology companies positions this candidate to excel in building scalable PyTorch training systems and optimizing foundation model pipelines for Adobe's Firefly generative AI efforts.

1. Build and optimize infrastructures that power large foundation model training on thousands of GPUs

Led development of high-scale ML serving systems handling massive models across distributed GPU and TPU clusters.

  • Led large-scale LLM inference and serving using PyTorch, JAX, and XLA on TPUs and GPUs at a major technology company.
  • Engineered high-throughput, low-latency inference for deep learning ranking and recommendation models up to hundreds of GBs on Nvidia A100 GPUs and custom accelerators at a prominent social media company.
  • Scaled distributed CPU and GPU inference solutions for realtime model serving and online training of PyTorch-based models up to hundreds of GBs.
  • Hosted systems across thousands of machines globally to serve billions of queries per second with 100-200ms latency.

2. Profile GPU utilization, trace inference and training runs and help craft strategies for optimizing our ML model latency

Optimized inference pipelines for low-latency, high-throughput performance on advanced GPU hardware.

  • Developed model validation and evaluation solutions for large PyTorch GPU models, ensuring optimal utilization and performance.
  • Implemented model freshness mechanisms and realtime inference for hundreds-of-GB models on Nvidia A100s.
  • Led optimizations for GPU inference on custom accelerators, achieving low-latency serving at massive scale.

3. Architect and optimize end-to-end ML pipelines, ensuring they're scalable, efficient, and robust

Architected distributed systems and indices supporting petabyte-scale data and high-traffic workloads.

  • Built disaggregated in-memory secondary indices using petabytes of RAM and Flash to host thousands of indices, serving 40% of online traffic at a prominent social media company.
  • Led team of 10 engineers in building advertiser bidding, budget pacing, and ads candidate generation infrastructure for millions of advertisers.
  • Scaled payment processing and reconciliation infrastructure to support 100x volume growth at a global fintech company.

4. Engage in architecture, design, deployment, and optimizations of ML models and systems throughout the product lifecycle

Demonstrated leadership in end-to-end infrastructure for AI/ML and high-scale services.

  • Led implementation of payments infrastructure integrated across major platforms.
  • Developed query processing engines handling one-quarter of online transaction processing queries.
  • Led AI/ML infrastructure teams focusing on model serving, validation, and deployment at scale.

Requirements & Candidate Alignment

Adobe RequirementCandidate Qualification
Education: Graduate, PhD, or postgraduate degree in Computer Science, Computer Engineering, or a related field—or equivalent experience.MS in Computer Science from a top research university, plus additional masters and incomplete PhD in Computer Science.
ML Engineering Experience: 2+ years ML Engineering experience, specializing in generative AI like LLMs.5+ years in AI/ML infrastructure roles, including large-scale LLM inference and serving with PyTorch on GPUs and TPUs.
Technical Skills: Strong Python and deep learning engineering skills, paired with experience in training and inferencing with PyTorch or TensorFlow.Deep PyTorch expertise in inference, training, validation, JAX, and serving for large-scale deep learning models up to hundreds of GBs.
Model Familiarity: Familiarity with distillation, transformers, and diffusion models. Experience with generative image and video is a plus.Strong transformers experience through large-scale LLMs, ranking, and recommendation models.
Deployment Knowledge: Knowledge of deployment technologies such as Docker, ML Ops, and ML services is valuable.ML Ops proficiency in high-throughput model serving, distributed inference, and realtime deployment pipelines.
Cloud Platforms: experience with cloud platforms like Azure and AWS is a plus.Distributed systems scaling across thousands of global machines for petabyte-scale ML workloads.
Problem-Solving: excellent problem-solving abilities and your capacity to analyze complex issues and drive solutions with a data-driven approach.Proven optimization leadership in low-latency inference and massive-scale indexing systems.
Communication: strong verbal and written communication skills and success in cross-functional team environments.Led multiple teams of 10+ engineers across AI infra, ads delivery, and payments projects.

If you would like to discuss this candidate or other critical roles, here is a link to my calendar to schedule a call:

Schedule a Call

Contact:

Jason Rath

TalentPros.AI

Finding the signal in the noise since 2005

512-993-8228

Jason@TalentPros.AI

Top