Candidate Profile — Candidate #1
Position: Figure — AI Training Infrastructure Engineer - Helix Team (San Jose, CA)
Candidate Location: Seattle Metro
Experience: 13.7 years experience

Relevant Experience & Education Highlights

Extensive experience building scalable AI training infrastructure, including distributed LLM training with PyTorch, GPU clusters managed by Kubernetes and Slurm, and MLOps tools, positions this candidate strongly for advancing Figure's Helix Team training systems.

1. Design, deploy, and maintain training clusters

Managed high-performance computing environments and GPU clusters optimized for deep learning workloads.

  • Built next-gen AI cloud infrastructure featuring GPU clusters orchestrated with Kubernetes at a stealth AI startup.
  • Developed ML serving platforms incorporating Ray, Kubernetes, and vector databases for large-scale model deployment.
  • Led construction of unified control planes for machine learning platforms supporting multi-GPU distributed training at a major ride-sharing company.

2. Architect and maintain scalable deep learning frameworks for training on massive robot datasets

Architected frameworks and services for training and inference of large language models and transformer architectures.

  • Implemented LLM distributed training, acceleration, and performance modeling using PyTorch, JAX, and Triton at a stealth AI startup.
  • Designed high-throughput ML inference services with Ray and Kubernetes, including scalable Retrieval-Augmented Generation applications with LLMs and vector databases at an AI/ML company.
  • Developed custom Triton kernels to accelerate LLM training and inference, alongside quantization and tensor parallelism optimizations.

3. Implement distributed training and parallelization strategies to reduce model development cycles

Explored and applied advanced parallelism techniques to enhance training efficiency across distributed systems.

  • Explored model parallelism strategies to accelerate LLM training and inference using PyTorch and JAX.
  • Optimized model inference through tensor parallelism, continuous batching, and quantization in production environments.
  • Built MLOps orchestration frameworks with Spark data processing, multi-GPU training, and job failure recovery at a major ride-sharing company.

4. Implement tooling for data processing, model experimentation, and continuous integration

Created developer tools, pipelines, and observability for streamlined AI model development workflows.

  • Developed CLI tools, workflow schedulers, and pipeline observability for model development at a major ride-sharing company.
  • Built A/B testing platforms and data analytics pipelines using Spark, Kafka, Cassandra, and Redis at a global information services company.
  • Researched capacity planning tools and performance modeling for LLMs, enhancing experimentation cycles.

Requirements & Candidate Alignment

Figure RequirementCandidate Qualification
Strong software engineering fundamentals13.7 years building reliable backend systems, including distributed ML platforms and high-throughput services
Bachelor's or Master's degree in Computer Science, Robotics, Engineering, or a related fieldMaster's degree in Computer & Electrical Engineering from a top research university; Bachelor's degree in Microelectronics from a leading engineering university in China
Experience with Python and PyTorchExtensive hands-on experience with Python and PyTorch for LLM distributed training, performance modeling, and deep learning frameworks
Experience managing HPC clusters for deep neural network trainingProven track record managing GPU clusters with Kubernetes and Slurm for LLM training and inference acceleration
Minimum of 4 years of professional, full-time experience building reliable backend systems13.7 years of professional experience across AI platforms, MLOps, and distributed systems engineering roles
Experience managing cloud infrastructure (AWS, Azure, GCP)Direct experience with AWS and GCP in building scalable AI infrastructure and services
Experience with job scheduling / orchestration tools (SLURM, Kubernetes, LSF, etc.)Strong expertise with Kubernetes, Slurm, and Ray for orchestrating distributed training and inference workloads

If you would like to discuss this candidate or other critical roles, here is a link to my calendar to schedule a call:

Schedule a Call
← Previous#1#2#3#4#5Next →

Contact:

Jason Rath

TalentPros.AI

Finding the signal in the noise since 2005

512-993-8228

Jason@TalentPros.AI

Top