Candidate Profile — Candidate #4
Position: Figure — AI Training Infrastructure Engineer - Helix Team (San Jose, CA)
Candidate Location: San Francisco Bay Area
Experience: 15.7 years experience

Relevant Experience & Education Highlights

Proven AI infrastructure engineer with extensive experience scaling HPC clusters for massive deep learning models and implementing distributed GPU training systems, directly aligning with Figure's Helix team requirements for training infrastructure.

1. Design, deploy, and maintain training clusters

Excelled in architecting and scaling high-performance computing clusters to support exascale AI training workloads.

  • Scaled HPC cluster infrastructure from petaFLOPs for 100M parameter models to several exaFLOPs for 70B-175B parameter foundational models at a leading wafer-scale AI hardware company.
  • Designed HPC infrastructure plans including networking layouts, file system configurations, and GPU integrations to meet high-throughput demands in semiconductor image processing at a major manufacturing firm.
  • Supported Slurm and Kubernetes-based GPU stacks with HGX-H100 fleets, handling customer onboarding, data migrations, and scale-related networking challenges at a stealth GPU cloud startup.
  • Triaged production reliability issues across GPU hardware, scheduling, networking, and storage via trace and log analysis for optimal training and inference performance.

2. Implement distributed training and parallelization strategies

Mastered parallel programming and distributed systems to accelerate large-scale model training cycles.

  • Applied parallel algorithms, MPI, OpenMP, and CUDA expertise in high-performance computing environments for deep neural network training.
  • Assisted in early prototype cluster design, provisioning, monitoring, deployment, and customer-site support for wafer-scale systems handling massive models.
  • Optimized training/inference stacks across GPU fleets using load balancing and RoCEv2 networking to overcome scale challenges.

3. Implement tooling for data processing, model experimentation, and continuous integration

Developed internal tools and solutions engineering practices to enhance developer productivity and reliability in AI workflows.

  • Built internal tooling and documentation for platform dependencies, data processing, and customer production issue resolution at a stealth GPU infrastructure startup.
  • Shaped product roadmaps through customer learnings, driving GPU-based cloud infrastructure improvements for sustained high-performance training.
  • Leveraged Python alongside Linux, simulations, and high-performance computing skills for backend systems reliability.

Requirements & Candidate Alignment

Figure RequirementCandidate Qualification
Strong software engineering fundamentalsDeep expertise in algorithms, data structures, parallel programming, C/C++, Python, and high-performance computing
Bachelor's or Master's degree in Computer Science, Robotics, Engineering, or a related fieldPhD in Computer Science from a top-tier research university; Bachelor's in Information Technology from a prominent engineering university
Experience with Python and PyTorchHands-on experience with Python in AI infrastructure and deep learning systems
Experience managing HPC clusters for deep neural network trainingExtensive track record scaling HPC/GPU clusters from petaFLOPs to exaFLOPs for foundational models up to 175B parameters
Minimum of 4 years of professional, full-time experience building reliable backend systems15.7 years building reliable HPC, AI infrastructure, and backend systems across multiple roles
Experience with job scheduling / orchestration tools (SLURM, Kubernetes, LSF, etc.)Proven work with SLURM for GPU cluster management and orchestration
Experience managing cloud infrastructure (AWS, Azure, GCP)Familiarity with for cloud-based HPC and AI workloads

If you would like to discuss this candidate or other critical roles, here is a link to my calendar to schedule a call:

Schedule a Call

Contact:

Jason Rath

TalentPros.AI

Finding the signal in the noise since 2005

512-993-8228

Jason@TalentPros.AI

Top