Relevant Experience & Education Highlights
Proven AI infrastructure engineer with extensive experience scaling HPC clusters for massive deep learning models and implementing distributed GPU training systems, directly aligning with Figure's Helix team requirements for training infrastructure.
1. Design, deploy, and maintain training clusters
Excelled in architecting and scaling high-performance computing clusters to support exascale AI training workloads.
- Scaled HPC cluster infrastructure from petaFLOPs for 100M parameter models to several exaFLOPs for 70B-175B parameter foundational models at a leading wafer-scale AI hardware company.
- Designed HPC infrastructure plans including networking layouts, file system configurations, and GPU integrations to meet high-throughput demands in semiconductor image processing at a major manufacturing firm.
- Supported Slurm and Kubernetes-based GPU stacks with HGX-H100 fleets, handling customer onboarding, data migrations, and scale-related networking challenges at a stealth GPU cloud startup.
- Triaged production reliability issues across GPU hardware, scheduling, networking, and storage via trace and log analysis for optimal training and inference performance.
2. Implement distributed training and parallelization strategies
Mastered parallel programming and distributed systems to accelerate large-scale model training cycles.
- Applied parallel algorithms, MPI, OpenMP, and CUDA expertise in high-performance computing environments for deep neural network training.
- Assisted in early prototype cluster design, provisioning, monitoring, deployment, and customer-site support for wafer-scale systems handling massive models.
- Optimized training/inference stacks across GPU fleets using load balancing and RoCEv2 networking to overcome scale challenges.
3. Implement tooling for data processing, model experimentation, and continuous integration
Developed internal tools and solutions engineering practices to enhance developer productivity and reliability in AI workflows.
- Built internal tooling and documentation for platform dependencies, data processing, and customer production issue resolution at a stealth GPU infrastructure startup.
- Shaped product roadmaps through customer learnings, driving GPU-based cloud infrastructure improvements for sustained high-performance training.
- Leveraged Python alongside Linux, simulations, and high-performance computing skills for backend systems reliability.
Requirements & Candidate Alignment
| Figure Requirement | Candidate Qualification |
|---|---|
| Strong software engineering fundamentals | Deep expertise in algorithms, data structures, parallel programming, C/C++, Python, and high-performance computing |
| Bachelor's or Master's degree in Computer Science, Robotics, Engineering, or a related field | PhD in Computer Science from a top-tier research university; Bachelor's in Information Technology from a prominent engineering university |
| Experience with Python and PyTorch | Hands-on experience with Python in AI infrastructure and deep learning systems |
| Experience managing HPC clusters for deep neural network training | Extensive track record scaling HPC/GPU clusters from petaFLOPs to exaFLOPs for foundational models up to 175B parameters |
| Minimum of 4 years of professional, full-time experience building reliable backend systems | 15.7 years building reliable HPC, AI infrastructure, and backend systems across multiple roles |
| Experience with job scheduling / orchestration tools (SLURM, Kubernetes, LSF, etc.) | Proven work with SLURM for GPU cluster management and orchestration |
| Experience managing cloud infrastructure (AWS, Azure, GCP) | Familiarity with for cloud-based HPC and AI workloads |
If you would like to discuss this candidate or other critical roles, here is a link to my calendar to schedule a call:
Schedule a CallContact:
Finding the signal in the noise since 2005