Candidate Profile — Candidate #3
Position: Figure — AI Training Infrastructure Engineer - Helix Team (San Jose, CA)
Candidate Location: San Francisco Bay Area
Experience: 15.6 years experience

Relevant Experience & Education Highlights

Extensive SRE and cloud architecture background with hands-on expertise in HPC clusters, Kubernetes orchestration, PyTorch, and scalable AI infrastructure positions this candidate as a strong fit for advancing Figure's Helix team training systems.

1. Design, deploy, and maintain training clusters

Proven track record managing large-scale HPC and cloud clusters optimized for high-performance AI workloads.

  • Managed HPC clusters and high availability infrastructure for machine learning and data science environments at a leading AI/ML platform company.
  • Architected Kubernetes clusters including EKS, GKE, and AKS administration, upgrades, and troubleshooting across AWS, GCP, and Azure.
  • Implemented VMware ESXi farms scaling to 200 nodes and 1800+ VMs, alongside big data clusters in multi-region startup environments.
  • Oversaw datacenter operations with proactive monitoring using Datadog, Grafana, and Logstash for reliability and uptime.

2. Architect and maintain scalable deep learning frameworks for training on massive robot datasets

Deep expertise in building scalable frameworks supporting PyTorch, CUDA, and Large Language Models for distributed AI training.

  • Developed infrastructure for Large Language Models leveraging PyTorch, CUDA, TensorFlow, and TensorRT at enterprise SaaS and consulting firms.
  • Executed multi-cloud migrations from AWS to Azure using Terraform and Kubernetes, achieving 30% cost savings and scalable CI/CD pipelines.
  • Automated infrastructure provisioning with Terraform, Packer, Jenkins, and GitOps across AWS, GCP, and hybrid cloud setups.
  • Maintained production environments with MySQL, PostgreSQL, Redis, and networking components like CDNs, firewalls, and ingress controllers.

3. Implement distributed training and parallelization strategies to reduce model development cycles

Strong foundation in distributed systems, orchestration, and DevOps practices accelerating AI model development.

  • Led incident response and root cause analysis for production outages in EKS, GKE, AKS, databases, and networking, ensuring high reliability.
  • Deployed job orchestration with Kubernetes, Jenkins, and Chef for automated scaling and high-performance computing workloads.
  • Optimized system performance through Linux administration, shell scripting, and tools like Vault for security in SRE roles.

4. Implement tooling for data processing, model experimentation, and continuous integration

Hands-on experience creating developer tools and CI/CD pipelines for data-intensive AI and ML pipelines.

  • Built CLI and web tools for automating Dev, staging, and production environments including AWS scaling and monitoring at a cloud content management company.
  • Integrated data analytics pipelines with AWS S3, RDS, VPC, and Azure DevOps for machine learning and data science projects.
  • Established continuous integration practices using Git, JIRA, Agile methodologies, and Terraform for reliable backend systems.

Requirements & Candidate Alignment

Figure RequirementCandidate Qualification
Bachelor's or Master's degree in Computer Science, Robotics, Engineering, or a related fieldBachelor's degree in Electronics and Communications Engineering from a technological university
Experience with Python and PyTorchProficient with Python and PyTorch for machine learning and Large Language Models
Experience managing HPC clusters for deep neural network trainingManaged HPC clusters with High Performance Computing, CUDA, and high availability for AI/ML workloads
Minimum of 4 years of professional, full-time experience building reliable backend systems15.6 years building reliable backend systems as Sr SRE/Cloud Architect across multiple roles
Strong software engineering fundamentalsDemonstrated through SDLC, GitOps, Docker, shell scripting, and Linux system administration
Experience managing cloud infrastructure (AWS, Azure, GCP)Extensive experience managing AWS, Azure, and GCP including EC2, EKS, AKS, GKE, S3, and VPC
Experience with job scheduling / orchestration tools (SLURM, Kubernetes, LSF, etc.)Expertise with Kubernetes, Jenkins, and orchestration for scalable clusters and CI/CD
Experience with configuration management tools (Ansible, Terraform, Puppet, Chef, etc.)Hands-on with Terraform, Chef, Ansible, and Packer for infrastructure automation and provisioning

If you would like to discuss this candidate or other critical roles, here is a link to my calendar to schedule a call:

Schedule a Call

Contact:

Jason Rath

TalentPros.AI

Finding the signal in the noise since 2005

512-993-8228

Jason@TalentPros.AI

Top