Candidate Profile — Candidate #2
Position: Figure — AI Training Infrastructure Engineer - Helix Team (San Jose, CA)
Candidate Location: Milpitas, California
Experience: 19.8 years experience

Relevant Experience & Education Highlights

Demonstrates strong alignment with Figure's Helix team needs through extensive experience building scalable ML infrastructure, data pipelines, cloud clusters, and deep learning tools using PyTorch, Python, Kubernetes, and Terraform.

1. Design, deploy, and maintain training clusters

Managed and scaled cloud-based clusters and infrastructure for high-volume data and ML workloads.

  • Architected cloud-based data backend using Azure AKS, Airflow, and Snowflake at a large educational institution serving hundreds of users.
  • Built and scaled real-time data pipelines at a major enterprise SaaS company to support data stakeholders.
  • Deployed infrastructure with Kubernetes, Docker, and Terraform for reliable backend systems.
  • Leveraged AWS and Azure for cloud infrastructure management.

2. Architect and maintain scalable deep learning frameworks for training on massive robot datasets

Developed frameworks and tools for machine learning models, data processing, and large-scale experimentation.

  • Built frameworks for batch and real-time data pipelines, machine learning, experimentation, and data analytics at a major enterprise SaaS company.
  • Constructed ML infrastructure supporting production-grade, product-powered models with PyTorch and TensorFlow.
  • Applied deep learning techniques including Convolutional Neural Networks, GANs, and Reinforcement Learning in computational projects.
  • Processed massive datasets using Apache Spark, Pandas, and Snowflake for data engineering workflows.

3. Implement distributed training and parallelization strategies to reduce model development cycles

Implemented scalable pipelines and orchestration to accelerate ML development and deployment.

  • Owned and scaled online-to-offline real-time data pipelines for efficient model serving at a major enterprise SaaS company.
  • Utilized Kubernetes and Airflow for job orchestration and distributed processing.
  • Integrated CI/CD practices with Git and Docker to streamline experimentation and deployment cycles.
  • Applied Terraform for configuration management in cloud environments.

4. Implement tooling for data processing, model experimentation, and continuous integration

Created developer tools for data handling, testing, and integration in ML systems.

  • Developed data cleaning, wrangling, and ETL pipelines using Python, SQL, and PostgreSQL.
  • Implemented A/B testing and experimentation frameworks alongside ML tools.
  • Collaborated on data governance and secure data systems across teams.
  • Utilized Scikit-Learn, XGBoost, and PyTorch for model experimentation and evaluation.

Requirements & Candidate Alignment

Figure RequirementCandidate Qualification
Strong software engineering fundamentalsStrong fundamentals across languages including Python, C++, Go, Rust, and Java, with full-stack backend development experience
Bachelor's or Master's degree in Computer Science, Robotics, Engineering, or a related fieldPhD in Biophysics from a top-tier research university; Bachelor’s in Physics with Computer Science minor from a top-tier research university
Experience with Python and PyTorchExtensive experience with Python and PyTorch in building ML infrastructure, data pipelines, and deep learning applications
Experience managing HPC clusters for deep neural network trainingManaged scalable clusters using Azure AKS, Kubernetes, and Airflow for large-scale data and ML workloads
Minimum of 4 years of professional, full-time experience building reliable backend systems19.8 years of professional experience building reliable backend systems, data pipelines, and ML infrastructure
Experience managing cloud infrastructure (AWS, Azure, GCP)Hands-on experience managing AWS, Azure, and related cloud services including Azure Databricks
Experience with job scheduling / orchestration tools (SLURM, Kubernetes, LSF, etc.)Proficient with Kubernetes, Airflow, and batch processing orchestration
Experience with configuration management tools (Ansible, Terraform, Puppet, Chef, etc.)Experienced with Terraform for infrastructure as code and configuration management

If you would like to discuss this candidate or other critical roles, here is a link to my calendar to schedule a call:

Schedule a Call

Contact:

Jason Rath

TalentPros.AI

Finding the signal in the noise since 2005

512-993-8228

Jason@TalentPros.AI

Top