Candidate Profile — Candidate #5
Position: Figure — AI Training Infrastructure Engineer - Helix Team (San Jose, CA)
Candidate Location: San Francisco Bay Area
Experience: 7.3 years experience

Relevant Experience & Education Highlights

Offers 7.3 years of software engineering experience building scalable ML infrastructure, distributed systems, and Kubernetes-orchestrated clusters, aligning strongly with Figure's needs for AI Training Infrastructure Engineer on the Helix Team.

1. Design, deploy, and maintain Figure's training clusters

Built and optimized infrastructure for high-performance ML and backend operations across demanding environments.

  • Developed scalable infrastructure ensuring efficient service operations at a major cryptocurrency trading platform.
  • Designed and implemented robust ML model infrastructure MVP at the same platform.
  • Orchestrated Docker containers and scaled via Kubernetes for an online learning marketplace on Google Cloud Platform.
  • Automated deployments with CI/CD pipelines using Jenkins and Spinnaker on Google Cloud Platform.

2. Architect and maintain scalable deep learning frameworks for training on massive robot datasets

Engineered systems for large-scale data processing and ML workflows supporting complex model development.

  • Developed core dashboard systems using big data pipelines and DAGs at a leading financial services company.
  • Built internal NLP system converting requirements into MIS pipelines utilized across multiple teams.
  • Contributed to adaptive learning platforms and personalized recommendation engines linking student data to hiring needs.
  • Led machine learning projects as team incharge, organizing technical events and implementations.

3. Implement distributed training and parallelization strategies to reduce model development cycles

Optimized distributed systems and databases for performance, resilience, and scalability in production.

  • Redesigned core backend flows for offline compatibility minimizing network disruptions at a rapid grocery delivery service.
  • Improved database scalability through partitioning and indexing, reducing CPU utilization from 60% to 25%.
  • Added health checks capturing celery heartbeats and configured Kubernetes liveness/readiness probes to reduce downtimes.
  • Leveraged distributed tools including Apache Spark, Apache Flink, Apache Kafka, and PySpark for scalable processing.

4. Implement tooling for data processing, model experimentation, and continuous integration

Delivered DevOps and MLOps tooling for reliable operations, monitoring, and developer productivity.

  • Designed initial payouts system for internal partners at a rapid grocery delivery service.
  • Utilized Ansible for configuration management and Jenkins for orchestration in infrastructure workflows.
  • Developed full-stack systems with Spring Boot, PostgreSQL, and system monitoring practices.
  • Created content and instructed on Python, supporting ML and software development teams.

Requirements & Candidate Alignment

Figure RequirementCandidate Qualification
Education: Bachelor's or Master's degree in Computer Science, Robotics, Engineering, or a related fieldMaster's in Computer Software Engineering from a top-tier research university; Bachelor's in Computer Science from a leading engineering institute
Experience: Minimum of 4 years of professional, full-time experience building reliable backend systems7.3 years of professional experience developing scalable backend and ML infrastructure systems
Technical Skills: Experience with Python and PyTorchPython expertise through development of NLP systems, instruction, and ML workflows; PyTorch experience aligns with MLOps and large language models background
Infrastructure: Experience managing HPC clusters for deep neural network trainingKubernetes and Docker orchestration for cluster scaling and ML model infrastructure at cloud scale
Bonus: Experience managing cloud infrastructure (AWS, Azure, GCP)Google Cloud Platform hands-on with CI/CD, container orchestration, GCP, and automated deployments
Bonus: Experience with job scheduling / orchestration tools (SLURM, Kubernetes, LSF, etc.)Kubernetes for health checks, scaling, liveness/readiness probes, and Docker orchestration
Bonus: Experience with configuration management tools (Ansible, Terraform, Puppet, Chef, etc.)Ansible for configuration management in DevOps and infrastructure workflows
Fundamentals: Strong software engineering fundamentalsDistributed systems, MLOps, and DevOps proficiency with tools like Apache Spark, Kafka, Flink, Jenkins

If you would like to discuss this candidate or other critical roles, here is a link to my calendar to schedule a call:

Schedule a Call

Contact:

Jason Rath

TalentPros.AI

Finding the signal in the noise since 2005

512-993-8228

Jason@TalentPros.AI

Top