Relevant Experience & Education Highlights
Offers 7.3 years of software engineering experience building scalable ML infrastructure, distributed systems, and Kubernetes-orchestrated clusters, aligning strongly with Figure's needs for AI Training Infrastructure Engineer on the Helix Team.
1. Design, deploy, and maintain Figure's training clusters
Built and optimized infrastructure for high-performance ML and backend operations across demanding environments.
- Developed scalable infrastructure ensuring efficient service operations at a major cryptocurrency trading platform.
- Designed and implemented robust ML model infrastructure MVP at the same platform.
- Orchestrated Docker containers and scaled via Kubernetes for an online learning marketplace on Google Cloud Platform.
- Automated deployments with CI/CD pipelines using Jenkins and Spinnaker on Google Cloud Platform.
2. Architect and maintain scalable deep learning frameworks for training on massive robot datasets
Engineered systems for large-scale data processing and ML workflows supporting complex model development.
- Developed core dashboard systems using big data pipelines and DAGs at a leading financial services company.
- Built internal NLP system converting requirements into MIS pipelines utilized across multiple teams.
- Contributed to adaptive learning platforms and personalized recommendation engines linking student data to hiring needs.
- Led machine learning projects as team incharge, organizing technical events and implementations.
3. Implement distributed training and parallelization strategies to reduce model development cycles
Optimized distributed systems and databases for performance, resilience, and scalability in production.
- Redesigned core backend flows for offline compatibility minimizing network disruptions at a rapid grocery delivery service.
- Improved database scalability through partitioning and indexing, reducing CPU utilization from 60% to 25%.
- Added health checks capturing celery heartbeats and configured Kubernetes liveness/readiness probes to reduce downtimes.
- Leveraged distributed tools including Apache Spark, Apache Flink, Apache Kafka, and PySpark for scalable processing.
4. Implement tooling for data processing, model experimentation, and continuous integration
Delivered DevOps and MLOps tooling for reliable operations, monitoring, and developer productivity.
- Designed initial payouts system for internal partners at a rapid grocery delivery service.
- Utilized Ansible for configuration management and Jenkins for orchestration in infrastructure workflows.
- Developed full-stack systems with Spring Boot, PostgreSQL, and system monitoring practices.
- Created content and instructed on Python, supporting ML and software development teams.
Requirements & Candidate Alignment
| Figure Requirement | Candidate Qualification |
|---|---|
| Education: Bachelor's or Master's degree in Computer Science, Robotics, Engineering, or a related field | Master's in Computer Software Engineering from a top-tier research university; Bachelor's in Computer Science from a leading engineering institute |
| Experience: Minimum of 4 years of professional, full-time experience building reliable backend systems | 7.3 years of professional experience developing scalable backend and ML infrastructure systems |
| Technical Skills: Experience with Python and PyTorch | Python expertise through development of NLP systems, instruction, and ML workflows; PyTorch experience aligns with MLOps and large language models background |
| Infrastructure: Experience managing HPC clusters for deep neural network training | Kubernetes and Docker orchestration for cluster scaling and ML model infrastructure at cloud scale |
| Bonus: Experience managing cloud infrastructure (AWS, Azure, GCP) | Google Cloud Platform hands-on with CI/CD, container orchestration, GCP, and automated deployments |
| Bonus: Experience with job scheduling / orchestration tools (SLURM, Kubernetes, LSF, etc.) | Kubernetes for health checks, scaling, liveness/readiness probes, and Docker orchestration |
| Bonus: Experience with configuration management tools (Ansible, Terraform, Puppet, Chef, etc.) | Ansible for configuration management in DevOps and infrastructure workflows |
| Fundamentals: Strong software engineering fundamentals | Distributed systems, MLOps, and DevOps proficiency with tools like Apache Spark, Kafka, Flink, Jenkins |
If you would like to discuss this candidate or other critical roles, here is a link to my calendar to schedule a call:
Schedule a CallContact:
Finding the signal in the noise since 2005