Relevant Experience & Education Highlights
Extensive SRE and cloud architecture background with hands-on expertise in HPC clusters, Kubernetes orchestration, PyTorch, and scalable AI infrastructure positions this candidate as a strong fit for advancing Figure's Helix team training systems.
1. Design, deploy, and maintain training clusters
Proven track record managing large-scale HPC and cloud clusters optimized for high-performance AI workloads.
- Managed HPC clusters and high availability infrastructure for machine learning and data science environments at a leading AI/ML platform company.
- Architected Kubernetes clusters including EKS, GKE, and AKS administration, upgrades, and troubleshooting across AWS, GCP, and Azure.
- Implemented VMware ESXi farms scaling to 200 nodes and 1800+ VMs, alongside big data clusters in multi-region startup environments.
- Oversaw datacenter operations with proactive monitoring using Datadog, Grafana, and Logstash for reliability and uptime.
2. Architect and maintain scalable deep learning frameworks for training on massive robot datasets
Deep expertise in building scalable frameworks supporting PyTorch, CUDA, and Large Language Models for distributed AI training.
- Developed infrastructure for Large Language Models leveraging PyTorch, CUDA, TensorFlow, and TensorRT at enterprise SaaS and consulting firms.
- Executed multi-cloud migrations from AWS to Azure using Terraform and Kubernetes, achieving 30% cost savings and scalable CI/CD pipelines.
- Automated infrastructure provisioning with Terraform, Packer, Jenkins, and GitOps across AWS, GCP, and hybrid cloud setups.
- Maintained production environments with MySQL, PostgreSQL, Redis, and networking components like CDNs, firewalls, and ingress controllers.
3. Implement distributed training and parallelization strategies to reduce model development cycles
Strong foundation in distributed systems, orchestration, and DevOps practices accelerating AI model development.
- Led incident response and root cause analysis for production outages in EKS, GKE, AKS, databases, and networking, ensuring high reliability.
- Deployed job orchestration with Kubernetes, Jenkins, and Chef for automated scaling and high-performance computing workloads.
- Optimized system performance through Linux administration, shell scripting, and tools like Vault for security in SRE roles.
4. Implement tooling for data processing, model experimentation, and continuous integration
Hands-on experience creating developer tools and CI/CD pipelines for data-intensive AI and ML pipelines.
- Built CLI and web tools for automating Dev, staging, and production environments including AWS scaling and monitoring at a cloud content management company.
- Integrated data analytics pipelines with AWS S3, RDS, VPC, and Azure DevOps for machine learning and data science projects.
- Established continuous integration practices using Git, JIRA, Agile methodologies, and Terraform for reliable backend systems.
Requirements & Candidate Alignment
| Figure Requirement | Candidate Qualification |
|---|---|
| Bachelor's or Master's degree in Computer Science, Robotics, Engineering, or a related field | Bachelor's degree in Electronics and Communications Engineering from a technological university |
| Experience with Python and PyTorch | Proficient with Python and PyTorch for machine learning and Large Language Models |
| Experience managing HPC clusters for deep neural network training | Managed HPC clusters with High Performance Computing, CUDA, and high availability for AI/ML workloads |
| Minimum of 4 years of professional, full-time experience building reliable backend systems | 15.6 years building reliable backend systems as Sr SRE/Cloud Architect across multiple roles |
| Strong software engineering fundamentals | Demonstrated through SDLC, GitOps, Docker, shell scripting, and Linux system administration |
| Experience managing cloud infrastructure (AWS, Azure, GCP) | Extensive experience managing AWS, Azure, and GCP including EC2, EKS, AKS, GKE, S3, and VPC |
| Experience with job scheduling / orchestration tools (SLURM, Kubernetes, LSF, etc.) | Expertise with Kubernetes, Jenkins, and orchestration for scalable clusters and CI/CD |
| Experience with configuration management tools (Ansible, Terraform, Puppet, Chef, etc.) | Hands-on with Terraform, Chef, Ansible, and Packer for infrastructure automation and provisioning |
If you would like to discuss this candidate or other critical roles, here is a link to my calendar to schedule a call:
Schedule a CallContact:
Finding the signal in the noise since 2005