Relevant Experience & Education Highlights
Extensive experience building scalable AI training infrastructure, including distributed LLM training with PyTorch, GPU clusters managed by Kubernetes and Slurm, and MLOps tools, positions this candidate strongly for advancing Figure's Helix Team training systems.
1. Design, deploy, and maintain training clusters
Managed high-performance computing environments and GPU clusters optimized for deep learning workloads.
- Built next-gen AI cloud infrastructure featuring GPU clusters orchestrated with Kubernetes at a stealth AI startup.
- Developed ML serving platforms incorporating Ray, Kubernetes, and vector databases for large-scale model deployment.
- Led construction of unified control planes for machine learning platforms supporting multi-GPU distributed training at a major ride-sharing company.
2. Architect and maintain scalable deep learning frameworks for training on massive robot datasets
Architected frameworks and services for training and inference of large language models and transformer architectures.
- Implemented LLM distributed training, acceleration, and performance modeling using PyTorch, JAX, and Triton at a stealth AI startup.
- Designed high-throughput ML inference services with Ray and Kubernetes, including scalable Retrieval-Augmented Generation applications with LLMs and vector databases at an AI/ML company.
- Developed custom Triton kernels to accelerate LLM training and inference, alongside quantization and tensor parallelism optimizations.
3. Implement distributed training and parallelization strategies to reduce model development cycles
Explored and applied advanced parallelism techniques to enhance training efficiency across distributed systems.
- Explored model parallelism strategies to accelerate LLM training and inference using PyTorch and JAX.
- Optimized model inference through tensor parallelism, continuous batching, and quantization in production environments.
- Built MLOps orchestration frameworks with Spark data processing, multi-GPU training, and job failure recovery at a major ride-sharing company.
4. Implement tooling for data processing, model experimentation, and continuous integration
Created developer tools, pipelines, and observability for streamlined AI model development workflows.
- Developed CLI tools, workflow schedulers, and pipeline observability for model development at a major ride-sharing company.
- Built A/B testing platforms and data analytics pipelines using Spark, Kafka, Cassandra, and Redis at a global information services company.
- Researched capacity planning tools and performance modeling for LLMs, enhancing experimentation cycles.
Requirements & Candidate Alignment
| Figure Requirement | Candidate Qualification |
|---|---|
| Strong software engineering fundamentals | 13.7 years building reliable backend systems, including distributed ML platforms and high-throughput services |
| Bachelor's or Master's degree in Computer Science, Robotics, Engineering, or a related field | Master's degree in Computer & Electrical Engineering from a top research university; Bachelor's degree in Microelectronics from a leading engineering university in China |
| Experience with Python and PyTorch | Extensive hands-on experience with Python and PyTorch for LLM distributed training, performance modeling, and deep learning frameworks |
| Experience managing HPC clusters for deep neural network training | Proven track record managing GPU clusters with Kubernetes and Slurm for LLM training and inference acceleration |
| Minimum of 4 years of professional, full-time experience building reliable backend systems | 13.7 years of professional experience across AI platforms, MLOps, and distributed systems engineering roles |
| Experience managing cloud infrastructure (AWS, Azure, GCP) | Direct experience with AWS and GCP in building scalable AI infrastructure and services |
| Experience with job scheduling / orchestration tools (SLURM, Kubernetes, LSF, etc.) | Strong expertise with Kubernetes, Slurm, and Ray for orchestrating distributed training and inference workloads |
If you would like to discuss this candidate or other critical roles, here is a link to my calendar to schedule a call:
Schedule a CallContact:
Finding the signal in the noise since 2005