Relevant Experience & Education Highlights
Deep expertise in GPU kernel development, quantization for LLMs, large-scale TensorFlow infrastructures, and performance optimizations aligns strongly with building scalable PyTorch training stacks and GPU-accelerated pipelines for Adobe's Firefly generative AI foundation models.
1. You'll build and optimize infrastructures that power large foundation model training on thousands of GPUs
Extensive experience scaling deep learning infrastructures for massive models and GPU-heavy workloads.
- Developed infrastructures and performance optimizations including quantization and GPU kernels for serving large language models at a major technology company with cloud services.
- Built end-to-end large-scale deep learning frameworks for image classification, object detection, and segmentation using distributed TensorFlow and Apache Spark at a prominent enterprise technology firm.
- Accelerated distributed neural network model training by 5x in Apache SystemML deep learning library using PySpark, TensorFlow, and GPUs during internship at the same firm.
- Enabled Spark with HDFS to support multidimensional array data formats, optimizing storage and query performance by 10x at a federal space agency research center.
2. You will profile GPU utilization, trace inference and training runs and help craft strategies for optimizing our ML model latency
Proven proficiency in profiling, benchmarking, and end-to-end performance tuning for inference and training.
- Improved inference speed of natural language processing services by profiling model operations, benchmarking, and optimizing data pipelines at a prominent enterprise technology firm.
- Conducted R&D on optimizing transformer-based models for natural language understanding and generation, including GPU kernel development in PTX and quantization at a major technology company.
- Developed forward and back propagation functions with test cases for deep learning library, enhancing prediction tasks like breast cancer proliferation scoring using GPUs.
3. We'll work together to architect and optimize end-to-end ML pipelines, ensuring they're scalable, efficient, and robust
Strong track record architecting scalable ML pipelines with framework contributions and cloud deployment.
- Contributed over 59K lines of code as TensorFlow committer, fixing issues and adding features to support distributed training and inference at a prominent enterprise technology firm.
- Wrapped deep learning models as cloud services, integrating with MLOps practices at the same firm.
- Progressed to Principal Software Engineer role, focusing on LLM serving infrastructures with custom optimizations at a major technology company.
Requirements & Candidate Alignment
| Adobe Requirement | Candidate Qualification |
|---|---|
| Education: Graduate, PhD, or postgraduate degree in Computer Science, Computer Engineering, or a related field—or equivalent experience. | PhD in Computer Science from a top-tier research university. |
| 2+ Years ML Engineering experience, specializing in generative AI like LLMs. | 9.5 years ML engineering experience, specializing in LLMs through infrastructures for serving and optimizing transformer-based models. |
| Strong Python and deep learning engineering skills, paired with experience in training and inferencing with PyTorch or TensorFlow, will be essential. | Deep learning engineering expertise with TensorFlow, including distributed training/inference, TensorFlow committer contributions exceeding 59K lines,-aligned LLM optimizations. |
| Familiarity with distillation, transformers, and diffusion models. Experience with generative image and video is a plus. | Expertise in transformers and distillation techniques like quantization, with R&D on transformer-based LLMs for generation and large-scale image classification/detection pipelines. |
| Knowledge of deployment technologies such as Docker, ML Ops, and ML services is valuable | Deployment technologies experience with MLOps, and ML services, demonstrated by wrapping deep learning models as scalable cloud services. |
| Experience with cloud platforms like Azure and AWS is a plus. | Cloud platforms experience including AWS, through LLM product infrastructures and cloud service deployments at major technology providers. |
| Excellent problem-solving abilities and your capacity to analyze complex issues and drive solutions with a data-driven approach. | Proven problem-solving with data-driven optimizations, including 5x training speedups, 10x storage/query gains, and inference profiling for production services. |
| Strong verbal and written communication skills and success in cross-functional team environments | Cross-functional success with communication excellence, evidenced by Best Demo Paper Award at a major conference and presentations like Spark+AI Summit. |
If you would like to discuss this candidate or other critical roles, here is a link to my calendar to schedule a call:
Schedule a CallContact:
Finding the signal in the noise since 2005