Candidate Profile — Candidate #4
Position: Metropolis — .io - Senior Machine Learning Engineer, Computer Vision (Los Angeles, CA or Seattle, WA)
Candidate Location: Seattle Metro
Experience: 8.7 years experience

Relevant Experience & Education Highlights

Brings 8.7 years of hands-on expertise in computer vision, multimodal learning, video understanding, and production model deployment, strongly aligning with the demands of designing and deploying advanced CV models at Metropolis.

1. Design, develop, and deploy advanced computer vision models for real-world applications, including object detection, tracking, OCR, image search, and scene understanding

Tackled core computer vision challenges in production environments, delivering measurable improvements in accuracy and efficiency.

  • Fine-tuned 7B parameter vision-language models on proprietary video datasets at a major streaming service division of a large technology company, boosting precision by 70% for movie scene content descriptor prediction.
  • Developed object detection models using YOLO algorithm with PyTorch at a startup, enabling accurate identification of stationery items.
  • Built deep learning models with TensorFlow and Keras for artist classification and style prediction from image patches at an AI company serving the art market.
  • Achieved 7% improvement over state-of-the-art on internal long-form video datasets and 2% on ActivityNet through scalable models handling up to 3-minute videos at a major technology company.

2. Explore and integrate multi-modal approaches, leveraging visual, textual, and other data modalities for robust solutions

Integrated vision, language, and multimodal data to solve complex real-world problems like content moderation and product matching.

  • Productionized multi-level profane content detection combining LLMs and model ensembles for textual and visual modalities at a major streaming service, reducing operator review time by 75% and achieving 90% accuracy across multiple territories.
  • Conducted representation learning for joint embedding of text sequences, images, and videos in story-illustration tasks at a research lab within a top-tier university.
  • Performed unsupervised product normalization using image and text-based models, increasing coverage by 4% across large catalogs at an e-commerce product platform.

3. Build and optimize deep learning models, ensuring high accuracy, performance, and scalability for deployment in production environments

Optimized large-scale models for domain-specific tasks, enabling production deployment with significant performance gains.

  • Instruction-fine-tuned vision-language models for maturity rating and content descriptor prediction on Prime Video datasets at a major technology company.
  • Implemented weakly supervised instance segmentation for video understanding, with work published at AMLC 2024.
  • Led full SDLC for deep learning projects, including model training and deployment, while training teams in recommendation systems at a startup.

Requirements & Candidate Alignment

Metropolis RequirementCandidate Qualification
PhD in Computer Science, Engineering, or a related field, or equivalent work experienceMS in Computer Science from a top-tier research university, complemented by 8.7 years of equivalent hands-on expertise
5+ Years of hands-on experience in machine learning and computer vision, with a strong track record of deploying models into production8.7 years including production deployment of vision-language models, content moderation systems, and video understanding pipelines at a major technology company
Proficiency in Python and ML frameworks (PyTorch/TensorFlow/ONNX/TensorRT)Proficiency in Python, PyTorch, TensorFlow, and Keras demonstrated through fine-tuning large vision-language models and YOLO-based detection
Strong experience with model optimization (e.g., quantization, pruning) and deployment on various platforms (cloud, edge, or mobile)Optimized and deployed scalable deep learning models for long-form video and scene analysis, achieving state-of-the-art improvements and zero human intervention in production
Familiarity with cloud platforms (AWS, GCP, or Azure), containerization (Docker), and orchestration (ECS, Kubernetes)Productionized ML models at scale in high-throughput video processing environments at a major technology company
Proven experience in building and maintaining data pipelines (e.g., Airflow)Developed end-to-end data preprocessing pipelines for image datasets and multimodal training at multiple organizations
Strong understanding of the agile development process and CI/CD pipelines and tools (e.g., Github Actions, Jenkins)Managed full software development lifecycle, including interviews and team training, in fast-paced startup environments
Excellent communication skills, capable of presenting complex technical information clearlyLed cross-functional teams and trained interns and coworkers on ML systems and recommendation technologies

If you would like to discuss this candidate or other critical roles, here is a link to my calendar to schedule a call:

Schedule a Call

Contact:

Jason Rath

TalentPros.AI

Finding the signal in the noise since 2005

512-993-8228

Jason@TalentPros.AI

Top