Candidate Profile — Candidate #1
Position: Contextual AI — Member of Technical Staff Data Platform (Mountain View, CA)
Candidate Location: San Francisco Bay Area
Experience: 6.2 years experience

Relevant Experience & Education Highlights

Deep expertise in GPU kernel optimization, Kubernetes deployments, vLLM inference servers, and distributed training infrastructure aligns strongly with building scalable data services and ML integrations for Contextual AI's Data Platform.

1. Design and implement scalable services, APIs, and databases to support the processing and ingestion of petabytes of information daily

Delivered high-performance, reliable infrastructure for AI workloads at scale.

  • Architected Kubernetes-based high-throughput inference service for vast datasets during capstone project at a top research university, implementing dynamic batching and scheduling for 85% throughput improvement.
  • Led vLLM-based inference server architecture with custom PagedAttention components using CUDA kernels at a national transfer network organization, achieving 60% reduction in end-to-end P99 latency.
  • Developed suite of internal APIs and tools using Python Flask and React for inference management platform serving over 10,000 DAU with 99.9% uptime at a software development company.
  • Engineered backend components for ML inference platform with Redis caching and semantic similarity matching at a financial services technology firm.

2. Architect and build streaming infrastructure, data orchestration systems, vector databases

Optimized low-latency data processing and memory management for production AI systems.

  • Developed low-latency C++ services for high-throughput industrial IoT data streams at a manufacturing company, achieving 50% improvement in data processing pipelines.
  • Built and stabilized flash-attention CUDA kernels with 12+ hour nvcc compilation pipelines for GPU infrastructure at a biotech AI company.
  • Implemented deterministic container build pipelines and GPU memory management for distributed training systems using Docker and Kubernetes.

3. Ensure seamless integration with machine learning models and pipelines, enabling efficient model deployment and management

Integrated custom kernels and frameworks for optimized LLM training and inference.

  • Designed and optimized high-performance CUDA, ROCm, and Triton kernels using C++, PTX, and GPU assembly for production-scale LLM workloads, delivering 85% performance improvements.
  • Integrated PyTorch, JAX, NCCL, PyTorch DDP, and Ray for distributed GPU training and vLLM-compatible inference engines.
  • Profiled GPU performance with Nsight, nvprof, and CUPTI tools for hardware-aware optimizations and low-precision algorithms.

Requirements & Candidate Alignment

Contextual AI RequirementCandidate Qualification
Education: At least a Bachelor's degree in Computer Science, Software Engineering, or related fieldMaster's degree in Information Systems Management from a top research university; Bachelor of Engineering in Electronics and Instrumentation Engineering from a premier engineering institute
Kubernetes services:Kubernetes expertise architecting high-throughput inference services with dynamic batching, intelligent scheduling, Docker, and deterministic deployments
Distributed queuing systems:Distributed queuing proficiency with NCCL, PyTorch DDP, and Ray for multi-node GPU training infrastructure
Streaming infrastructure:Streaming infrastructure experience developing low-latency C++ services for high-throughput IoT data streams and real-time systems
Proven ability to diagnose distributed vector databases and design systems for low-latency retrieval of image, text, audio, and video vectors:Low-latency retrieval systems designed via vLLM PagedAttention, custom CUDA/Metal kernels, and GPU memory management for multimodal AI workloads
Machine Learning: Familiarity with machine learning concepts and frameworks, including dense information retrieval, document understanding/parsing models, and vision language modelMachine learning frameworks including PyTorch, JAX, Triton Inference Server, vLLM, and custom kernels for LLM inference, training, and document processing optimization
Problem-Solving: Strong problem-solving skills and the ability to work effectively in a fast-paced, collaborative environmentStrong problem-solving demonstrated by 85% throughput gains, 60% latency reductions, and 50% pipeline improvements across GPU systems and ML projects
Experience:6.2 years in GPU systems engineering, ML infrastructure, runtime performance, and kernel optimization

If you would like to discuss this candidate or other critical roles, here is a link to my calendar to schedule a call:

Schedule a Call
← Previous#1#2#3Next →

Contact:

Jason Rath

TalentPros.AI

Finding the signal in the noise since 2005

512-993-8228

Jason@TalentPros.AI

Top