Candidate Profile — Candidate #1
Position: Contextual AI — Member of Technical Staff Data Platform (Mountain View, CA)
Candidate Location: San Francisco Bay Area
Experience: 6.2 years experience
Relevant Experience & Education Highlights
Deep expertise in GPU kernel optimization, Kubernetes deployments, vLLM inference servers, and distributed training infrastructure aligns strongly with building scalable data services and ML integrations for Contextual AI's Data Platform.
1. Design and implement scalable services, APIs, and databases to support the processing and ingestion of petabytes of information daily
Delivered high-performance, reliable infrastructure for AI workloads at scale.
- Architected Kubernetes-based high-throughput inference service for vast datasets during capstone project at a top research university, implementing dynamic batching and scheduling for 85% throughput improvement.
- Led vLLM-based inference server architecture with custom PagedAttention components using CUDA kernels at a national transfer network organization, achieving 60% reduction in end-to-end P99 latency.
- Developed suite of internal APIs and tools using Python Flask and React for inference management platform serving over 10,000 DAU with 99.9% uptime at a software development company.
- Engineered backend components for ML inference platform with Redis caching and semantic similarity matching at a financial services technology firm.
2. Architect and build streaming infrastructure, data orchestration systems, vector databases
Optimized low-latency data processing and memory management for production AI systems.
- Developed low-latency C++ services for high-throughput industrial IoT data streams at a manufacturing company, achieving 50% improvement in data processing pipelines.
- Built and stabilized flash-attention CUDA kernels with 12+ hour nvcc compilation pipelines for GPU infrastructure at a biotech AI company.
- Implemented deterministic container build pipelines and GPU memory management for distributed training systems using Docker and Kubernetes.
3. Ensure seamless integration with machine learning models and pipelines, enabling efficient model deployment and management
Integrated custom kernels and frameworks for optimized LLM training and inference.
- Designed and optimized high-performance CUDA, ROCm, and Triton kernels using C++, PTX, and GPU assembly for production-scale LLM workloads, delivering 85% performance improvements.
- Integrated PyTorch, JAX, NCCL, PyTorch DDP, and Ray for distributed GPU training and vLLM-compatible inference engines.
- Profiled GPU performance with Nsight, nvprof, and CUPTI tools for hardware-aware optimizations and low-precision algorithms.
Requirements & Candidate Alignment
| Contextual AI Requirement | Candidate Qualification |
|---|---|
| Education: At least a Bachelor's degree in Computer Science, Software Engineering, or related field | Master's degree in Information Systems Management from a top research university; Bachelor of Engineering in Electronics and Instrumentation Engineering from a premier engineering institute |
| Kubernetes services: | Kubernetes expertise architecting high-throughput inference services with dynamic batching, intelligent scheduling, Docker, and deterministic deployments |
| Distributed queuing systems: | Distributed queuing proficiency with NCCL, PyTorch DDP, and Ray for multi-node GPU training infrastructure |
| Streaming infrastructure: | Streaming infrastructure experience developing low-latency C++ services for high-throughput IoT data streams and real-time systems |
| Proven ability to diagnose distributed vector databases and design systems for low-latency retrieval of image, text, audio, and video vectors: | Low-latency retrieval systems designed via vLLM PagedAttention, custom CUDA/Metal kernels, and GPU memory management for multimodal AI workloads |
| Machine Learning: Familiarity with machine learning concepts and frameworks, including dense information retrieval, document understanding/parsing models, and vision language model | Machine learning frameworks including PyTorch, JAX, Triton Inference Server, vLLM, and custom kernels for LLM inference, training, and document processing optimization |
| Problem-Solving: Strong problem-solving skills and the ability to work effectively in a fast-paced, collaborative environment | Strong problem-solving demonstrated by 85% throughput gains, 60% latency reductions, and 50% pipeline improvements across GPU systems and ML projects |
| Experience: | 6.2 years in GPU systems engineering, ML infrastructure, runtime performance, and kernel optimization |
If you would like to discuss this candidate or other critical roles, here is a link to my calendar to schedule a call:
Schedule a CallContact:
Finding the signal in the noise since 2005