Relevant Experience & Education Highlights
Brings deep hands-on expertise in MLIR dialect design, PyTorch compiler integration via Torch-MLIR, and building end-to-end compilation stacks for heterogeneous ML accelerators, directly aligning with extending Mythic's MLIR ecosystem and PyTorch infrastructure.
1. Direct, hands-on experience with MLIR, including dialect design, compiler passes, or lowering pipelines
Led MLIR-based compiler development to bridge high-level frameworks to hardware targets.
- Co-created and maintained Tensor Compute Primitives (mlir-tcp), an open-source mid-level MLIR dialect for ML programs, at a leading autonomous vehicle technology company.
- Co-maintained Torch-MLIR, providing first-class compiler support from PyTorch to MLIR ecosystems.
- Tech-led ML compiler frontend connecting state-of-the-art PyTorch models to MLIR for optimization toward heterogeneous compute platforms.
- Built MLIR-enabled compilation stack simplifying ML model deployments onto commodity and in-house ML accelerators.
2. Experience architecting complete MLIR flows: from frontend dialects down to hardware-aware dialects
Architected full-stack compiler infrastructures optimizing PyTorch workloads for production inference.
- Created GraFX, a source-to-source compiler framework accelerating PyTorch model development, optimization, and deployment lifecycles.
- Developed lowering pipelines from high-level PyTorch operations through MLIR middle-end to latency-critical accelerator targets.
- Worked at intersection of compilers and deep learning algorithms for accelerated ML deployments on autonomous systems.
3. Familiarity with PyTorch compiler technologies, especially Torch-MLIR and integration paths with PyTorch 2.0
Delivered seamless PyTorch-to-MLIR connectivity enabling efficient hardware targeting.
- Integrated PyTorch ecosystem with MLIR compiler passes for optimized inference on diverse accelerators.
- Led frontend development ensuring compatibility with advanced PyTorch models and compiler technologies.
- Developed advanced quantization algorithms like TQT for hardware-friendly fixed-point inference at a programmable logic company specializing in FPGAs and AI Engines.
4. Background in heterogeneous or specialized accelerators
Optimized ML workloads for varied compute platforms including in-house and commodity accelerators.
- Targeted compilation stacks for autonomous vehicle inference on heterogeneous platforms.
- Explored novel architectures and algorithms for efficient ML acceleration on FPGAs and AI Engines in cloud and edge use cases.
- Advanced algorithms for ML acceleration achieving high-accuracy quantized representations for constrained hardware.
Requirements & Candidate Alignment
| Mythic Requirement | Candidate Qualification |
|---|---|
| 3+ Years of experience in compiler or high-performance systems development. | 15.9 years across staff ML compiler engineer and machine learning engineer roles focused on compilers and accelerators. |
| Proficiency in modern C++ (C++14/17/20) and Python. | Proficiency in Python with skills in Jupyter Notebook; hands-on modern C++ aligning with high-performance compiler development. |
| Direct, hands-on experience with MLIR, including dialect design, compiler passes, or lowering pipelines. | Led MLIR dialect design including co-creation of mlir-tcp and maintenance of Torch-MLIR; built full MLIR compilation stacks. |
| Strong understanding of compiler IRs and transformations, with the ability to reason about lowering from high-level ops to hardware-aware representations. | Architected PyTorch-to-MLIR lowering pipelines and source-to-source compilers optimizing to heterogeneous hardware IRs. |
| Experience architecting complete MLIR flows: from frontend dialects down to hardware-aware dialects, including conversion to and from existing IRs. | Built end-to-end MLIR flows from PyTorch frontend through middle-end optimizations to accelerator backends. |
| Familiarity with PyTorch compiler technologies, especially Torch-MLIR and integration paths with PyTorch 2.0 (TorchDynamo, TorchInductor). | Co-maintained Torch-MLIR for PyTorch ecosystem integration; tech-led PyTorch frontend connectivity to MLIR. |
| Knowledge of dataflow architectures, scheduling, and memory orchestration. | Optimized deep learning workloads for latency-critical inference involving scheduling and memory management on accelerators. |
| Background in heterogeneous or specialized accelerators (e.g., analog compute, NPUs, GPUs, DSPs). | Developed for FPGAs, AI Engines, commodity ML accelerators, and in-house heterogeneous platforms. |
If you would like to discuss this candidate or other critical roles, here is a link to my calendar to schedule a call:
Schedule a CallContact:
Finding the signal in the noise since 2005