Relevant Experience & Education Highlights
Extensive compiler engineering background with direct hands-on MLIR and PyTorch expertise positions this candidate to excel in extending Mythic's MLIR dialects and PyTorch interoperability for hardware-aware representations.
1. 3+ years of experience in compiler or high-performance systems development
Led compiler development across quantum, AI, and high-performance computing projects at leading technology firms.
- Architected AI compilers for PyTorch using OpenAI Triton and MLIR at an AI research company.
- Developed quantum compilers with LLVM/MLIR backends and AST to MLIR lowering at a major research lab.
- Retargeted LLVM/Clang for ARM64 supercomputing architectures at a major semiconductor company.
- Maintained and extended GCC, Clang/LLVM for SPARC and Intel at enterprise software and OS companies.
2. Direct, hands-on experience with MLIR, including dialect design, compiler passes, or lowering pipelines
Designed and implemented MLIR-based pipelines for quantum and AI workloads, bridging high-level ops to hardware-specific IRs.
- Built MLIR flows for distributed AI inference/training with PyTorch and OpenMPI at an AI research firm.
- Created quantum compiler toolchains using LLVM/MLIR on RISC-V and FPGA at quantum computing companies.
- Implemented AST to MLIR lowering and OpenQASM 3.0 front-end with C++17 at a premier research institution.
- Integrated LLVM/MLIR/QEMU/TVM for specialized RISC-V64 compilers with modulo scheduling at a circuit design firm.
3. Proficiency in modern C++ (C++14/17/20) and Python
Applied advanced C++ standards and Python in production compiler infrastructures for AI and quantum systems.
- Developed OpenQASM 3.0 compiler front-end in C++17 with Flex and Bison for quantum ISAs.
- Engineered PyTorch C++ compilers with MLIR and Triton kernels using modern C++ at FPGA and AI firms.
- Contributed to Apache C++ Standard Library and BOOST while maintaining GCC/Clang toolchains.
- Utilized Python in PyTorch-based AI compilers alongside CUDA, CUTLASS, and NVSHMEM for distributed training.
4. Familiarity with PyTorch compiler technologies, especially Torch-MLIR and integration paths with PyTorch 2.0
Integrated PyTorch with MLIR and Triton for AI model compilation and inference acceleration.
- Developed AI compiler for PyTorch leveraging OpenAI Triton Compiler and MLIR at an AI research company.
- Implemented PyTorch C++ with NVIDIA Triton Inference Server and LLaMA-CPP for distributed inference.
- Built backends for Apple METAL and CUDA in Triton compilers supporting PyTorch workflows.
Requirements & Candidate Alignment
| Mythic Requirement | Candidate Qualification |
|---|---|
| 3+ Years of experience in compiler or high-performance systems development | 23.6 years across AI, quantum, and supercomputing compiler roles at technology leaders |
| Proficiency in modern C++ (C++14/17/20) and Python | Advanced expertise in C++17 for quantum compilers and Python for PyTorch/MLIR AI pipelines |
| Direct, hands-on experience with MLIR, including dialect design, compiler passes, or lowering pipelines | Deep hands-on work building MLIR flows for PyTorch AI, quantum RISC-V, and AST lowering |
| Strong understanding of compiler IRs and transformations, with the ability to reason about lowering from high-level ops to hardware-aware representations | Proven mastery via LLVM/MLIR retargeting for ARM64, SPARC, quantum ISAs, and TVM Relay models |
| Experience architecting complete MLIR flows: from frontend dialects down to hardware-aware dialects, including conversion to and from existing IRs | Architected full flows with LLVM/MLIR/QEMU/TVM on RISC-V64 and quantum FPGA backends |
| Familiarity with PyTorch compiler technologies, especially Torch-MLIR and integration paths with PyTorch 2.0 (TorchDynamo, TorchInductor) | Strong integration of PyTorch with MLIR/Triton compilers for CUDA/METAL inference and training |
| Background in heterogeneous or specialized accelerators (e.g., analog compute, NPUs, GPUs, DSPs) | Extensive experience with FPGA (AMD Xilinx, Intel Altera), ARM64 ThunderX, quantum CPUs, and CUDA GPUs |
If you would like to discuss this candidate or other critical roles, here is a link to my calendar to schedule a call:
Schedule a CallContact:
Finding the signal in the noise since 2005