Candidate Profile — Candidate #2
Position: Mythic — Compiler Engineer – MLIR / PyTorch Infrastructure (Remote)
Candidate Location: Not specified
Experience: 15.9 years experience

Relevant Experience & Education Highlights

Brings deep hands-on expertise in MLIR dialect design, PyTorch compiler integration via Torch-MLIR, and building end-to-end compilation stacks for heterogeneous ML accelerators, directly aligning with extending Mythic's MLIR ecosystem and PyTorch infrastructure.

1. Direct, hands-on experience with MLIR, including dialect design, compiler passes, or lowering pipelines

Led MLIR-based compiler development to bridge high-level frameworks to hardware targets.

  • Co-created and maintained Tensor Compute Primitives (mlir-tcp), an open-source mid-level MLIR dialect for ML programs, at a leading autonomous vehicle technology company.
  • Co-maintained Torch-MLIR, providing first-class compiler support from PyTorch to MLIR ecosystems.
  • Tech-led ML compiler frontend connecting state-of-the-art PyTorch models to MLIR for optimization toward heterogeneous compute platforms.
  • Built MLIR-enabled compilation stack simplifying ML model deployments onto commodity and in-house ML accelerators.

2. Experience architecting complete MLIR flows: from frontend dialects down to hardware-aware dialects

Architected full-stack compiler infrastructures optimizing PyTorch workloads for production inference.

  • Created GraFX, a source-to-source compiler framework accelerating PyTorch model development, optimization, and deployment lifecycles.
  • Developed lowering pipelines from high-level PyTorch operations through MLIR middle-end to latency-critical accelerator targets.
  • Worked at intersection of compilers and deep learning algorithms for accelerated ML deployments on autonomous systems.

3. Familiarity with PyTorch compiler technologies, especially Torch-MLIR and integration paths with PyTorch 2.0

Delivered seamless PyTorch-to-MLIR connectivity enabling efficient hardware targeting.

  • Integrated PyTorch ecosystem with MLIR compiler passes for optimized inference on diverse accelerators.
  • Led frontend development ensuring compatibility with advanced PyTorch models and compiler technologies.
  • Developed advanced quantization algorithms like TQT for hardware-friendly fixed-point inference at a programmable logic company specializing in FPGAs and AI Engines.

4. Background in heterogeneous or specialized accelerators

Optimized ML workloads for varied compute platforms including in-house and commodity accelerators.

  • Targeted compilation stacks for autonomous vehicle inference on heterogeneous platforms.
  • Explored novel architectures and algorithms for efficient ML acceleration on FPGAs and AI Engines in cloud and edge use cases.
  • Advanced algorithms for ML acceleration achieving high-accuracy quantized representations for constrained hardware.

Requirements & Candidate Alignment

Mythic RequirementCandidate Qualification
3+ Years of experience in compiler or high-performance systems development.15.9 years across staff ML compiler engineer and machine learning engineer roles focused on compilers and accelerators.
Proficiency in modern C++ (C++14/17/20) and Python.Proficiency in Python with skills in Jupyter Notebook; hands-on modern C++ aligning with high-performance compiler development.
Direct, hands-on experience with MLIR, including dialect design, compiler passes, or lowering pipelines.Led MLIR dialect design including co-creation of mlir-tcp and maintenance of Torch-MLIR; built full MLIR compilation stacks.
Strong understanding of compiler IRs and transformations, with the ability to reason about lowering from high-level ops to hardware-aware representations.Architected PyTorch-to-MLIR lowering pipelines and source-to-source compilers optimizing to heterogeneous hardware IRs.
Experience architecting complete MLIR flows: from frontend dialects down to hardware-aware dialects, including conversion to and from existing IRs.Built end-to-end MLIR flows from PyTorch frontend through middle-end optimizations to accelerator backends.
Familiarity with PyTorch compiler technologies, especially Torch-MLIR and integration paths with PyTorch 2.0 (TorchDynamo, TorchInductor).Co-maintained Torch-MLIR for PyTorch ecosystem integration; tech-led PyTorch frontend connectivity to MLIR.
Knowledge of dataflow architectures, scheduling, and memory orchestration.Optimized deep learning workloads for latency-critical inference involving scheduling and memory management on accelerators.
Background in heterogeneous or specialized accelerators (e.g., analog compute, NPUs, GPUs, DSPs).Developed for FPGAs, AI Engines, commodity ML accelerators, and in-house heterogeneous platforms.

If you would like to discuss this candidate or other critical roles, here is a link to my calendar to schedule a call:

Schedule a Call

Contact:

Jason Rath

TalentPros.AI

Finding the signal in the noise since 2005

512-993-8228

Jason@TalentPros.AI

Top