Skip to content
View joepothiboot's full-sized avatar

Block or report joepothiboot

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
joepothiboot/README.md

Joe Pothiboot 👋

I build compilers and kernels for accelerators: MLIR and LLVM pipelines, tiling for CPUs, GPUs and scratchpad DSP/NPU machines, and kernels that are checked bit for bit against a reference.

🎯 Current focus

  • MLIR/LLVM compilers: custom dialects, Transform-dialect schedules, lowering to machine code
  • Kernels for CPU SIMD, GPU and scratchpad/DMA accelerators, measured against real ceilings and vendor libraries
  • Correctness first: every optimized kernel must match an unoptimized reference exactly

🛠️ Projects

An out-of-tree MLIR compiler for a small tensor DSL. A dsp dialect is lowered through Linalg to LLVM, tiled and vectorized by Transform-dialect schedules generated from a hardware target model, for ARM NEON, x86 AVX2 and Hexagon HVX. Scheduled code is bit-exact with unscheduled code and a scalar C++ reference. A Mojo kernel library implements the same ops on CPU and GPU; its register-blocked matmul reaches 1.9 TFLOP/s on an NVIDIA T4, 49–58% of cuBLAS from 512³ up.

A cycle-level simulator and assembler in C++17 for MN1, a fictional vector accelerator with a 64 KiB scratchpad and an async DMA engine. It traps on DMA hazards and writes Perfetto traces, so double buffering and pipeline stalls show up as cycle counts.

A browser tool (Rust → WebAssembly) that reads MLIR pass pipelines and analyzes GPU memory access in MLIR and Triton IR. Its predictions, written before the run, matched all 22 Nsight Compute counters measured on an NVIDIA T4.

Also

  • device-atlas: a catalog of CPUs, GPUs, DSPs and NPUs with a planner that tiles the same GEMM for each
  • compiler-field-guide: my notes on compiler concepts across LLVM, MLIR, Triton and Mojo

🧰 Tools

C++ · MLIR / LLVM · Mojo · Python · Rust

📫 Contact

Pinned Loading

  1. nano-dsp-mlir nano-dsp-mlir Public

    Small MLIR compiler: a dsp tensor dialect lowered to linalg, then tiled and vectorized by Transform-dialect schedules derived from a target model (NEON, AVX2). Includes int8 qmatmul, SIMD Mojo kern…

    C++ 1

  2. vizmlir vizmlir Public

    See how MLIR runs on the GPU, in your browser: threads, warps and memory spaces drawn from mlir-opt output, with proven coalescing and bank-conflict verdicts, Triton IR support and pass-by-pass dif…

    JavaScript 1

  3. json-schema-mlir json-schema-mlir Public

    An out-of-tree MLIR dialect that compiles JSON Schema (Draft 2020-12) into native validators: hand-written lexer, parser and importer, constraint-lattice canonicalization, lowering to LLVM, and a M…

    C++ 1