I'm based in Vancouver, where I work on GPU systems and compilers for language models. I'm building FindTensor, an experimental compiler and runtime. Previously, I worked with Tor Aamodt at UBC on GPU architecture simulation.
Manuscripts, with code and supporting material:
- The Work a Verifier Needs — token verification in language models.
- A reachable-state separation between attention compression and its certificates
- SQ learning and dimension complexity
- Memory return in quantum machines
I also study Erdős problems in Lean.
- FlashInfer: A MoE routing fix was merged, and my dispatch diagnostics and configuration API were incorporated upstream. I also reported and reproduced an FP8 calibration bug that was fixed.
- NVIDIA CCCL / CUB: Merged changes to DeviceMergeSort, DeviceMerge and DeviceMemcpy.
- Accel-Sim: Merged container update for Ubuntu 24.04 and CUDA 13.1.
Other contributions and proposals
- Further FlashInfer and CUB / Thrust changes under review.
- SGLang: Verification, planning and serving fixes, under review.
- GPU MODE QR benchmark checks and torchcomms extension packaging, under review.
- Accel-Sim application compatibility, under review.
- Triton FP8 dot product and GPGPU-Sim CUDA 13 proposals, closed without merge.
- CAIRN — experimental language for CPU and GPU programs, designed for AI agents.
- KernelIndex — an index of GPU performance measurements.
- B200 kernels — CUDA C++, CuTe DSL and Triton implementations.
- Command A+ benchmarks — vLLM serving measurements on two H100s.
- H100 serving estimator — estimates GPU time per request from 91 benchmark runs.
- SmolLM2 CPU checks — numerical comparisons with Transformers.
- Tensor parallel reference — a CPU implementation for studying execution across processes.



