AI Systems Performance Engineering
A 1,062-page O’Reilly field guide and active open source lab for GPU architecture, CUDA, PyTorch, distributed training, and modern inference serving.
1.7K+ GitHub stars · 2,700+ commits
Hi, I’m
AI systems performance engineer · Product leader · Founder · Advisor · 3× O’Reilly author
I build and explain high-performance AI systems. My work spans GPU kernels, distributed training, high-throughput inference, and the product systems that move AI into production.
Recent work
The current work connects low-level performance engineering with the product discipline needed to ship useful AI.
A 1,062-page O’Reilly field guide and active open source lab for GPU architecture, CUDA, PyTorch, distributed training, and modern inference serving.
1.7K+ GitHub stars · 2,700+ commits
An evidence-first workbench for evaluating agent prompts, tools, transcripts, and traces without losing the receipts behind each conclusion.
Prompts · Traces · Evals · Evidence
A set of agent workflows for profiling and improving GPU inference, including benchmarking, quantization, speculative decoding, and performance reports.
32 focused workflows
Books
From production data science to generative AI and the performance of the full system beneath it.
Optimize training and inference across GPU hardware, CUDA kernels, compilers, PyTorch, networking, storage, and multinode systems.
Build context-aware multimodal applications with foundation models, fine-tuning, reinforcement learning, RAG, and production deployment patterns.
Implement end-to-end machine learning pipelines with data engineering, model training, tuning, and production deployment on AWS.
Course
A practical course built with DeepLearning.AI and AWS for people who want to understand the full generative AI lifecycle.
The course covers transformers, model selection, scaling laws, fine-tuning, evaluation, reinforcement learning, inference, and deployment. It includes 47 lessons and three graded assignments.
Talks and media
Talks, interviews, and technical sessions from the last two years.
AI Performance Engineering
A systems-level look at the hardware and software choices that shape AI performance.
Watch or listenNVIDIA GTC recap
Recent advances in inference, disaggregated prefill and decode, and the systems work behind faster models.
Watch or listenSuperDataScience · Episode 973
A conversation about GPU computing, full-stack profiling, and what engineers need to understand beneath modern AI frameworks.
Watch or listenMLOps Community · Episode 363
How PyTorch, CUDA, and NVIDIA GPUs fit together when performance and cost matter in production.
Watch or listenAI Performance Engineering
Practical techniques for adapting inference systems as request patterns, memory pressure, and latency targets change.
Watch or listenO’Reilly AI Superstream
Performance patterns for agentic workloads using DeepSeek, NVIDIA Dynamo, vLLM, CUDA, and PyTorch.
Watch or listenCommunity
I cohost monthly technical sessions with Antje Barth for engineers who build the systems behind modern AI. Our global network reaches more than 100,000 people worldwide. The flagship Meetup group has hosted 375 past events, and the YouTube channel has 193 videos with more than 780,000 views.