Hi, I’m

Chris Fregly

AI systems performance engineer · Product leader · Founder · Advisor · 3× O’Reilly author

I build and explain high-performance AI systems. My work spans GPU kernels, distributed training, high-throughput inference, and the product systems that move AI into production.

Chris Fregly
1,062pages in the 2025 book
1.7K+GitHub stars
440K+course enrollments
100K+community members worldwide
780K+YouTube views

Recent work

Systems depth, packaged for builders

The current work connects low-level performance engineering with the product discipline needed to ship useful AI.

Open source · 2026

Agent Harness Optimization

An evidence-first workbench for evaluating agent prompts, tools, transcripts, and traces without losing the receipts behind each conclusion.

Prompts · Traces · Evals · Evidence

Open source · 2026

GPU Performance Tuning

A set of agent workflows for profiling and improving GPU inference, including benchmarking, quantization, speculative decoding, and performance reports.

32 focused workflows

Books

Three O’Reilly books across the AI stack

From production data science to generative AI and the performance of the full system beneath it.

Cover of Generative AI on AWS

O’Reilly · 2023 · New translations in 2024 and 2025

Generative AI on AWS

Build context-aware multimodal applications with foundation models, fine-tuning, reinforcement learning, RAG, and production deployment patterns.

Cover of Data Science on AWS

O’Reilly · 2021

Data Science on AWS

Implement end-to-end machine learning pipelines with data engineering, model training, tuning, and production deployment on AWS.

Course

Generative AI with Large Language Models

A practical course built with DeepLearning.AI and AWS for people who want to understand the full generative AI lifecycle.

Co-instructor · Intermediate

Learn the full generative AI lifecycle

The course covers transformers, model selection, scaling laws, fine-tuning, evaluation, reinforcement learning, inference, and deployment. It includes 47 lessons and three graded assignments.

  • 440K+enrollments
  • 4.8learner rating
  • 24languages

Talks and media

Recent conversations from kernels to clusters

Talks, interviews, and technical sessions from the last two years.

Talk · June 2026

AI Performance Engineering

Hacking AI Accelerators

A systems-level look at the hardware and software choices that shape AI performance.

Watch or listen

Talk · March 2026

NVIDIA GTC recap

Performance Highlights from GTC 2026

Recent advances in inference, disaggregated prefill and decode, and the systems work behind faster models.

Watch or listen

Podcast · March 2026

SuperDataScience · Episode 973

AI Systems Performance Engineering

A conversation about GPU computing, full-stack profiling, and what engineers need to understand beneath modern AI frameworks.

Watch or listen

Interview · February 2026

MLOps Community · Episode 363

Software and Hardware Codesign

How PyTorch, CUDA, and NVIDIA GPUs fit together when performance and cost matter in production.

Watch or listen

Talk · July 2025

AI Performance Engineering

Dynamic Inference Tuning with CUDA and vLLM

Practical techniques for adapting inference systems as request patterns, memory pressure, and latency targets change.

Watch or listen

Talk · June 2025

O’Reilly AI Superstream

High-Performance Agentic AI Inference Systems

Performance patterns for agentic workloads using DeepSeek, NVIDIA Dynamo, vLLM, CUDA, and PyTorch.

Watch or listen

Community

AI performance is a team sport

I cohost monthly technical sessions with Antje Barth for engineers who build the systems behind modern AI. Our global network reaches more than 100,000 people worldwide. The flagship Meetup group has hosted 375 past events, and the YouTube channel has 193 videos with more than 780,000 views.

100K+worldwide members
375past events
193videos
780K+views