In 2022, I wrote about the damning fall of events tech company Pollen, founded by Callum Negus-Fancey. The short of it: Pollen seemed to have pulled off the improbable feat of building a business in the notoriously low margin industry of events, surviving Covid-19, and building a solid
GSD Task Manager is a private Eisenhower Matrix for web, iPhone, iPad, and Mac. Sort urgent from important, keep tasks on your device, and start without an account.
Fill out your own 2026 World Cup bracket — group winners, the best third-placed teams, and every knockout tie to the final — then see what the crowd predicts.
A puzzle game about tracing constellations in a living night sky.
Deciding whether to join Antler's startup incubator / accelerator? Learn key scenarios to consider: the VC-funded lifestyle, startup basics, and when you're simply stuck.
Modern GPU Programming For MLSys Contents Modern GPU Programming For MLSys# Machine learning systems sit at the heart of modern AI workloads. In these systems, performance often comes down to the quality of a small number of GPU kernels. Attention kernels, LLM prefill and decode kernels, low-precision block-scaled GEMMs, fused MoE layers, and other large fused kernels all directly shape end-to-end speed in both training and serving. To make these kernels fast, however, we need more than a list of optimization tricks. Modern GPUs are no longer simple variations of the same old design. Recent architectures introduce richer memory spaces, new access patterns, and increasingly specialized execution units. To program them well, we need both a clear mental model of the hardware and a practical understanding of how high-performance kernels are built. This book is about developing both. The book follows a simple progression: first understand the GPU hardware, then learn the programming model we will use, and finally build state-of-the-art kernels step by step. Our main target is the Blackwell generation, and our main running examples are fast matrix multiplication (GEMM) and FlashAttention. Along the way, we will also study the core ingredients behind GPU optimization: data layout, asynchronous data movement, and asynchronous coordination. The material grows out of the Machine Learning Systems course series at Carnegie Mellon University. To make the ideas easier to study and easier to run, this book uses the TIRx Python DSL to build real GPU kernel examples step by step. TIRx stays close to the hardware, which lets us reason about low-level control while still learning through runnable code. How This Book Is Organized# Part I, Understanding the GPU. This part introduces the overall organization of the GPU, general recipes for writing fast kernels, and key concepts such as data layout, asynchronous memory operations, and coordination. It builds the hardware intuition that the rest of the book relies on. Part II, TIRx Overview. This part introduces the key elements of TIRx, which serve as the foundation for the code examples throughout the book. Part III, GEMM: Tiled to SOTA. A complete guide to optimizing a tiled GEMM, built up through TMA pipelining, persistent scheduling, warp specialization, and 2-CTA clusters. Part IV, Flash Attention 4. A complete attention kernel built from the Part III techniques: two MMAs with softmax between them, online-softmax rescaling, causal masking, and GQA. Reference. TIRx language reference and compiler internals. Part I, Understanding the GPU GPU Execution Model What Makes a Kernel Fast Data Layout and Its Notation Tensor Core Operand Layouts Across GPU Generations Async Data Movement: TMA Tensor Cores: tcgen05 Special Memory: TMEM Async Coordination: mbarriers Advanced: Cluster Launch Control Part II, TIRx Overview Introduction to TIRx TIRx Layout API Part III, GEMM: Tiled to SOTA Building a Tiled GEMM GEMM Optimization Path Step 1: Sequential Single-Tile GEMM Step 2: K-Loop Accumulation Step 3: Spatial Tiling (Multi-CTA) Exercises Pipelining GEMM with TMA Step 4: TMA Async Load Step 5: Software Pipeline (PIPE_DEPTH=2) Step 6: Persistent Kernel + Tile Scheduler Exercises Scaling GEMM with Warp Specialization and Clusters Step 7: Warp Specialization + Pipeline Step 8: 2-CTA Cluster Step 9: Multi-Consumer Warp Specialization End-to-End Result Exercises Part IV, Flash Attention 4 Flash Attention 4 Algorithm Shape Tile-Primitive Graph Warp Roles and Scopes Reading the Fragments The Two MMA Phases TMEM Layout and Reuse How Barriers Connect the Roles Pipelining Structure Rescaling and Writeback Causal Masking GQA Support Tile Scheduling Compile and Verify Differences from GEMM Exercises Reference Reference Debugging Warp-Specialized Kernels Compiler Internals TIRx Language Reference Contents
The first portrait arrives
0 READ OURPROMISE TO AMERICA America is stronger than our politics. Politics forces false choices between extremes on right and left. We reject them. Too many Americans have not benefited from the last half-century of economic growth. Americans want an economy that lowers costs, expands opportunity, and rewards the people who work hard every day. Americans want safe communities and institutions that solve problems. We believe in building more to lower costs and expand opportunity. We believe Democrats succeed when we speak to the whole country and value persuasion over purity.We, the undersigned, make this Promise To America. THE PROMISE TO AMERICA Growth, Competition, and Broad ProsperityWe are capitalist, not socialist. We believe in a growing, fair, and competitive economy that rewards hard work, innovation, entrepreneurship, and ownership. Full-time work should make it possible to own a home, raise a family, afford healthcare, and retire with dignity. Economic, permitting, and tax policy should expand opportunity and lower costs for workers, families, entrepreneurs, and those striving to join the middle class, not disproportionately favor those already at the top. Safety, Security, and Human DignityWe want safety, not lawlessness.We believe Americans deserve secure borders, safe communities, honest government, and an orderly immigration system that protects the country, strengthens the economy, and treats people with dignity. We believe America remains indispensable to global stability, democratic values, international security, and strong alliances. In a dangerous and uncertain world, America must lead with strength, purpose, and partnership. Fiscal DisciplineWe are responsible, not reckless.A generation has passed since our party balanced the budget. We will prioritize tackling the national debt honestly. We must pay our bills, and not leave our children in debt. Government That WorksWe believe government should solve problems, not create them.Government matters, but it must work. Public institutions should be competent, accountable, easier to navigate, and capable of building, innovating, and delivering results people experience in everyday life. Government should make life easier. Free Speech, Respect, AND Common PurposeWe are mainstream, not extreme.We believe Americans can disagree without division. We stand for moderation, acceptance, respect, free expression, and democratic pluralism. We embrace a politics of persuasion over purity, contempt, and cultural division. Confident Patriotism and National RenewalWe are proud, not ashamed of America.We believe America’s story is one of extraordinary achievement and unfinished work. We honor America’s strengths and exceptional character while striving to build a freer, stronger, more prosperous, and more perfect union. SIGN THE PROMISE
A modular engine that runs real vendor detection logic from reverse-engineered EDR components against live or replayed Windows telemetry.