Residual Connections Rethought: How Kimi’s ‘Attention Residuals’ Fixed a 10-Year-Old Transformer Flaw

Residual connections feature image showing attention-based depth routing in transformer layers

Standard residuals Fixed, uniform weights Embedding h₁ Layer 1 Layer 2 Layer 3 Each layer only sees the accumulated sum Attention Residuals Learned, input-dependent Embedding h₁ Layer 1 Layer 2 Layer 3 Layer 3 selectively attends to any earlier layer Residual connections are one of those rare ideas in deep learning that became so successful, … Read more

ChatGPT Physics Breakthrough Explained: How GPT-5.2 Broke The “Zero” Rule, And What Didn’t Change

ChatGPT Physics feature image: GPT-5.2 “zero rule” loophole shown as a kinematic wall in a lab scene

Introduction Some days in theoretical physics feel like mountain climbing. You spend hours inching upward through algebra, you finally reach a viewpoint, and the “beautiful simple formula” everyone promised turns out to be hiding behind a boulder labeled “one more identity.” Then there are days when a language model strolls by, points at your pile … Read more

Drifting Models: The One-Step Image Generator That Trains The “Diffusion Steps” Away

Drifting Models feature image showing diffusion steps compressed into one-step generation.

Drifting Models: The One-Step Image Generator That Trains The “Diffusion Steps” Away Play Introduction Most generative models feel like they’re doing the same ritual: start with noise, take many tiny steps, hope you land on something beautiful before your GPU fan files for overtime. The paper Generative Modeling via Drifting (arXiv: 2602.04770v1) proposes a different … Read more