摘要
A 16 n m FinFET transformer-based diffusion model processor chip is fabricated for supporting class-conditional DiT-XL/2 and text-to-image PixArt-a with 98 ms and 134 ms generation time per step with 7.37TOPS and 1477 mW at 400 MHz. Classifier-free guidance batching reduces external memory access (EMA) for weights by 50%, and weight-reordered quantization brings another 37% reduction. Hybrid microscaling blocking reduces featuremap buffer size by 37%, while approaching floating-point quality.