Preprint

APT accelerator reports faster 4K AI image generation

Preprint: APT pairs attention-guided pruning with mixed precision and reports major gains at high resolution, while its hardware evaluation uses simulation and modeled energy.

A software and hardware system for diffusion-transformer image generation has reported an end-to-end speedup of up to 8.16 times over an NVIDIA A100 at 4K resolution. In comparisons with the EXION accelerator, the maximum reported gain was 3.01 times. Called APT, the design combines attention-probability-guided pruning and adaptive precision with hardware support for irregular sparsity and dual-precision execution.

APT uses attention probabilities as a unified importance signal. It uses that signal to decide which calculations can be pruned, or skipped, and which should use higher or lower numerical precision. Its design combines adaptive dual-threshold pruning, a timestep-aware attention method called TAFA, and hardware support for irregular sparsity and dual-precision execution.

How APT chooses what to keep

The pruning decisions change across diffusion timesteps. APT's APDT method derives two thresholds for each attention head and timestep from profiled attention-probability standard deviations. Elements above the quantization threshold are assigned 12-bit execution, those between the thresholds use 6-bit execution, and lower elements are pruned. Masks are refreshed every 7 to 10 steps, with a head-wise refresh when the profiled standard deviation drops by more than 5% to 10% of its observed range.

TAFA reuses normalization statistics from the previous timestep to predict probabilities for each tile. In the reported validation, consecutive-step attention factors had similarity above 0.95, while execution masks were more than 97% similar to those produced by full Softmax. At 4K, the associated on-chip memory requirement was up to 103.4 times lower than with naive probability materialization.

Quality held in the reported 1K test

The quality check used COCO 2017 and was run at 1K resolution without fine-tuning or retraining. APT was reported to remain comparable to an FP16 baseline on three measures: inception score, CLIP score and PSNR. Variation in inception score and CLIP score stayed within 2.5%. For PixArt-α, the paper reports FP16-level quality with 57.34% pruning and 9.06% of computation using 12-bit precision.

For self-attention, the reported speedup at 1K was 4.35 to 5.27 times over A100 and 1.55 to 1.81 times over EXION. At 4K, it reached up to 9.71 times over A100 and 3.70 times over EXION.

Across the full pipeline, the reported A100 speedup was up to 3.89 times at 1K and 8.16 times at 4K. The computational savings also grew with resolution: at 4K, pruning alone reduced bit operations, or BOPs, by 53.7% on average, while pruning combined with dual precision reached up to 76.4%.

A strong result with an engineering caveat

APT was implemented in SystemVerilog at the register-transfer level and synthesized in a 14 nm process. The evaluation used an 800 MHz HBM2E system providing 2 TB/s of bandwidth, a cycle-level simulator, modeled HBM2E energy, 3,072 SD-MPUs and a 16 MiB buffer. The architecture includes mask management, address translation, 16 dual-precision MAC lines supporting 12-bit and 6-bit execution, a transpose controller, and a vector-processing unit for Softmax and normalization.

The paper therefore presents results for a synthesized and simulated design rather than measured performance from fabricated APT silicon. Across the tested models and resolutions, APT was reported to deliver up to 14.98 times higher energy efficiency than A100 and 2.04 times higher than EXION. Reported area was 160.11 mm2 and power was 128.26 W at 0.8 V and 800 MHz.

The GPU comparison adds a warning about sparse execution. On an A100 at 1K, even structured sparsity remained slower than dense FlashAttention up to 50% sparsity. The paper attributes that result to kernel-launch, synchronization and scattered-memory overheads, and expects runtime-unstructured sparsity to perform worse on a GPU.

What the preprint leaves open

The evaluation covered PixArt-α, Stable Diffusion 3 and FLUX.1-dev at 1K, 2K and 4K, but the reported quality comparison was at 1K. Tests used batch size 1, and the number of prompts, generated images and repeated runs was not reported. That leaves open how stable the quality and speed figures would be across other prompts, batch sizes, sampling settings and models.

APT was compared with NVIDIA A100, EXION and APT-Base, which used the same hardware without pruning or quantization; EXION was reimplemented, while A100 used FP16 FlashAttention. The study reports summary metrics rather than inferential tests or confidence intervals, so ordinary run-to-run variation is not known. Independent measurements on more models and datasets, plus matched hardware tests, could help show how much of the reported gain comes from APDT, TAFA and the accelerator architecture.

The item is labeled a preprint. Its front matter identifies arXiv:2608.25380v1 dated Aug. 26, 2026, and includes an ACM reference to ICCAD '26 scheduled for Nov. 8 to 12, 2026. The authors report partial support from an IITP grant funded by Korea's MSIT and from Samsung Electronics.

Paper data and sources

Original title: APT: Accelerating Diffusion Transformers via Attention Probability-Guided Pruning and Quantization
Authors: Sungyeob Yoo, Seeyeon Kim, Joonyong Park et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: 10.1145/3831252.3834102
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.