AHA · LOOPED TRANSFORMERSubmit a paper ↗
← Back to the collection
Full-stack recurrence ·

DeepLoop

DeepLoop: Depth Scaling for Looped Transformers

Shuzhen Li, Yifan Zhang, Jiacheng Guo, Quanquan Gu, Mengdi Wang

Revisits residual scaling for repeated parameter use, focusing on stable training of deeper looped networks.

Inside the method

[ Residual-scaled shared layers ] × R

Simplified conceptual schematic. Consult the paper for the complete architecture.

Recurrence family
Full-stack recurrence
Depth control
Deep recurrence
KV / state strategy
See paper

Reading note

No public weights were confirmed in this review.

Sources checked 2026-09-15. This catalog does not imply independent reproduction.

Cite this work

@misc{deeploop2026,
  title = {DeepLoop: Depth Scaling for Looped Transformers},
  author = {Shuzhen Li and Yifan Zhang and Jiacheng Guo and Quanquan Gu and Mengdi Wang},
  year = {2026},
  eprint = {2607.13491},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2607.13491}
}

Suggest a correction ↗