← Back to the collection Suggest a correction ↗
Full-stack recurrence ·
DeepLoop
DeepLoop: Depth Scaling for Looped Transformers
Revisits residual scaling for repeated parameter use, focusing on stable training of deeper looped networks.
Inside the method
[ Residual-scaled shared layers ] × R
Simplified conceptual schematic. Consult the paper for the complete architecture.
- Recurrence family
- Full-stack recurrence
- Depth control
- Deep recurrence
- KV / state strategy
- See paper
Reading note
No public weights were confirmed in this review.
Sources checked 2026-09-15. This catalog does not imply independent reproduction.
Cite this work
@misc{deeploop2026,
title = {DeepLoop: Depth Scaling for Looped Transformers},
author = {Shuzhen Li and Yifan Zhang and Jiacheng Guo and Quanquan Gu and Mengdi Wang},
year = {2026},
eprint = {2607.13491},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2607.13491}
}