URM 与 UT 是否已被前沿模型采用?还是被遗忘在论文堆里?
Have URMs and UTs been integrated into frontier models? Or did they disappear into the dustbin of forgotten papers? [D]
Reddit 用户讨论 Universal Reasoning Model(URM)与 Universal Transformer(UT)是否已被大型科技公司的前沿模型采用。
UT-based small models, despite being trained from scratch on these tasks without internet-scale pre-training, consistently outperform most standard Transformer-based Large Language models (LLMs) by a significant margin. 👈
The Universal Transformer (UT) extends the standard Transformer by introducing recurrent computation over depth. Instead of stacking L distinct layers, the UT applies a single transition block repeatedly to refine token representations. For an input sequence x with embedding matrix H0 ∈ Rn×d , the UT updates states as
Ht+1 = LayerNorm(Ht + MHA(Ht ))
followed by a shared position-wise transition function
Ht+1 ← LayerNorm( Ht+1 + Transition(Ht+1 )), t = 0, . . . , T − 1
where Transition is either a feed-forward network or separable convolution. To encode both position and refinement depth, UT adds 2-D sinusoidal embeddings at each step.
https://i.imgur.com/S8TAh3f.png
(below) Figure 2 : Illustration of our Universal Reasoning Model (URM) architecture. The left shows a standard Transformer layer stack, while the right illustrates the URM with fixed loops, ACT loops, and the ConvSwiGLU module. For illustrative purposes, components such as embeddings, residual connections, RMSNorm, positional encodings, and other modules are omitted, x in right figure represents the first x loops of the inner loop in forward-only mode, TBPTL represents our proposed Truncated Backpropagation Through Loops.
https://i.imgur.com/uEuhAeC.png
Have frontier labs at big tech companies already implemented the enhancements of URMs and UTs? Or is this cutting-edge research still waiting for its day in the sun?
Below are the original paper, a simplified blog, and a talk.
Paper : https://arxiv.org/pdf/2512.14693
blog : https://bdtechtalks.substack.com/p/inside-urm-the-architecture-beating
youtube talk : https://www.youtube.com/watch?v=RxNPFCYCFBU
submitted by /u/moschles
[link] [留言]
来源:r/MachineLearning · reddit.com