Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs Paper • 2609.29845 • Published 17 days ago • 105
GigaChat 3.5 Collection GigaChat 3.5 is a large-scale Mixture-of-Experts (MoE) Hybrid model with 432B total parameters • 8 items • Updated Sep 10 • 20
The More Popular, The Harder to Forget: Adaptive Popularity for LLM Unlearning Paper • 2608.14229 • Published Aug 14 • 17