Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference Paper • 2609.05275 • Published 4 days ago • 15
Unlocking Lossless Speedups in LLMs via Discrete Diffusion Paper • 2609.04010 • Published 5 days ago • 8
Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference Paper • 2609.05275 • Published 4 days ago • 15
Calibrating Beyond English: Language Diversity for Better Quantized Multilingual LLM Paper • 2601.18306 • Published Jan 26
Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation Paper • 2608.15062 • Published 13 days ago • 11
Cerebras REAP Collection Sparse MoE models compressed using REAP (Router-weighted Expert Activation Pruning) method • 30 items • Updated Feb 25 • 152
Cerebras REAP Collection Sparse MoE models compressed using REAP (Router-weighted Expert Activation Pruning) method • 30 items • Updated Feb 25 • 152