Untie the Knots: An Efficient Data Augmentation Strategy for Long-Context Pre-Training in Language Models Paper • 2409.04774 • Published Sep 7, 2024 • 1
Gated Slot Attention for Efficient Linear-Time Sequence Modeling Paper • 2409.07146 • Published Sep 11, 2024 • 21
Stuffed Mamba: State Collapse and State Capacity of RNN-Based Long-Context Modeling Paper • 2410.07145 • Published Oct 9, 2024 • 3
Pre-training Dataset Samples Collection A collection of pre-training datasets samples of sizes 10M, 100M and 1B tokens. Ideal for use in quick experimentation and ablations. • 15 items • Updated Apr 2 • 18
view article Article Meta is back with Muse Glimmer: local, agentic, multimodal, and open source +2 pcuenq, merve, burtenshaw, ariG23498 • 24 days ago • 110
📚 LLM pretraining datasets Collection A collection of datasets for LLM pretraining • 9 items • Updated May 5, 2025 • 30
Benchmarking the Spectrum of Agent Capabilities Paper • 2109.06780 • Published Sep 14, 2021 • 1
LoopMoE: Unifying Iterative Computation with Mixture-of-Experts for Language Modeling Paper • 2606.04438 • Published Jun 3 • 2
EMO: Pretraining Mixture of Experts for Emergent Modularity Paper • 2605.06663 • Published May 7 • 13
LTX-2.3 Creative Lab Collection LoRAs and IC-LoRAs, trained on the LTX-2.3 model • 26 items • Updated Jul 28 • 90
Ornith-1.0 Collection Ornith-1.0 is  a family of open-source LLMs specialized for agentic coding. • 8 items • Updated 15 days ago • 391
ReFT: Representation Finetuning for Language Models Paper • 2404.03592 • Published Apr 4, 2024 • 101
EasyControl: Adding Efficient and Flexible Control for Diffusion Transformer Paper • 2503.07027 • Published Mar 10, 2025 • 30
Speed by Simplicity: A Single-Stream Architecture for Fast Audio-Video Generative Foundation Model Paper • 2603.21986 • Published Mar 23 • 125
view article Article 🎨 VAE (Variational Autoencoders) — When AI learns to dream! 🌈🧠RDTvlokip • Oct 19, 2025 • 4