microsoft/VibeVoice-ASR-Streaming-1.5B Automatic Speech Recognition • 3B • Updated 3 days ago • 715 • 30
Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090 Paper • 2608.27370 • Published 10 days ago • 36
Muse Glimmer Collection Muse Glimmer 30B: multimodal agentic model for local deployment. BF16 weights, GGUF k-quants, ExecuTorch builds, DFlash drafter. • 4 items • Updated 27 days ago • 107
ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF Image-Text-to-Text • 27B • Updated 4 days ago • 297k • 429
Running on CPU Upgrade Featured 3.29k The Smol Training Playbook 📚 3.29k The secrets to building world-class LLMs
QUASAR-QAT/Qwen3.8-27B-QUASAR-NVFP4 Image-Text-to-Text • 28B • Updated 9 days ago • 18.9k • 109
Granite 4.2 Language Models Collection Efficient reasoning and thinking language models for multilingual generation, coding, and AI assistant workflows. • 3 items • Updated 12 days ago • 37
MoE-SpAc: Efficient MoE Inference Based on Speculative Activation Utility in Heterogeneous Edge Scenarios Paper • 2603.09983 • Published Feb 12 • 4
FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution Paper • 2608.16157 • Published 20 days ago • 107