view article Article The Transformers modeling backend serves DeepSeek-R1 at native speed hmellor • 9 days ago • 2
view article Article The Transformers modeling backend is now as fast as native vLLM hmellor • 27 days ago • 1