Distilled models from paper Thinking at the Right Size: Amortized Distillation Across Post-Trained LLMs (https://arxiv.org/abs/2608.22854).
-
Harvard-DCML/ADAPT-Qwen3-2.3B-Instruct
Text Generation • 0.6B • Updated • 469 -
Harvard-DCML/ADAPT-Olmo3-4.3B-Instruct
Text Generation • 1B • Updated • 454 -
Harvard-DCML/ADAPT-Llama3.1-4.8B-Instruct
Text Generation • 1B • Updated • 469 -
Harvard-DCML/ADAPT-Qwen3-8.5B
Text Generation • 0.5B • Updated • 310