Direct Preference Optimization: Your Language Model is Secretly a Reward Model Paper • 2305.18290 • Published May 29, 2023 • 68
My models: daily driver rotation Collection A rotating list of models I created and currently use as daily drivers. From my many models, these are the ones I’m actively using. • 7 items • Updated 9 days ago • 10
Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models Paper • 2606.03748 • Published Jun 2 • 19
SafeDiffusion-R1: Online Reward Steering for Safe Diffusion Post-Training Paper • 2605.18719 • Published May 18 • 7
view article Article Chitos: From Detection to Proof — An Autonomous Security AI That Actually Exploits FINAL-Bench • 26 days ago • 19
Refusal in Language Models Is Mediated by a Single Direction Paper • 2406.11717 • Published Jun 17, 2024 • 15