maybe you can train a model and i can use it as a benchmark π
Emin Temiz PRO
etemiz
AI & ML interests
Alignment
Recent Activity
repliedto their post about 4 hours ago
maybe one of the hard parts of my kind of fine tuning is evals.
where do you want model to go? how do you find true answer of a hardly debated issue?
instead of manually writing answers in many domains (which i can't do, i don't know many answers in many domains) i rely on a few tricks:
- find other aligned llms and get ideas from them
- rank many llms in AHA leaderboard and get ideas from top ones also rejecting the worst ones
- do mixture of agents of the above to get a collective answer
- and lately, compare the answers coming from my fine tunes with base models, assume my fine tune is preferred if there is a difference
these still can't find perfect answers but they kick the model in the right direction. and that may be a big deal. can't claim my fine tune (ostrich) knows every truth. but it may make more sense to claim if the base differs with fine tune most probably fine tune is the better answer.
making better evals ends up training better models. and produce better AHA ranking. which further sharpens evals. this feedback loop is going to be useful for a while.
liked a model about 22 hours ago
JonathanColetti/Qwen3.8-27B-Uncensored liked a model 1 day ago
Qwen/Qwen3.8-27B