-
JudgeLM: Fine-tuned Large Language Models are Scalable Judges
Paper • 2310.17631 • Published • 35 -
Prometheus: Inducing Fine-grained Evaluation Capability in Language Models
Paper • 2310.08491 • Published • 57 -
Generative Judge for Evaluating Alignment
Paper • 2310.05470 • Published • 1 -
Calibrating LLM-Based Evaluator
Paper • 2309.13308 • Published • 12
Andrew Reed
andrewrreed
AI & ML interests
Applied ML, Practical AI, Inference & Deployment, LLMs, Multi-modal Models, RAG
Organizations
Curated resources that support the use of LLMs to serve as automatic evaluators of other LLM outputs.
Eval Leaderboards
- Running5k
Arena Leaderboard
🏆5kView the LMArena model performance leaderboard
- Running on CPU Upgrade14.1k
Open LLM Leaderboard
🏆14.1kTrack, rank and evaluate open LLMs and chatbots
- Running on CPU Upgrade7.68k
MTEB Leaderboard
📊7.68kEmbedding Leaderboard
- RunningAgentsFeatured590
LLM-Perf Leaderboard
🏆590Compare LLM hardware performance and find the best model
AI x Audio
Hallucination Detection
-
vectara/hallucination_evaluation_model
Text Classification • 0.1B • Updated • 170k • 364 -
notrichardren/HaluEval
Viewer • Updated • 35k • 180 -
TRUE: Re-evaluating Factual Consistency Evaluation
Paper • 2204.04991 • Published • 1 -
Fine-grained Hallucination Detection and Editing for Language Models
Paper • 2401.06855 • Published • 4
Small, but mighty chat models
Awesome Spaces
- Running on ZeroAgents120
StableDesign
🏆120Generate a furnished interior from an empty room photo
- Running on ZeroAgentsFeatured5.45k
IllusionDiffusion
👁5.45kGenerate stunning high quality illusion artwork
- Running on ZeroAgentsFeatured1.6k
InstantMesh
📚1.6kCreate a 3D model from an image in 10 seconds!
- Runtime errorAgentsFeatured184
Sing an idea ➡️ Music
🔥184Bring song ideas to life
LLM as a Judge
Curated resources that support the use of LLMs to serve as automatic evaluators of other LLM outputs.
-
JudgeLM: Fine-tuned Large Language Models are Scalable Judges
Paper • 2310.17631 • Published • 35 -
Prometheus: Inducing Fine-grained Evaluation Capability in Language Models
Paper • 2310.08491 • Published • 57 -
Generative Judge for Evaluating Alignment
Paper • 2310.05470 • Published • 1 -
Calibrating LLM-Based Evaluator
Paper • 2309.13308 • Published • 12
Hallucination Detection
-
vectara/hallucination_evaluation_model
Text Classification • 0.1B • Updated • 170k • 364 -
notrichardren/HaluEval
Viewer • Updated • 35k • 180 -
TRUE: Re-evaluating Factual Consistency Evaluation
Paper • 2204.04991 • Published • 1 -
Fine-grained Hallucination Detection and Editing for Language Models
Paper • 2401.06855 • Published • 4
Eval Leaderboards
- Running5k
Arena Leaderboard
🏆5kView the LMArena model performance leaderboard
- Running on CPU Upgrade14.1k
Open LLM Leaderboard
🏆14.1kTrack, rank and evaluate open LLMs and chatbots
- Running on CPU Upgrade7.68k
MTEB Leaderboard
📊7.68kEmbedding Leaderboard
- RunningAgentsFeatured590
LLM-Perf Leaderboard
🏆590Compare LLM hardware performance and find the best model
Small, but mighty chat models
AI x Audio
Awesome Spaces
- Running on ZeroAgents120
StableDesign
🏆120Generate a furnished interior from an empty room photo
- Running on ZeroAgentsFeatured5.45k
IllusionDiffusion
👁5.45kGenerate stunning high quality illusion artwork
- Running on ZeroAgentsFeatured1.6k
InstantMesh
📚1.6kCreate a 3D model from an image in 10 seconds!
- Runtime errorAgentsFeatured184
Sing an idea ➡️ Music
🔥184Bring song ideas to life