AGO AI Quality Gate: Evidence-First Release Decisions for Retrieval-Augmented Generation
Abstract
Enterprises adopting retrieval-augmented generation (RAG) face a recurring operational decision: promote, revise, or block a system version. The evidence is incomplete and the metrics come from fallible LLM judges. We report on AGO AI Quality Gate (AGO), an evidence-first quality-gate framework deployed in industrial RAG assessment engagements. AGO integrates four key components: a four-state decision model that treats missing data and judge errors as explicit outcomes; layered scoring combining deterministic checks, local guardrails, and structured LLM evaluation; a stratified beta-binomial gate that quantifies regression risk probabilistically; and a mandatory meta-evaluation protocol to validate the LLM judge before it influences decisions. Since engagement data is proprietary, we evaluate the judge layer on RAGBench, a public benchmark of 100k annotated RAG traces across 12 datasets. On identical stratified test samples (N=1200 per judge), a low-cost judge (gpt-4.1-nano) detects non-adherent answers barely above chance (AUROC 0.603 [0.570, 0.634]), despite producing flawless protocol output, while gpt-4o reaches 0.783 [0.756, 0.807] -- yet its per-domain performance still ranges from 0.62 to 0.88. A fixed-seed gate study spanning regression, no change, and improvement quantifies unsafe promotion, false-alarm cost, and improvement throughput. Under regression, the decision-grade profile reduces unsafe promotion to 22.2%-35.1%, against 29.3%-41.8% for a naive gate. These results support the design choices that judge quality must be measured per engagement and that point estimates alone are not a release decision.
Community
Most RAG evaluation tools stop at a score. In enterprise engagements we needed a release decision we could defend: promote, manual review, block, or "not evaluable" when the evidence is missing.
AGO AI Quality Gate is the evidence-first gate we built and deploy at Protom. Three takeaways:
- A cheap LLM judge can fail silently: on RAGBench (N=1200 per judge), gpt-4.1-nano reached AUROC 0.603 [0.570, 0.634] while producing perfectly well-formed output. gpt-4o reached 0.783, yet ranged from 0.62 to 0.88 across domains. Judges must be validated per engagement, not certified once.
- Point estimates are not release decisions: under regression, a stratified beta-binomial gate cut unsafe promotion to 22.2%-35.1%, versus 29.3%-41.8% for a naive gate, at an explicit false-alarm cost.
- Missing evidence is a first-class outcome, never a silent pass.
Accepted at NFMCP 2026 (ECML PKDD 2026 Workshops). Feedback welcome!
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Why RAGs Hallucinate: Penalty-Aware Evaluation of Retrieval-Augmented Generation Systems with Knowledge-Gap Canaries (2026)
- CARGO: Context-Aware Retrieval-Gated Evaluation of Agentic AI in Production (2026)
- Re:CAP - Auditing Retrieval Coverage in Production RAG Pipelines (2026)
- RAGScope: A Leakage-Controlled, Cost-Aware Evidence-Gating Protocol for RAG Hallucination Triage (2026)
- Dear Algo: A Precision-First Agentic Intent Layer for Unified Search and Recommendation (2026)
- TWIST: A Proposed Benchmark for Intervention Quality in Conversational Memory, with a Human-Validated Draft-Alignment (2026)
- Same Agent, Different Answers: A Repeat-Aware Audit of Corpus-Induced Answer Churn in Retrieval-Augmented QA (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2610.01218 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper