Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
ginigen-ai
PRO
ginigen-ai
8
22
150
Follow
dipankarsarkar's profile picture
MetaMorpheusG's profile picture
AbstractPhil's profile picture
56 followers
Β·
126 following
AI & ML interests
None yet
Recent Activity
liked
a dataset
about 1 hour ago
FINAL-Bench/AX-RAY
liked
a Space
about 1 hour ago
FINAL-Bench/AX-RAY
reacted
to
SeaWolf-AI
's
post
with π₯
about 1 hour ago
AX-Ray: Safety Diagnostics for AI/AX Models AI models can no longer be evaluated only by capability scores. As models move into public services, enterprise workflows, scientific research, and administrative decision support, we need a second layer of evaluation: whether the model behaves safely, structurally, and consistently under real deployment conditions. VIDRAFT AX-Ray is a public AI/AX safety diagnostic initiative powered by FINAL-Bench Diagnostics. AX-Ray evaluates models across a structured guideline framework, including model-level safety, AX deployment readiness, and agent/service operation risks. The public diagnostic catalog contains 117 diagnostic items, mapped to legal, regulatory, ethical, and religious-law governance contexts so that safety review can be discussed in a form closer to real institutional responsibility. A central finding of AX-Ray is causal leakage: a structural defect where information that should not influence an earlier reasoning state appears to affect model behavior. AX-Ray presents a public case of diagnosing, reproducing, and demonstrating causal leakage in two general-purpose public models. This matters because such defects are not exposed by ordinary benchmark scores. A model can appear capable while still carrying hidden safety or integrity risks. Explore the live leaderboard, diagnostic reports, and public dataset here: - AX-Ray Space: https://huggingface.co/spaces/FINAL-Bench/AX-RAY - AX-Ray Dataset: https://huggingface.co/datasets/FINAL-Bench/AX-RAY - Technical Article: https://huggingface.co/blog/FINAL-Bench/ax-ray AX-Ray is intended as a practical guideline for moving AI evaluation beyond βhow smart is the model?β toward βcan this model be trusted, governed, and deployed safely?β
View all activity
Organizations
ginigen-ai
's activity
All
Models
Datasets
Spaces
Buckets
Papers
Collections
Community
Posts
Upvotes
Likes
Articles
published
an
article
about 1 month ago
view article
Article
Does Your LLM Know *When It's About to Be Wrong*?
ginigen-ai
β’
Jul 1
β’
21