What should you check before deploying an AI model?
A benchmark score does not tell you deployment risk. AX-RAY diagnoses 117 risk items across 11 categories on three axes - the model, its operating environment and agent autonomy - and maps each item to regulation in seven jurisdictions.
A performance benchmark measures how well a model does; a safety diagnostic measures what can go wrong. AX-RAY checks 117 risk items across 11 categories on three axes - the model itself (MODEL-SCAN), the operating environment (AX-SCAN) and agent autonomy (AGENT-SCAN) - grades them with a DHS Score, and shows which law, regulation or guidance each item maps to in Korea, the EU, the US, Japan, China, the UAE and Saudi Arabia. Per-model grades are visible on a public leaderboard.
What does AX-RAY diagnose?
MODEL-SCAN covers risks that come from the weights - capability, reliability, causal safety, robustness and safety. AX-SCAN covers risks that change with where and how the same model is deployed. AGENT-SCAN is the newest axis and covers what happens once a model acts on its own: tool permissions, resistance to hijacking, runaway loops and memory.
Why is a performance score not enough?
What breaks in production is often not accuracy. A model accepts a false premise, invents an answer where it should say it does not know, fails to refuse a harmful request, or over-refuses harmless ones. None of that shows up in an accuracy number.
What does the output look like?
Each of the 117 risk items is judged, then rolled up into category grades. The public leaderboard shows per-model grades and a category heatmap. The full definitions of all 117 items are provided only under contract or NDA - if every item were public, a model could be tuned to pass those specific items, and the diagnostic would stop meaning anything.
How does this relate to regulation and export?
The same defect carries different consequences in different markets. In one jurisdiction an item may lead to fines or a sales restriction; in another it stays advisory. The point is to know what you are exposed to before you ship or export. This mapping is reference material to support compliance work - it is not legal advice, and a final determination needs review by counsel in each market.
Frequently asked questions
- How is AX-RAY different from a performance benchmark?
- A performance benchmark measures how well a model does, using metrics like accuracy. AX-RAY measures what can go wrong. Neither replaces the other; a deployment decision needs both.
- Are all 117 items published?
- No. Representative items are shown per category, and the full item detail is provided under contract or NDA. If every item were public, a model could be tuned to pass exactly those items and the diagnostic would lose its meaning.
- Is a grade a final verdict on the model?
- No. It is a measurement under specified items and conditions, and results can change when conditions change. The regulatory mapping is reference material, not legal advice.
Related
This article is based on VIDRAFT public and measured data. Diagnostic results are measurements under the stated items and conditions and can vary by environment. The regulatory mapping is reference material, not legal advice.