VIDRAFT.
VIDRAFT / Insights / Third-party Verification
Third-party Verification

Where can you check AI benchmark results that a third party actually scored?

Self-reported numbers and externally scored records are not the same thing. This page collects only results where someone other than us did the scoring - a government-run leaderboard, public blind benchmarks and organiser re-verification.

Published 2026-08-153 min readby VIDRAFT
Quick answer

Most AI performance claims are self-measured. What can be checked is a record where someone else did the scoring. For VIDRAFT that means: the K-AI Leaderboard run by Korea's Ministry of Science and ICT and NIA, where JGOS-31B-Citizen ranked first overall; 16 first places on Polaris, a public blind benchmark for drug prediction; the Google x Hugging Face Fast Gemma Challenge, where the organisers re-ran submissions on a private prompt set before accepting a rank; and an Excellent rating on a NIPA-commissioned, KAIT-administered programme. Each of these can be checked outside our site.

Why are self-reported numbers hard to trust?

Because the scorer and the contestant are the same party. If the side choosing the conditions also produces the score, that is not verification.

The same model produces very different numbers depending on prompts, hardware and measurement method. When a company measures its own model under its own conditions, the result is a reference point, not a verification. Press coverage does not change that: an article reporting a company announcement is journalism, not independent scoring. What counts is a separate scorer and a result that stays visible outside the company.

What has a government body scored?

On the K-AI Leaderboard run by Korea's Ministry of Science and ICT and NIA, JGOS-31B-Citizen ranked first overall.

The K-AI Leaderboard evaluates Korean-language and public-sector capability under blind conditions - contestants do not see the answers in advance and the organising body does the scoring. On the same board, partner company Ourbox placed second in the 30B+ tier with Ourbox-31B-JGOS, as reported by several outlets. Separately, a GPU lease support programme commissioned by NIPA and administered by KAIT was rated Excellent.

What can be checked on public leaderboards?

16 first places on Polaris drug-prediction benchmarks, and the verified-record rank on the Google x Hugging Face Fast Gemma Challenge.

Polaris is a public blind benchmark for drug prediction spanning absorption, distribution, metabolism, excretion, toxicity, potency and kinase selectivity. Rankings are visible to anyone at polarishub.io. The Fast Gemma Challenge does not accept submitted numbers at face value: organisers re-run them on a private prompt set before a rank is recognised. VIDRAFT passed with 510.58 TPS at a quality bar of PPL 2.39.

Where does verification end and announcement begin?

The items above were scored externally. Other performance figures are our own measurements, and we label them that way.

Scientific-reasoning benchmark records and internal evaluations, for instance, are numbers we measured and published ourselves. Even when they appear in the press, if the source is our own announcement we do not classify them as independent verification. Keeping that line visible serves us better over time. Verified items link out so they can be checked directly; self-measured figures are published with their measurement conditions.

Frequently asked questions

If it was in the news, is it verified?
An article reports a fact but does not score performance. Coverage that relays a company announcement is journalism, not independent evaluation. What decides verification is who did the scoring.
Who runs the K-AI Leaderboard?
It is run by Korea's Ministry of Science and ICT together with the National Information Society Agency (NIA), under blind evaluation.
How do I check the Polaris rankings?
polarishub.io publishes rankings per benchmark, so anyone can compare directly.

Sources

Related

Inference acceleration
Which AI took the verified #1 spot in the Fast Gemma Challenge?
Science AI
Which AI model best predicts oral drug intestinal absorption?
AI Safety Diagnostics
What should you check before deploying an AI model?
↖ Home - vidraft.net

This page labels a result as verified only when an external body scored it or it can be checked on a public leaderboard. Self-measured figures are labelled as such. All figures are values under stated conditions and can vary by environment.