What is the best Korean LLM on the GPQA Diamond science benchmark?
VIDRAFT's Darwin-398B-JGOS reaches 90.9% on GPQA Diamond — 3rd in the world and #1 among Korean models (base-only, July 2026 snapshot).
As of July 2026, VIDRAFT's Darwin-398B-JGOS is the top-ranked Korean model on GPQA Diamond with 90.9%, placing 3rd globally behind only two Chinese frontier models — Kimi-K3 (93.5) and GLM-5.2 (91.2). This is a base-only, single-benchmark result.
What is the GPQA Diamond benchmark?
The questions span biology, chemistry, and physics and are authored by domain experts, then filtered so that even non-experts with web access still struggle to answer them. Because it is resistant to memorized or searchable answers, GPQA Diamond is widely used to measure the frontier of scientific reasoning in large language models.
How does Darwin-398B-JGOS rank on GPQA Diamond?
Only two Chinese frontier models score higher: Kimi-K3 (93.5) and GLM-5.2 (91.2). Darwin sits ahead of DeepSeek-V4-Pro (90.1), Alibaba Qwen3.5-397B (88.4), and Nvidia Nemotron-3-Ultra-550B (87.9). The next-best Korean model trails Darwin by roughly 4.6 points, making it the clear domestic leader on this benchmark.
How was Darwin-398B-JGOS built?
VIDRAFT combined models from the Gemma-4 and Qwen-3.5 lineages through an evolutionary merge process that searches for the best combination of weights, instead of training a new 398B-parameter model from zero. This lets a comparatively small compute budget (~24 GPUs) produce a top-3 GPQA Diamond result, illustrating that careful model composition can rival brute-force pretraining scale.
What are the limits of this GPQA Diamond result?
GPQA Diamond scores can shift meaningfully with the evaluation setup (prompting, sampling, and scoring choices), and this figure reflects base-model performance at a single point in time (July 2026). A strong GPQA Diamond score signals excellent science-reasoning ability, but it does not by itself establish overall superiority across coding, language, safety, or real-world tasks.
Frequently asked questions
- Is Darwin-398B-JGOS the #1 AI model in the world?
- No. Darwin ranks 3rd globally on GPQA Diamond with 90.9%. Two Chinese frontier models score higher — Kimi-K3 (93.5) and GLM-5.2 (91.2). Darwin is #1 among Korean models.
- Did VIDRAFT train Darwin-398B-JGOS from scratch?
- No. Darwin was built by evolutionary merging of open models from the Gemma-4 and Qwen-3.5 lineages on about 24 GPUs, not from-scratch pretraining — an approach VIDRAFT calls 'method over scale'.
- Does a 90.9% GPQA Diamond score mean Darwin is the best model overall?
- No. This is a base-only, single-benchmark snapshot. GPQA Diamond measures graduate-level science reasoning and scores shift with evaluation setup, so it is not a verdict on overall capability.
Related
This article is based on VIDRAFT public, measured data and external sources. Performance figures are measurements under the stated conditions and may vary by environment.