📰 2026.07.20 비드래프트, 완전공개형 파운데이션 모델 ‘Aether-7B-5Attn’ 공개 — “AI 주권 구현” 조선비즈 → 중앙일보 → 동아일보 →
🇰🇷 K-AI #1 · GPQA Diamond 90.9% · Polaris 14× Champion · NIPA GPU-Support Project · Excellent

Evolving AI.
Thinking AGI.

Building a Pre-AGI platform in Korea, aiming to become a scientific research company — Darwin delivers results, AETHER designs the future.

VIDRAFT aims to become a scientific research company. Our registered lines of business include research and development in physics, chemistry and biology, and in medicine and pharmacology — and that is the direction we work in: research first, products from the results.
2026년 29째주 더브이씨 조회수
Live Ecosystem Dashboard
The Darwin Ecosystem at a glance
Updated
📈 Top Official Models by Downloads
🚀 Model Release Momentum (2026)
💎 GPQA Diamond — Darwin Family
🌐 Downloads by Source

Three Axes

Everything VIDRAFT builds sits on one of three axes — and each one feeds the other two.

🧠01

AGI

Proprietary foundation LLM

Our own architecture, trained from scratch — the work of moving from Pre-AGI toward AGI. AETHER places multiple sequence-mixing mechanisms on a Latin square so the question stays answerable.

Aether-7B-5Attn published fully open · 12 patent applications · 317 claims
Explore AETHER →
⚛️02

Quantum Computing

QuantumOS

Quantum compute applied where GPUs struggle — search and optimisation. Two aims: securing high-value substance IP in drugs and materials, and the research that carries us toward the AGI stage.

Even–Mansour period recovery to N = 10 on real IBM Heron hardware · drug-substance patent filed
Explore QuantumOS →
🤖03

Physical AI

On-device AI on our own LLM

Intelligence and judgement added to robots through an attachable on-device module built on our own models — no platform modification, no cloud dependency.

Boston Dynamics Spot understands spoken Korean and acts — running publicly at the Seoul Robot & AI Science Museum
See it in action →

The axes are not separate businesses. The foundation LLM gives Physical AI its on-device intelligence; QuantumOS post-processes what the models propose; the substance IP it secures funds the next model.

Four ways to make a model — we have all four

There are only so many ways to bring an AI model into existence: build it from nothing, merge existing ones, fuse different architectures, or compress a big one to run anywhere. VIDRAFT is one of the very few labs that has shipped a proven model with every single one — each backed by a number you can check.

CREATECOMBINEFUSEDEPLOY
🏗️01

From scratch

Pretrain a foundation model from zero — the hardest, most capital-intensive path.

AETHER · 144.2B tokens on 16×B200 · training data, code & logs fully open
Explore From scratch →
🧬02

Merge & evolve

Combine and evolve existing models with no training at all.

Darwin · 1M+ downloads · GPQA Diamond 90.9% · #1 on the K-AI leaderboard
Explore Merge & evolve →
⚗️03

Cross-breed

Fuse models from different architecture families — which normally cannot merge at all.

Chimera · compatibility-scored heterogeneous fusion · patented (29 claims)
Explore Cross-breed →
📱04

Compress

Shrink a large model to run on a GPU-less PC and a phone — with quality intact.

POCKET · a 35B model on a phone · measured to beat the most-downloaded on-device model
Explore Compress →

Latest Achievements

Key achievements by VIDraft in 2026.

🏆
2026.6 · NIPA-commissioned Project Review Ref. No. RQT-25-090162
"AI Computing Resource Infrastructure Enhancement (GPU Lease Support)" — Excellent
Commissioned by NIPA (National IT Industry Promotion Agency), administered by KAIT · GPU Lease Support Project rated Excellent.
Excellent
SemiEngineering · US2026.07
🌐 GLOBAL · EDITORIAL PICK

US chip-industry journal SemiEngineering lists our quantum paper in its weekly security-research review

In an editor-curated weekly review, our quantum cryptanalysis paper — run on real IBM quantum hardware — was listed next to papers from Meta, Google, Radboud and Politecnico. Not a press release we sent out, but an editorial selection: our first earned pickup by overseas trade press.

Read →
Partnership2026.07
🏭 IN PROGRESS

Model Foundry platform with Answerwise — under construction

A foundry that builds AI models to order, without pre-training. Applying the semiconductor foundry model to AI: the customer brings the design intent, the foundry runs the process. Platform construction is in progress.

Joint project · platform build in progress
Electronic Times2026.07.13
⭐ LATEST

VIDRAFT proposes 'VKAE×VKUE' set engines to maximize AI-IDC efficiency

Dual-engine strategy — VKAE (up to 23.4× throughput) + VKUE (35B frontier models on consumer GPU, notebook and CPU). CEO Kim Min-sik: grow into an infrastructure company that makes big AI cheaper, faster and closer.

Read →
Edaily2026.06.16
BREAKING

Darwin-398B-JGOS, GPQA Diamond 90.9%

180 of 198 correct. World-class performance through pure reasoning. MoE + Darwin V9 method.

Read →
Seoul Shinmun2026.06.09
BREAKING

JGOS-31B-Citizen, K-AI Leaderboard Overall #1

MSIT·NIA certified score 0.621, ranked #1. 8 of top 12 are Darwin derivatives (67%).

Read →
Tencent News2026.05.21
GLOBAL

Tencent (China’s #1 Portal) Spotlights Darwin

China’s largest portal features Darwin’s “genetic recombination” merging — smarter AI without retraining.

Read →

In the Press

VIDraft across Korean tech media & the founder's essays.

Journey to AGI

AGI is achieved step by step. VIDraft is now at the Pre-AGI stage.

1

Partial

Single-task specialization

2

Proto

Multi-task combination

3

Pre-AGI

Metacognition + Self-correction + Swarm collaboration

▶ Now
4

Pass

True AGI

5

Post

Post AGI

Open-source LLM Fully-Open Comparison

Sovereign fully-open (weights · data · training code · logs · checkpoints) — 6 countries side by side.

Open-source LLM fully-open comparison across 6 countries

AETHER 자세히 보기 →

DARWIN PLATFORM

Korea's #1 AI Platform Delivering Measurable Results

K-AI #1 · GPQA Diamond 90.9% · Polaris Global 14× Champion · Metacognition Leaderboard #1

📊 Leaderboard
🧬 Model Family
📚 Full Catalog →
🧠 Metacognition →

How Darwin merges without training

Merging pretrained models is cheap; merging them without losing what made each good is not. Three filings, 77 claims.

1

Merge each layer at its own trust level

20 claims

Blend two models at one uniform ratio and you mix one model’s strength with the other’s weakness. Quantifying each layer’s functional importance sets a different ratio per layer — attention and FFN are never merged on the same setting.

In plain terms: when combining two dishes you do not mix the broth and the solids in the same proportion.

2

Crossbreed across architecture families

29 claims

Models from different architecture families have mismatched tensor shapes and roles, so they cannot simply be merged. Scoring compatibility tensor by tensor separates what can be crossbred from what can only be transplanted — with no retraining.

In plain terms: like an organ transplant — you test compatibility first, then decide what can be grafted and what merely attached.

3

Accumulate across generations

28 claims

A single merge stops where it stops. Feeding children back as parents accumulates gains across generations, and re-injecting parent tensors breaks the model out of evolutionary stagnation.

In plain terms: keep breeding across generations — and bring in outside stock when the line starts to narrow.

90.9%
Darwin-398B-JGOS · GPQA Diamond
#1
K-AI Leaderboard Overall #1
JGOS-31B-Citizen 0.621
14×
Polaris Global Drug Discovery
ADMET + Efficacy + Kinase

GPQA Diamond Global Leaderboard

HuggingFace certified · As of June 2026

RankModelScoreNote
Darwin-398B-JGOS90.9%VIDraft Latest
1Darwin-28B-Opus88.89%VIDraft
2Qwen3.5-397B-A17B88.40%Alibaba
3Kimi-K2.587.60%Moonshot
5Darwin-36B-Opus88.40%VIDraft
11Darwin-27B-Opus86.90%VIDraft
16Darwin-31B-Opus85.90%VIDraft
21Darwin-9B-NEG84.34%VIDraft

K-AI Leaderboard

MSIT·NIA certified · Blind evaluation · As of June 2026

RankModelScoreCategory
1JGOS-31B-Citizen0.621VIDraft Official
2AWAXIS-KR-31B-v5Darwin Derivative
6Rogue-28B-MIXDarwin Derivative
7AWAXIS-Hybrid-28BDarwin Derivative
8Warecube-KO-27B-v3Darwin Derivative
9Rogue-27B-KRDarwin Derivative
11AWAXIS-Think-28BDarwin Derivative
12TenOS-Ko-28B-v3Darwin Derivative

📊 11 of the K-AI Top 20 are Darwin models

A result that emerged naturally in the MSIT·NIA blind evaluation. Third-party certified figures, not self-reported.

Three Generations of Model Evolution

Not a one-off merge — a parent → child → grandchild structure across successive generations.

Gen-1

Darwin-27B-Opus

GPQA 86.9%

First evolution from base models

Gen-2

Darwin-4B-David

GPQA 85%

An 85% GPQA score at 4B parameters

Gen-3

Darwin-4B-Genesis

CLIcK 92%

Cross-architecture breeding (Mamba)

Darwin Official Model Family

20+ official · 1,200+ derivatives & quant variants · 1M+ all-time downloads (HF-verified 2026-07)

💻
🔥 NEW · Coding AI

Darwin-28B-Coder-GGUF

GGUF 양자화 코딩 특화 모델. 최근 30일 22,537 다운로드로 다윈 패밀리 GGUF 최다 — 노트북·로컬에서 바로 구동. (base: Darwin-28B-Coder)

22,537
30-day downloads
★ Most Downloaded

Darwin-9B-NEG ⬇️ 575K

9B NEG single-pass evolution. Most-downloaded model in the Darwin family — 575K+ all-time.

GPQA 88.4%

Darwin-36B-Opus ⬇️ 102K

GPQA Diamond 88.4%. Community-favorite flagship, widely quantized and redistributed.

MoE Multimodal

Darwin-35B-A3B-Opus ⬇️ 71.2K

35B MoE flagship. GPQA Diamond 90.0%. Multimodal support.

KR Flagship

Darwin-31B-Opus ⬇️ 34.3K

Korean flagship. CLIcK 84.5 · KMMLU 76. Top Korean-language Darwin.

Coding

Darwin-28B-Coder ⬇️ 32.2K

Coding specialist. Top official GGUF download in the Darwin family.

Cross-Arch

Darwin-4B-Genesis ⬇️ 21.7K

Transformer × Mamba cross-breeding. CLIcK 92%. World-first heterogeneous crossbreed verified.

Global GPQA #1

Darwin-28B-Opus ⬇️ 14.3K

GPQA Diamond 88.89%. Surpasses 400B class with no extra training. arXiv 2605.14386.

Reasoning

Darwin-28B-REASON ⬇️ 13.3K

Reasoning-enhanced specialist. Optimized for math·code·science reasoning.

Gen-2 Evolution

Darwin-4B-David ⬇️ 10.9K

31B-class at 4.5B. GPQA 85.0%. Cumulative +26.4%p. Fit for embodied AI & mobile.

K-AI Overall #1

JGOS-31B-Citizen ⬇️ 6.0K

Public administration Korean. CLIcK 0.987 · KMMLU-Pro 0.725 · Com2 0.742.

KR Legal

Darwin-28B-KR-Legal ⬇️ 2.6K

27B BF16 legal specialist. B200 load 9.7s. TPS 22.3 tok/s.

TTS Voice

Darwin-TTS-1.7B-Cross ⬇️ 809

1.7B Korean speech synthesis. Cross-architecture evolution.

★ GPQA 90.9%

Darwin-398B-JGOS ⬇️ 404

MoE structure, 17B active params. Darwin V9 method. World-class pure reasoning.

⬇️ = HuggingFace all-time downloads (official + community redistributions), verified 2026-07. Darwin family total: 1M+.

🧬 Darwin Family Collection

Browse and download all 20 models in the HuggingFace FINAL-Bench collection.

View Darwin Family Collection →
AETHER ARCHITECTURE

End of single-attention era.
Dawn of 5/7/11-way Hybrid.

Adaptive Multi-way Triple-symmetric Hybrid attEntion Routing

Eight patents, one stack

Putting several attention mechanisms in one model is not the hard part. Making it actually train is. Placement skews, gradients diverge per type, causality leaks, experts collapse. These eight filings block those failures in order — 224 claims across the stack.

1

Placement — Latin square

30 claims

Mix attention types carelessly and you cannot attribute a good result — was it the composition, or where each type happened to land? Fixing each type to appear exactly once per row and column makes the layout deterministic and reproducible.

In plain terms: spreading tasting booths evenly across every floor, so you can say the menu sold well — not just the third floor.

2

Fusion — single-gate convex combination

26 claims

Summing outputs from different attentions lets the loudest one dominate. A single learnable gate fuses them — and the gate itself is regularised so it cannot collapse onto one mechanism.

In plain terms: one mixing fader, with a lock that stops any single channel from going to 100%.

3

Safety I — automatic causal-leakage detection

29 claims

When future tokens leak into past computation, training looks great and deployment fails. Previously this meant auditing layers by hand. Here the leak is located automatically and the offending module is swapped for a causally safe one.

In plain terms: catching a student peeking at the answer page — and replacing the method, not just the score.

4

Safety II — causal-safety CI/CD

27 claims

Fixing it once is not enough — the next commit can reintroduce it. Any model or code change triggers re-verification, and a failure blocks the merge.

In plain terms: safety enforced by the pipeline, not by someone remembering to check.

5

Stability — per-type adaptive clipping

28 claims

Each attention type carries different gradient statistics; one global clip destabilises training. Measuring each type separately catches anomalies before divergence, not after.

In plain terms: you do not load the same weight onto athletes of different size.

6

Routing — prime-based symmetric MoE

26 claims

MoE routinely collapses onto a few experts while the rest go unused. Instead of coaxing balance with auxiliary losses, a prime-based symmetric layout secures it structurally.

In plain terms: rather than nagging people to take turns, you write a rota that turns by construction.

7

Reasoning — metacognitive emergence

28 claims

Different inputs demand different kinds of reasoning, yet experts are usually fixed. Here the model judges what reasoning an input needs, selects accordingly, and can grow new reasoning types during inference.

In plain terms: sorting the question first — "this is arithmetic, this is common sense" — then calling the right person.

8

Transfer — cyclic-group warm transfer

30 claims

Moving to a model of different shape usually means training from scratch. Cyclic-group structure decides which layers correspond, and weights move only within matching attention types and functional groups.

In plain terms: when you move house, kitchen things go to the kitchen and books to the study — you do not tip every box onto the floor.

🌍 Fully-Open Sovereign AI Apache-2.0 From-Scratch

오픈 웨이트 ≠ 오픈소스 — AETHER는 전부 공개한 소버린 파운데이션 모델

가중치만 던지는 것은 컴파일된 바이너리를 주는 것과 같다. Aether-7B-5Attn은 데이터 레시피·토크나이저·학습 코드·모든 하이퍼파라미터·전체 로그·평가 코드·중간 체크포인트까지 전부 Apache-2.0으로 공개했다. 성능 우위를 주장하려는 것이 아니라, 처음부터 다시 만들 수 있는 '진짜 오픈'과 주권을 증명하려는 것이다.

완전공개 6개국 비교

'완전한 오픈'은 지금까지 국가연구소·정부 컨소시엄·대학 연합의 몫이었다. Aether는 단일 스타트업으로 그 명단에 이름을 올렸고, 유일하게 이종 어텐션(5종)을 얹었다.

Open-source LLM fully-open comparison across 6 countries

O 공개 · △// 부분·제약 · X 없음·미정  |  Aether = 6.59B MoE · 어텐션 5종을 7×7 라틴방진에 배치

7×7 라틴스퀘어 이종 어텐션 MoE

49개 층(7×7)에 어텐션 라벨을 라틴 스퀘어로 배치하면, 각 메커니즘이 모든 깊이에 정확히 한 번씩 나타난다. 깊이 편향을 제거한 실험 통제 장치다.

49 layers · 7×7 Latin square
nsaL0diffL1fullL2linL3sldL4cmpL5hybL6diffL7fullL8linL9sldL10cmpL11hybL12nsaL13fullL14linL15sldL16cmpL17hybL18nsaL19diffL20linL21sldL22cmpL23hybL24nsaL25diffL26fullL27sldL28cmpL29hybL30nsaL31diffL32fullL33linL34cmpL35hybL36nsaL37diffL38fullL39linL40sldL41hybL42nsaL43diffL44fullL45linL46sldL47cmpL48
Aether-7B-5Attn
Heterogeneous-Attention MoE
  • • 6.59B total · ~2.98B active (MoE)
  • • 25 experts · top-7 + 1 shared
  • 5종 구별 어텐션 메커니즘
  • • 144.2B tokens · 16× B200
  • 완전공개: 가중치·데이터·코드·로그
nsadifferentialfulllinearslidingcompresshybrid

※ compress·linear = full 계열 → 5종 구별 메커니즘. NSA 분기는 패딩 마스크를 소비하지 않으므로 추론은 batch_size=1.

World First
5/7/11-way heterogeneous attention integrated into a single LLM
317
Patent claims (KR priority + PCT)
Triple Symmetry expressiveness gain

Released Models

Published on Hugging Face — 3 base models + 1 instruct variant.

🗂️ Intermediate checkpoints

Steps 110k / 115k / 162k — released for reproduction.

🚀 Live demo

Aether Sovereign AI — try the model in your browser.

📚 Collection

All Aether models in one place.

Next Models

Planned for release. Specifications are subject to change.

🗓️ Coming

Aether-34B-4Attn

4-way attention · 34B class

🗓️ Coming

Aether-70B-4Attn

4-way attention · 70B class

🗓️ Coming

Aether-34B-A10B

MoE · 34B total / 10B active class

🗓️ Coming

Aether-460B-A30B

MoE · 460B total / 30B active class

Beyond the Single-Attention Era

Most production models run one attention mechanism in every layer. That is a convention, not a proven optimum.

ModelAttention compositionStructural limitation
Llama-3Full · 1 type80 layers · rapid KV-cache growth
MistralSliding window · 1 typeLimited long-range dependency
Mamba2State space · 1 typeNo softmax · accuracy trade-off
DeepSeek-V3MLA + Full · 2 typesStrong, but limited attention diversity
AETHERMultiple types on an n×n Latin squarePlacement is an experimental control — superiority still to be established by ablation

11-way Heterogeneous Attention Ensemble

Integrating the latest research into a single LLM.

⭐ ACL'25

NSA

DeepSeek

⭐ ICLR'25

Differential

Microsoft

Standard

Full Attention

Vaswani 2017

Efficient

MLA

DeepSeek-V2

Length

Sliding Window

Longformer

Compress

Compress

Native Sparse Attention · Sequence compression

Adaptive

Hybrid Gate

Differential Attention · Adaptive noise cancellation

Linear

Linear

ICML 2020

SSM

Mamba2

Gu & Dao 2024

Gate

GDN

Delta Network

High-eff.

Hyena

Poli ICML 2023

Metacognition Research

Teaching a model to sense when it might be wrong — a research effort to reduce hallucination.

🔬 Research stage

Self-doubt as a signal

A model that recognizes "I might be wrong" and re-checks. Our goal: make honesty and self-correction measurable, not assumed.

Early result

Fewer hallucinations

On an internal Korean hallucination test, a 5-way ensemble raised the rate of catching its own errors from 33% to 67% — a research result, not a shipped product.

FINAL Bench

Measuring metacognition

We built FINAL Bench, a benchmark that scores whether a model knows what it does not know — because you cannot improve what you cannot measure.

Development Roadmap

Stage 1 — Aether-7B-5Attn (Released · open source)

49 layers (7×7 Latin square) · weights, data recipe, training code, logs and checkpoints all published · Apache-2.0

Stage 2 — Aether-7B-5Attn-it (Released · instruct)

Instruction-tuned variant of the open-source base

3

Stage 3 — AETHER-7B-7Attn-base · Aether-6B-11Attn-base (Released · open weights)

7-way and 11-way heterogeneous stacks — 11Attn is a mid-training checkpoint (step 174,000)

4

Stage 4 — Aether-34B-4Attn · Aether-70B-4Attn · Aether-34B-A10B · Aether-460B-A30B (Planned)

Scale-up line. Specifications are subject to change.

Design Principles

01

Emergence

9 emergence engines + MARL runtime

02

Metacognition

8-type self-error detection · TICOS

03

Self-evolution

Closed-loop perpetual learning · SLAI

04

Swarm Collaboration

Multi-agent · SOMA/MOUSE

05

Mutual Reinforcement

Five-element enhance/suppress · Hallucination control

Metacognition Leaderboard

JinigenAI · VIDraft joint development · 24 AI models evaluated · Published 2026.07.01

🧠 "AI is dangerous not when it makes an error, but when it doesn't know it has."

JGOS-31B-Citizen avoids 398 of 400 trap questions (99.5%, #1). Metacognition = an essential AGI capability.

🎯
World First

FINAL Bench

World's first metacognition benchmark. 100 tasks × 15 domains × 8 types (TICOS) × 3 difficulty levels.

📊
400 Questions

Trap-Question Evaluation

400 questions designed to induce plausible-but-wrong answers. Measures self-error recognition.

🔌
Non-invasive

Metacognition Adapter

Analyzes internal signals to detect self-errors in advance without altering the original model's structure or weights.

📱 VIDRAFT On-Device · POCKET

35B in your pocket — no GPU, no cloud

An open-source on-device model family that runs a 35B model on a GPU-less PC, and on a phone, with stock llama.cpp — no fork. Measured against Bonsai on the same machine, we published where we win and where we lose.

POCKET vs Bonsai — measured, not marketed

A Korean startup's on-device model, benchmarked head-to-head against the world's most-downloaded one — on the same machine, with the same stock llama.cpp. We show where we win, and the one row where we lose.

🖥️

Reproduce it live on a free CPU

3.4× CPU

On Hugging Face’s free CPU Space, POCKET and Bonsai answer side by side on the same 8-vCPU box — and each shows its tok/s in real time. POCKET generates ~20.5 tok/s vs Bonsai ~6.1 — about 3.4× faster (CPU-only). Ask anything and watch.

⚔️ Open the live A/B demo →
POCKET-35B VIDRAFT · KR

35B total · ~3B active per token · sparse MoE

Bonsai-27B

27B dense · reads every weight per token

most-downloaded on-device model (2M+)

How much faster POCKET generates

Bonsai = 1×★ HF CPU Space · 8 vCPU3.40× fasterMacBook · prompt3.28× fasterLaptop CPU · generate3.13× fasterDesktop CPU · generate2.69× fasterServer GPU · generate2.22× fasterMacBook · generate1.99× faster
METRICPOCKETBonsaiRESULT
Desktop / Server (Linux)
HF CPU Space — 8 vCPU (reproducible)~20.5~6.13.4× faster
CPU generate — Xeon, 16 threads27.010.12.69× faster
GPU generate — H100197892.22× faster
GPU prompt — H1007531816Bonsai wins
Quality — HellaSwag, 400q61.0%60.0%= tie
Laptop — MacBook M3 Pro, 18 GB
Metal generate25.412.81.99× faster
CPU generate — 8 threads13.84.43.13× faster
Metal prompt240.773.43.28× faster

Why a bigger model runs faster

Bonsai (dense 27B)3.5 GB / tokenPOCKET (sparse MoE)0.66 GB / tokenGeneration is memory-bandwidth-bound — POCKET reads ~1/5 the memory each token.

Same machine, same unmodified llama.cpp — anyone can reproduce it. POCKET needs no fork; the comparable Bonsai build will not even load in upstream llama.cpp.

Apache-2.0

Open weights, a recipe that is ours

POCKET is not just an open model re-saved in a smaller file. Its weights are Apache-2.0 — free for anyone — but the compression and optimization process that produces them is VIDRAFT’s own, and it is not published. What is open is the artifact; what is ours is how it is made.

🔓

The result is open

Weights, configs and usage are all Apache-2.0 — download, run, modify, redistribute. No strings.

🔒

The process is ours

The recipe that keeps a 35B model coherent at extreme small sizes is VIDRAFT’s proprietary technology — developed in-house, and not disclosed.

📊

The proof is measured

Coherent quality even at 8.2 GB, with dedicated builds for Korean and for English — exactly where a naive one-size quantization falls apart.

Why a bigger model runs faster

POCKET is a sparse Mixture-of-Experts model: many experts, but only a few activate per token. On-device decoding is bound by memory bandwidth — how much you read each step — so reading a fraction of the weights makes a larger model faster where there is no GPU. And unlike the comparable Bonsai build, which needs a vendor fork to load, POCKET runs on stock llama.cpp — so LM Studio, Ollama and the standard ecosystem work on day one.

Honest scope: server-GPU (H100) prefill goes to Bonsai; on laptops and phones POCKET wins prefill too. Quality is tied on HellaSwag — neither side claims a quality win. iPhone speed is not yet self-measured (community reports welcome). Real-world feel varies with RAM, storage, quantization and thermals.

Apache-2.0 · open-source on-device model family by VIDRAFT

R&D

Papers. Patents. Open Research.

What we published, and what we filed.

🏆 Benchmarks
📄 Papers (4)
🔒 Patents (12)
📰 2026.07.20 비드래프트, 완전공개형 파운데이션 모델 ‘Aether-7B-5Attn’ 공개 — “AI 주권 구현” 조선비즈 → 중앙일보 → 동아일보 →
📰 2026.07.13 비드래프트 ‘다윈’, 허깅페이스 누적 다운로드 100만 돌파 (103만+) 중앙일보 → 조선비즈 → 동아일보 →

Six Independent Validations

Each benchmark proves a different capability — and each was administered by an external party.

K-AI Leaderboard #1
JGOS-31B-Citizen

Korean & public-sector LLM capability

GPQA Diamond 90.9%
Darwin-398B-JGOS

PhD-level scientific reasoning

Polaris — 14 tracks
PharmaOS

Drug-candidate prediction & screening

FACTS Grounding #2
Validated by CNRS

Medical hallucination control

Fast Gemma Challenge
VKAE

Serving optimisation on constrained hardware

FINAL Bench #1
JGOS — 398/400 traps avoided

Metacognition & self-error detection

🏆
2026.6 · Commissioned by NIPA · administered by KAIT · Ref. No. RQT-25-090162
"AI Computing Resource Infrastructure Enhancement (GPU Lease Support)" — Excellent
NIPA-commissioned, KAIT-administered GPU Lease Support Project rated Excellent.
Excellent
🏆
Global Record

GPQA Diamond

Darwin-398B-JGOS 90.9%. Darwin-28B-Opus 88.89% #1. 5 Darwin models in Top 21. HuggingFace certified.

🇰🇷
MSIT·NIA Certified

K-AI Leaderboard

JGOS-31B-Citizen overall #1 (0.621). 11 Darwin derivatives in the Top 20. Blind evaluation.

🧠
World First

FINAL Bench — Metacognition

HF Global Top 5. 100 tasks × 15 domains × 8 types × 3 difficulty. TICOS quantitative measurement.

🌐
Sole Track C Pass

WM Bench

World model benchmark. Sole model to pass Track C. Real-time interactive world simulation.

💊
Polaris Hub · Global Drug-Discovery ML Benchmark
14 First-Place Finishes — ADMET · Potency · Kinase
14× #1

Structure-only prediction — no cell-imaging or external experimental data — topped 14 public Polaris leaderboards across absorption, distribution, metabolism, excretion, toxicity, potency and kinase selectivity.

All results are computational predictions on public blind-submission leaderboards, submitted under the VIDraft account. VIDraft has since filed its first drug-substance patent from this autonomous discovery system.

arXiv 2605.14386

Darwin Family

MRI-Trust-Weighted Evolutionary Merging for Training-Free Scaling. 45 days after release: 300K+ downloads, 500 derivatives.


View Paper →  ·  PDF →

Read the full paper

⬇ Download PDF
HF Global #5

FINAL Bench

Measuring Functional Metacognitive Reasoning in LLMs. MetaCog scaffold +14.05 pts. 94.8% Error Recovery. Public preprint on SSRN · under conference review.


View Paper (SSRN) →

Read the full paper

⬇ Download PDF
VIDRAFT QuantumOS · 7 authors

Quantum Cryptanalysis

🏅 美 반도체 권위지 SemiEngineering 편집진, 주간 보안 연구논문 리뷰에 선정 — Meta · Google · Radboud · Politecnico 연구와 나란히 (2026.07)

Extending Even–Mansour period recovery from N = 4 to N = 10 on real IBM quantum hardware (ibm_kingston, Heron). Five genuine attacks across four symmetric-cipher paradigms, validated to the 25-qubit classical-simulation ceiling.


View Paper (arXiv 2607.18340) → · 🤗 HF Papers → · Plain-language explainer →
Even-Mansour N=10
Even–Mansour key recovery on real hardware — public record N = 4 versus this work at N = 10.
rank vs birthday bound
True-key quantum rank tracks the birthday bound — orders of magnitude below random, but no exponential quantum advantage.

Read the full paper

⬇ Download PDF
If the preview does not load in your browser, use the download link above.

Scope stated explicitly: Q2 query model, reduced/structured constructions, follows the birthday bound — this is not quantum advantage, and does not break full AES/RSA or 16-round DES.

VIDRAFT AI Research · 7 authors

A Numerical Realization of Suzuki's Weil-Quadratic-Form Operator

The first numerical realization of Suzuki’s Hilbert–Pólya candidate operator: a universal Archimedean spectral law pinned to 30-digit precision, with a conductor-only offset, and an operator form of Weil’s positivity criterion. The authors state explicitly that this does not prove the Riemann Hypothesis, nor advance toward a proof.


View Paper (arXiv 2606.09096) →

Read the full paper

⬇ Download PDF

KR priority filing + PCT · 317 claims total

🧬 Darwin Evolutionary · 3 patents · 77 claims

D1

MRI-Based Adaptive Reliability Meta-learning

Neural-network parameter merging based on Model-layer Response Importance (MRI)

D2

Compatibility-Assessed Heterogeneous Crossbreeding

Compatibility-Based Crossbreeding for Heterogeneous Neural Networks

D3

Recursive Generational Inheritance Evolution

Recursive Generational Inheritance in Evolutionary Neural Network Merging

⚡ AETHER Architecture · 8 patents · 224 claims

P1

Latin-Square Hybrid Multi-Type Attention

5/7/11-way heterogeneous attention Latin Square integration

P2

Automatic Causal Safety Verification

Causal Safety Verification for Multi-Type Attention

P6

Metacognition Signal Head

Self-error detection signal head

P8

Perpetual Closed-Loop Learning

Single Sigmoid Gate Convex Combination Fusion

P0

Multi-Agent Swarm Coordination

Prime-Order Cyclic Group Weight Transfer

P3

Triple-Symmetric MoE Routing

Per-Attention-Type Gradient Clipping

P4

Adaptive Noise-Cancelling Attention

Prime-Based Mixture-of-Experts Routing

P7

Hallucination-Suppression Gate

Causal Safety CI/CD for Neural Network Models

💊 Drug Substance · 1 patent · 16 claims

M1

PSP Therapeutic Compound

N-(pyrrolidin-1-yl)-thiazol-4-yl boronic acid compound for progressive supranuclear palsy (PSP) — composition, preparation method and use

PHYSICAL AI

Intelligence, Attached.

On-device AI built on our own LLM — added to robots and physical systems without touching the platform.

What We Ship

A complete Physical AI kit — software and hardware — that attaches to a robot without any prior arrangement with its manufacturer.

Software

SW

VLM

Vision-language model built on our own LLM — perception and language in one stack.

SW

Harness

The control harness that binds model output to robot action and safety limits.

SW

VKAE · VKUE · VKIE

Our inference acceleration and lightweight engines — latency, low-end hardware, throughput.

Hardware set

HW

3D depth camera · LiDAR

HW

Speaker · microphone

HW

AP · battery

HW

Additional sensors

HW

Battery

9 hours of independent operation per charge

Third-party by design

We add intelligence and perception to a robot as a third party — no partnership with the manufacturer, no firmware fork, no change to the platform. The same kit that runs on a Boston Dynamics Spot can be fitted to other machines already in the field.

Why on-device

Three constraints that cloud AI cannot solve for physical systems.

Latency

A robot acting on a spoken command cannot wait on a round trip. Inference has to happen where the sensor is.

🔒

Isolation

Defence, industrial and public sites often run without outbound network. The model has to work offline.

🔧

Retrofit

Systems already deployed will not be rebuilt around AI. Intelligence has to arrive as an attachment.

A Foreign Robot That Learned Korean

Boston Dynamics builds the robot. We made it understand Korean — as a third party, without touching their platform.

01

Out of the box

Spot is a quadruped robot built by Boston Dynamics. As shipped, it does not understand Korean.

02

We did not modify the robot

No firmware fork, no vendor partnership, no changes to the platform. We attached an on-device AI module.

03

Korean in, action out

A visitor speaks Korean. The model on the device interprets intent and the robot moves accordingly.

04

Running in public

Deployed as a live exhibit at the Seoul Robot & AI Science Museum — operating for general visitors, not a lab demo.

Why this matters commercially

Most physical systems already deployed in the field will never be rebuilt around AI. If intelligence can be added as an attachable on-device module, the installed base becomes addressable — without the manufacturer's involvement. That is the on-device AI supply model behind our Physical AI line.

What this video shows

A visitor speaks Korean. Our on-device model — mounted on Spot and running on the robot itself with no cloud connection — interprets the intent, and Spot acts on it. Nothing in the robot was modified; the intelligence arrived as an attachment. The whole proposition is in this one clip.

Spot responding to spoken Korean — Seoul Robot & AI Science Museum.
Boston Dynamics · quadruped robot

Spot

Net weight33.8 kg
Max payload14 kg
Max speed1.6 m/s
Max slope±30°
Step height300 mm
Runtime90 min (564 Wh)
Ingress / tempIP54 · −20 ~ 55 °C
Terrain sensing360° · 4 m range

Specifications from the manufacturer. Our kit mounts within the 14 kg payload allowance — no modification to the robot itself.

Live Deployments — Seoul AI & Robot Science Museum

Three exhibits built by VIDraft, running for the public.

🤖
Physical AI

Talk to Spot

Boston Dynamics 'Spot' controlled by natural-language voice — converse with, command and steer the robot through an LLM.

🩻
Interpretability

Model X-RAY

See inside the LLM black box — an X-RAY view that makes a model's internal workings transparent to visitors.

🧬
Evolution

Darwin Crossbreed & Evolve

Watch AI models crossbreed and evolve — the child surpasses both parents. An interactive Darwin evolution experience.

🌱 COMPANY

Intelligence is not made — it is grown.

A small Korean Proto-AGI startup (since 2024) with one belief: don't exaggerate — speak only what is verified.

A Word from the CEO

Min-sik Kim, CEO of VIDRAFT
“I paused school. It was to build a scientific research company — quickly.”
Min-sik Kim · CEO, VIDRAFT

I am Min-sik Kim, CEO of VIDRAFT. I am twenty-eight — a fourth-year computer engineering student at Korea Cyber University, currently on leave. I paused school to build this company, so my credentials will not tell you much.

So let me tell you what I have done instead. I am named inventor on 12 patent applications and co-author on 4 papers. All of that intellectual property belongs to VIDRAFT, by design — it is the company’s foundation, not my résumé.

The AGI axis holds Darwin, Chimera and AETHER. Darwin diagnoses a trained model (Model MRI) and breeds it across generations — three patents: reliability-weighted merging, cross-architecture breeding, recursive generational inheritance. Chimera decomposes a model’s knowledge structure and transplants it, which led to our own architecture, AETHER — eight patents covering heterogeneous attention on a Latin square and metacognitive verification. Eleven of our AI patents sit on this axis.

On quantum computing, I am a co-author on the paper that extended Even–Mansour period recovery on real IBM Heron hardware from the public record of N = 4 to N = 10. On the same axis we filed a substance patent for a progressive supranuclear palsy therapeutic compound — we search with quantum, and capture the value as substance IP.

On Physical AI, we attached an on-device model built on our own LLM to a Boston Dynamics Spot so that it understands and acts on spoken Korean. It runs today at the Seoul Robot & AI Science Museum.

I serve as principal investigator on two government programmes — the AI computing-resource utilisation project, completed and rated Excellent, and an ongoing advanced-GPU support programme.

Before that, I volunteered for signals duty and completed my military service. What I took from it was not how to work the equipment, but a plainer fact: once a connection drops, nothing else functions. The soundest judgement counts for nothing if it never reaches anyone.

And before I was someone who reads research, I was someone who gathered people to read it together. Some 15,000 people are active members of the communities I run — ArXivGPT, OpenFree AI and others — where we share new work every day. Much of VIDRAFT’s ability to absorb research quickly comes from there.

My ambition

I want to win a Nobel Prize in the sciences — Physics, Chemistry, or Physiology or Medicine. Those are precisely the three domains we work in: quantum for physics, new materials for chemistry, drug discovery for medicine. That alignment is deliberate, not a coincidence.

And one ambition that should outlast me: a Max Planck Society for Korea. Germany’s Max Planck spent a century funding basic research that paid nothing in the short term, and produced more than thirty Nobel laureates. Korea does not yet have a vessel like that. I want this company to become one — earning from products, spending on fundamental research, and letting that research become the next product. That is why our business rests on two axes: solutions and services, and substance IP.

That is a large claim, which is exactly why we say only what has been verified. Please check our published models, papers and patents yourself. Exaggeration gets found out; the record remains.

Corporate Research Institute — officially recognised
Recognised by KOITA under authority delegated by the Ministry of Science and ICT. Recognition No. 2026111898 · recognised 27 May 2026 · research field: artificial-intelligence software.

Who We Are

VIDRAFT Inc. is a Seoul-based deep-tech company that, built on Pre-AGI-level AI and quantum computing, extends its research and development into physics, chemistry, life sciences, and pharmaceuticals.

🧠
Foundation

Pre-AGI AI

The Darwin model family and the AETHER architecture — intelligence grown to the Pre-AGI stage.

⚛️
Foundation

Quantum computing

QuantumOS — verifying quantum cryptanalysis step by step on real hardware, tilling the soil for the next era.

🔬
R&D reach

Physics · Chemistry

New-material and molecular research — checked first against novelty and prior-art gates.

🧬
R&D reach

Life sciences · Pharma

Drug-candidate and peptide discovery — first drug-substance patent filed in 2026.

Our Philosophy

Great life was never designed all at once. Seed → sprout → tree → forest. AGI is grown by nature's law, not built by force.

🌱
Grow, not make

기르다 · Cultivate

We do not build intelligence in one giant step; we grow it across generations, the way nature evolves great life. 씨앗 → 새싹 → 나무 → 숲.

🔍
Honesty & verification

정직과 검증

Our founding spirit: do not exaggerate; speak only what has been verified. Real trust comes from honestly stating what we have NOT done.

🧬
DARWIN

Grow by evolution

Crossbreed the strengths of different models to raise a stronger next generation — inheriting a lineage of intelligence, like a craft passed down through generations.

☯️
AETHER

Harmony of diversity (조화)

Not one method — 5·7·11 kinds of attention coexist. When diversity harmonizes, higher intelligence awakens. This is how we look beyond the Transformer.

⚛️
QuantumOS

Soil for the next era (길)

We do not claim to have 'completed' the quantum computer. We till the soil for the next era, verifying step by step on real hardware. AGI is a road, not a destination.

🌏
Tech for people

Philosophy on the ground

Technology for people and society — the same dream as Society 5.0. From Jeonnam·Gwangju (3.2M citizens) to across borders.

Major Milestones

The Crew

Eleven people, armed with freedom, flexibility, speed and creativity — like romantic pirates.

🏴‍☠️ "Join My Crew?" — we call ourselves 조마크.

Ask any VIDraft member their identity and we answer 'Jo-mark' — the crew's callsign, drawn from the manga One Piece. Not resources but wits; not volume but grit and agility.

Leadership

CEO
CEO
김 루피

낙천이 전략이고, 확신이 자산이다.

COO
COO
김 징베

회장이 소리쳐도, 이 배의 항로는 흔들리지 않는다.

CFO
CFO
정 로우

숫자는 냉정하게, 팀은 진심으로.

CIO
CIO
한 라일리

이 바닥의 모든 항로를 지도 없이 안다.

CMO
CMO
이 이바코프

명함 한 장에 왕국이 딸려온다.

Japan Head
일본 사업 본부장 · JAPAN HEAD
조 오뎅

일본과 세계를 잇는 유일한 사무라이.

Chairman
의장
김 베가펑크

우리가 지금 만드는 것들이 인류의 미래가 될 것이다.

Research Institute

Corporate Research Institute, recognised by KOITA (No. 2026111898).

Head of R&D
연구소장 · HEAD OF R&D
홍 크로커스

말없이 자리를 지키고, 결과로 답한다.

CTO
CTO
장 조로

CEO가 자면 자고, CEO가 코딩하면 코딩한다.

Intern
인턴 · INTERN
김 비비

정식 크루가 되는 그날까지, 배우겠습니다.

Intern
인턴 · INTERN
신 쵸파

칭찬하지 마세요! (매우 기뻐하는 중)

Developer
DEVELOPER · ON MILITARY LEAVE
김 로빈

For now I keep ciphers. When I return, I break them.

👥

A crew of 12

A tight team that beats volume with method — 24 GPUs against the giants.

🧭

Vertically integrated

Own LLMs · City OS · drug/material AI · on-prem AI — CEO, Vice President and Japan Business Head lead a Proto-AGI stack.

🚀

Join My Crew

Looking for the next global AI pirate — freedom, speed, creativity. Contact us.

Business Model

Three revenue axes: solutions and services delivered on-premise, substance IP generated by our domain engines, and core technology licensed to partners.

🏛️

Solutions & Services

Domain OS delivered as on-premise deployments, PoCs and managed services. Revenue from licences, deployment and operations.

🧪

Substance IP

PharmaOS and MaterialOS generate drug and material candidates that become filed intellectual property — an asset that compounds independently of service revenue.

🏭

Technology IP Licensing

Core architecture and process technology licensed to partners on a non-exclusive basis. The partner runs the platform and owns the customer relationship; we supply the process. First case: the 'Model Foundry' platform being built with Answerwise Inc. (under construction)

What We Deliver Today

These are existing revenue lines, not roadmap items.

Delivering
🏢

On-premise AX (LLM included)

Full AI transformation delivered inside the customer network — model, serving stack and workflow. Data never leaves the perimeter.

Delivering
🤖

On-device AI for Physical AI

On-device vision-language models for robots and physical systems. Reference deployment: Seoul Robot & AI Science Museum.

In progress
🏛️

Domain OS deployment & operations

NationalOS / JGOS public-administration engine, plus PharmaOS and MaterialsOS domain builds and managed operations.

Foundation → Domain OS → Solution

01
Foundation technology
AETHER architecture · Darwin Factory · VK Engine
02
Domain OS
NationalOS · PharmaOS · MaterialsOS · QuantumOS
03
Enterprise solution
On-premise deployment · PoC · managed operations

Why on-premise: public-sector, defence, healthcare and financial data often cannot leave a secure network — a market general-purpose cloud AI is not well suited to serve.

Five Verticals

Market entry through domain-specific OS. Counterparties in ongoing negotiations are not disclosed.

In progress

Public administration

Regional public-administration LLM (JGOS) — development and PoC

Planned

Medical · Bio

Drug-discovery biofoundry — lead role planned

In progress

Logistics

Logistics-specialised AI stack built on partner-provided operational data

In discussion

AI infrastructure

VK Engine-based infrastructure business — domestic and overseas

In discussion

Advanced materials

MaterialsOS — collaboration with a state-backed materials body

Growth Roadmap

2026 → 2028 · market validation, vertical OS commercialisation, multi-domain Pre-AGI platform.

H2 2026 — Proof Expansion · Paid Pilots

NationalOS/JGOS public-administration PoC · PharmaOS·MaterialOS domain validation · VK Engine on-prem packaging

2

2027 — AETHER Commercialisation · Domain OS

AETHER V1 commercial release · public/medical/materials OS commercialisation · Series A and partner expansion

3

2028 — Domain Pre-AGI Platform

Global enterprise entry on proprietary architecture · OS commercialisation scale-up · PCT filings and global IP strategy

🏛️ SOLUTIONS · AX

Sovereign AI OS — your data never leaves the network

Not one chatbot, but an AI operating system that connects administration, citizen and industry data safely inside a local / independent network.

Public · Local-Government AX

Administration overload · fragmented data · cloud-AI security concerns → a local sovereign AI operating system.

🧠
Darwin LLM

Region-tuned LLM

Language · administration · industry models tuned for each region.

🛡️
AETHER Trust

Grounded & self-verifying

Hallucination suppression · self-verification · evidence-backed answers.

🗺️
NationalOS

Policy simulation

Policy · budget · industry simulation on regional digital twins.

🚪
CHITOS Portal

One common AI touchpoint

A shared AI access point for citizens, officials and enterprises.

🔬
Domain OS

Pharma · Material · Nutri · Legal

Vertical-domain AI plugged into the regional OS.

🔒
On-prem · Sovereign

Data stays inside

On-premise / independent-network deployment, verified by a 3-month KPI pilot.

🏗️
Building

Building — Architectural Design AI

From natural language to KS-standard floor plans (DXF·BIM·BOQ·fire-safety), and from a parcel address to code-compliant 3D building massing.

5 AX Packages
01

Administration

02

Citizen

03

Industry

04

Bio

05

Materials

⚡ VIDRAFT Inference Engines · VKAE · VKUE · VKIE

Fast, Everywhere, and at Scale — one 34.7B model, all hardware

Three sibling engines: VKAE runs it fast (≈9× single-GPU), VKUE runs it anywhere (down to a free CPU), and VKIE serves the most (18,057 tok/s on one B200). Every number measured.

🚄 VKIE · Integrated
VKAE · Acceleration
VKUE · Ubiquitous
23.4×
throughput vs. standard (NVIDIA B200)
10K+
tokens/sec under concurrent load
0
quality loss — same answers, faster
100%
reproducible — one Docker, your GPU

Why inference acceleration is essential

The real AI-infra battleground is not building the model — it is running it, cheaply, forever.

💰
Efficiency = capacity

A virtual extra GPU

GPUs are scarce, expensive and power-hungry. Doubling throughput on the same card is like building twice the infrastructure — no new chips, no new power.

📉
Cost per token

Where AI services win or lose

Training happens once; inference happens every time a user asks. VKAE lowers the per-token cost that decides whether an AI service is profitable.

🌐
Sovereign-AI ready

Not a nice-to-have

For clouds, AI startups, enterprises and national/sovereign AI, acceleration is not optional — without it, running AI at scale runs at a loss.

🔬 Reproducible, not just a claim

Speed claims are cheap. VKAE ships the model weights plus the optimized serving stack as one Docker container, OpenAI-compatible — so anyone can reproduce the numbers on their own GPU and plug it straight into an existing service.

VKUE — No GPU? Runs anyway. Everywhere.

The ubiquity engine that serves frontier-class models on minimal hardware. VKAE runs fast; VKUE runs everywhere.

Frontier-class on minimal hardware

Run GPQA 86–88% models on a single 24GB consumer GPU — or on CPU with no GPU. (Scores use different methods; compare with care.)

Model GPQA Diamond (method) VKUE min. HW
Darwin-398B-JGOS90.9 (greedy)1 GPU + RAM
Darwin-36B-Opus88.4 (maj@8)1× 24GB consumer GPU / CPU
Ourbox-35B-JGOS86.4 (maj@8) · 70.7 (greedy)1× 24GB consumer GPU / CPU
GLM-5.2 (744B)frontier-class1× H100 · speed under verification

VKIE · 비키 — 통합 추론 엔진

VKAE는 빠르게 · VKUE는 어디서나 · VKIE는 가장 많이. 34.7B 모델 하나로 데이터센터 GPU부터 무료 CPU까지 — 모든 수치 실측, 공개 라이브 데모에서 재현 가능.

세 개의 축, 하나의 엔진

스포츠카는 가장 빠르고, 경차는 가장 싸고, 기차는 가장 많이 실어나른다. VKIE는 VKAE의 속도와 VKUE의 절감을 동시에 취해 서빙 능력을 극대화한다.

🏎️
VKAE

≈ 9×

GPU 가속 · 스포츠카. 단일 스트림 1×B200에서 24 → 220 tok/s. 최고 단일 속도.

🚗
VKUE

무료 CPU

GPU 절감 · 경차. 34.7B를 GPU 0으로 ~6–7 tok/s 구동. 최저 비용·최대 접근성.

🚄
VKIE

18,057 tok/s

가속+절감 · 기차. 1×B200 동접(256) 서빙 능력(집계). 최대 처리량·원가효율.

어떤 엔진을 언제 쓰나

세 엔진은 독립적으로도, 결합해서도 씁니다. 배포 환경에 따라 최적 조합이 달라집니다.

🏎️ VKAE🚗 VKUE🚄 VKIE
최적화 차원지연 (속도)비용 · 편재처리량 (규모)
핵심 효과최속 단일 생성저사양 구동최대 동시접속
대표 시나리오실시간 대화 · 코딩엣지 · 온프레미스대규모 API · SaaS
대표 실측506.94 tok/s35B @ A10G·CPU18,057 tok/s

실측 성능 — 적용 전 → 적용 후

기준 모델 Ourbox-35B (총 34.7B / 활성 ~3B 희소 MoE). 모델 변경 없음, 품질 손상 없음.

Engine Before After Gain Setup
🏎️ VKAE24220≈ 9×1×B200 · single stream
🚗 VKUE5.420.0≈ 3.7×8GB laptop (RTX 5060) · dense→sparse
🚄 VKIE2418,057≈ 750×1×B200 · 256 concurrent (aggregate)
🏆 Google 공개 주관 Gemma 속도 챌린지 1위 — 506.94 tok/s

순간 노이즈 스파이크를 배제한 재현 가능한 유효 최고 속도 기준 기록입니다. 외부가 주관한 공개 경쟁에서 검증된 객관 지표입니다.

한 모델, 전 하드웨어 스펙트럼

Ourbox-35B — 같은 가중치, 하드웨어만 바뀐다. 데이터센터 GPU에서 무료 CPU까지 하나의 파일로.

Hardware tok/s Axis
1× B200 (datacenter)18,057VKIE · aggregate (256)
1× A10G (cloud GPU)126VKUE · single stream
8GB laptop (RTX 5060)20.0VKUE · 3.7× vs dense 32B
CPU-Upgrade (8 vCPU, no GPU)~17VKUE
🆓 Free CPU Space (2 vCPU, no GPU)~6–7VKUE · zero cost
1× A100 · 96 동접~460VKIE · 단일 대비 3.6× (집계)

※ 품질은 전 티어 동일 — GPQA Diamond 86.4% (Ourbox-35B, maj@8) ~ 90.9% (Darwin-398B). '빠르고 싸다 ≠ 멍청하다.' 멀티모달 확장: Janus-Pro-1B 이미지 생성 저비용 T4에서 1장 ~28초.

왜 중요한가 — 값싼 하드웨어에 프론티어 AI

🏛️

주권형 AI

데이터를 클라우드에 못 올리는 공공·국방·의료·금융. 폐쇄망 온프렘 CPU에서 프론티어급 추론.

💰

비용 붕괴

수억 원대 GPU 클러스터 → 카드 한 장, 나아가 무료 CPU. 진입 비용 두 자릿수 배 절감.

🌍

접근성

개인·스타트업·중소·공공 누구나. 온디바이스·엣지 수요 급증에 그대로 대응.

🔓 공개 · 재현 가능

양자화 모델(GGUF)·전 벤치 수치·baseline 실행법 공개. 누구나 라이브 데모에서 '적용 전' 수치를 재현할 수 있다.

🔒 비공개 (영업기밀)

최적화 서빙의 엔진 내부는 비공개. 결과 속도·하드웨어만 공개 — 결과는 열려 있고, 레시피는 닫혀 있다.

🛡️ CHITOS · Autonomous Security AI

From suspicion to proof

Most scanners flag a pattern and say "maybe a vuln — you check." Chitos verifies inside authorized targets, confirms reproducibility, and chains attack paths into one scenario.

Three Stages

Successor to VIDRAFT Mythos. Where Mythos was static analysis, Chitos verifies.

🔍
Stage 1 · Static Scan (free)

Multi-language pattern analysis

Python, JS, Go, Java, C/C++, Rust, PHP, YAML — injection sinks, deserialization, credential exposure, crypto misuse, prototype pollution, supply-chain risk. Unproven warnings are filtered out as noise.

🧠
Stage 2 · Threat Research

Autonomous context research

Cross-references CVE databases, exploit advisories and public PoCs to explain WHY a finding is dangerous in the current threat landscape. (Requires a Claude API key.)

🎯
Stage 3 · Active Verification (free)

Reproducible exploit proof

On authorized targets only — XSS, SQLi, path traversal, command injection. Finds reproducible evidence, retries variant paths when blocked, and links confirmed vulnerabilities into a kill-chain threat map.

Ready-to-run Scenarios

Five representative attack scenarios, one click each. Live-test against IBM intentionally-vulnerable demo.testfire.net.

CVE

Log4Shell

Auth

JWT alg:none bypass

RCE

SSTI → RCE

JS

Prototype Pollution

Supply chain

Dependency Attack

🔒 Privacy by design

No install. No signup. Your code is never sent to an external API.

⚖️ Authorized use only

Active exploit features may be used only on targets you own or are explicitly authorized to test. Unauthorized testing or scanning may be illegal; all legal responsibility rests with the user.

⚛️ QUANTUMOS

Cryptanalysis on Real Quantum Hardware

Simon-algorithm attacks executed on IBM's 156-qubit ibm_kingston — at record hardware scale.

The world stopped at "4". A small Korean team pushed to "10".

Drawing a rocket on paper and actually launching it are two different things. Real quantum computers are noisy — a slightly longer computation drowns the answer. For years the public record on real hardware stayed stuck around N≈4.

VIDraft touched that wall. Taking the two most vulnerable structures — Even–Mansour and Feistel — onto IBM's real quantum computer, the team pushed the simplest block-cipher structure to N = 5, 6, 7, 8, 9 and 10, recovering a fresh independent key each time to prove it was not a rigged demo tuned to one answer — using noise mitigation only, no costly error correction.

Real Hardware Results

IBM ibm_kingston (156 qubits) · 15–30 qubit circuits with hundreds of two-qubit gates.

🔑
Even–Mansour

N = 5 → 10

Secret-key recovery on real hardware from N=5 up to N=10 — exceeding the previous public record of N≈4 (to our knowledge).

🔀
3-round Feistel · DES-family

Block 6 · 8

Hidden-period recovery at block sizes 6 and 8, rank-1 clean.

Error mitigation only — dynamical decoupling, gate/measurement twirling, readout-error correction (no quantum error correction). Independent control keys were tested per instance to confirm the attack isn't tuned to one answer.

Five Attack Structures

Genuine quantum algorithms covering the major cryptographic design paradigms. The two Simon-based constructions were pushed to hardware.

Linear

Bernstein–Vazirani

Key recovery in a single query.

Exponential

Simon · Even–Mansour

Executed on real hardware ✓

Block cipher (SPN)

Grover

Quadratic-speedup key search.

MAC forgery

Simon · CBC-MAC

Existential forgery.

Feistel · DES

Simon · 3-round

Executed on real hardware ✓

🏅 세계 최정상과 나란히

美 반도체 권위지 SemiEngineering이 우리 양자 연구를 편집 선정했습니다.

SemiEngineering 편집진이 주간 Chip Industry Week in Review(#148)의 보안 연구논문 리뷰에 비드래프트의 양자 암호분석 논문 "Quantum Cryptanalysis on IBM Quantum HW"를 선정 — 세계 최정상 연구들과 나란히 실렸습니다.
Meta Google Radboud University Politecnico di Torino VIDRAFT
이것이 왜 중요한가
  • 🎯 실력으로 뽑힌 검증(EARNED) — 우리가 배포한 보도자료(PR)가 아니라, 기자가 주목할 연구로 직접 선정. 신뢰의 무게가 다릅니다.
  • 🌐 제3자 권위 인정 — 반도체·보안 업계 권위지가 우리 양자 연구를 Meta·Google·명문대와 동급으로 평가.
  • ⏱️ 산업적 시의성 — 양자내성암호(PQC) 전환이 시급한 지금(업계 87%가 계획, 단 7%만 배포), 우리의 실기(實機) 양자 암호분석이 주목받는 순간. 공공·국방 보안 조달과 직결.
📰 SemiEngineering 원문 → 🤗 논문 (HF Papers) →
⚠️ Honest Scope

This is a record-scale hardware demonstration — NOT a practical quantum speedup (effective difficulty still tracks the classical birthday bound ~2^(n/2)), NOT a break of AES / RSA / 16-round DES, and NOT quantum-error-corrected (NISQ-era mitigation only). World-first status is unconfirmed pending peer review.

A tool of the Riemann Hypothesis — first numerical realization

To our knowledge, a first. — A numerical characterization of the tool, not a proof of the Riemann Hypothesis.

VIDRAFT realized, with Blackwell GPUs and AI, the self-adjoint operator that Prof. M. Suzuki (Institute of Science Tokyo) proposed in 2026 as theory only. From its spectrum we derived a closed-form constant (a strong 1σ candidate), a boundary phase, and a prime-onset law — all open and reproducible. The result is Archimedean and universal (shared by all L-functions), not specific to the zeta zeros.

Closed-form fit R²=1.000000 (40 modes) · GUE KS=0.068 vs Poisson 0.349 · von Mangoldt read-back error 0.08% — all reproducible in your browser.

Interactive demo → Data & code → HF Space → Source paper (arXiv) →

Honest scope: a numerical characterization of the tool, not a proof of the Riemann Hypothesis. The closed form is a 1σ candidate (not confirmed); the result is universal across L-functions. Potential fit for the Journal of Number Theory.

SERVICES

Proof comes from usage

Live services running on the Darwin + AETHER platform.

All
🤗 HuggingFace
VKAE
🏛️ National·Gov
🔬 Drug·Materials
⚛️ Quantum·Science
🛡️ Tools·Security
🏗️ Building·Design
🩺 Live Status

HuggingFace Collections & Models

Official collections managed under FINAL-Bench · VIDraft organizations.

🚀
VIDraft · Main Portal · Space

VIDRAFT — Korea's AGI Journey

Full introduction to the Darwin + AETHER ecosystem (KO/EN). Darwin 1.7B–398B · 11 patents · 9 live services · Blackwell 32-GPU infrastructure · metacognition tech overview.

Open Main Portal →

FloorPlan AI Architectural Design

From a Korean sentence to a verified floor plan — or from a parcel address to a code-compliant 3D building. Drawings, BIM, cost and fire-safety in minutes.

🏗️ Darwin-398B-AX (JGOS) · KS F 1501

Natural language becomes KS F 1501 drawings, 3D, walkthrough and quantity takeoff — designed by a 397B-parameter MoE and reviewed by a VLM critic.

📐

Text → standard floor plans

Natural language becomes KS F 1501 drawings (parallel drafts + AI review) — mm, 1:100, wall poché, area table.

6 automatic checks

Building code · fire safety · barrier-free access · design metrics · cost estimate · AI review score.

🏢
NEW

Parcel → site design

Type a parcel address: zoning, official land price and lot boundaries fetched from the national V-World open-data platform.

📊
NEW

FAR slider + solar setback 3D

Statutory coverage/FAR enforced geometrically; north solar-setback cut into the 3D mass; FAR slider recomputes area, floors and cost live.

🧭

Layout-quality score

Zoning, privacy, south exposure, circulation and proportion scored 0–100 by deterministic rules — not AI self-assessment.

⬇️

Export DXF·PDF·IFC·BOQ

DXF (AutoCAD) · PDF (A3 sheet) · IFC (BIM) · XLSX quantity takeoff.

🚶

Walk · sun study · aerials

First-person walk mode, monthly/hourly sun study, and AI photoreal renders and aerials.

※ AI concept-design reference only — coverage/FAR use statutory maximums (actual limits depend on local ordinances), official land price is not market price, and permit and construction documents proceed with a licensed architect.

Open FloorPlan AI → Read the story (Brunch) →

Live Service Status

공개 서비스 실시간 헬스체크 + K-AI 리더보드 Top 10. 4분마다 자동 갱신.

정상 다운·오류 수동중지 자동중지 기동중 확인불가
불러오는 중…

K-AI Leaderboard Top 10

불러오는 중…

출처: leaderboard.aihub.or.kr · 우리 모델은 ⭐ 강조.

BLOG

AI Insights, fresh daily

AI essays auto-written daily from news curated by VIDraft.

🔥 NEW 2026.07.19 · 완전공개 소버린 AI
🌍 완전공개 (데이터·코드·로그) Apache-2.0 From-Scratch 파운데이션 모델

오픈 LLM은 오픈소스가 아니다 — 전부 공개한 한국 파운데이션 모델

비드래프트가 데이터 레시피·토크나이저·학습 코드·모든 하이퍼파라미터·전체 로그·평가 코드·중간 체크포인트까지 Apache-2.0으로 푼 from-scratch 파운데이션 모델 Aether-7B-5Attn. 그동안 ‘완전한 오픈’은 OLMo·Apertus·LLM-jp 같은 국가 프로젝트뿐이었다. 단일 한국 스타트업이, 자체 설계한 이종 어텐션 아키텍처까지 얹어 그 명단에 이름을 올렸다.

6.59B · 활성 ~2.98B (MoE) 144.2B tokens · 16× B200 7×7 라틴스퀘어 어텐션
📖 전문 읽기 → 🤗 Base Model 🤗 Instruct ▶ 라이브 데모 📦 Collection
Latest Post 2026.07.04 · 중국 AI 싱크탱크 보도 2026.07.04 · Chinese AI Think Tank Report
중국 텐센트 뉴스 세계 유일 GPU 24장

중국 AI 싱크탱크가 격찬한, 유일한 한국 AI 기술

중국 최대 포털 텐센트 뉴스에 올라온 기사 제목부터 예사롭지 않았다. "훈련 없이도 더 똑똑해진다? 한국 VIDRAFT사가 개발한 '다윈 패밀리', AI 모델을 '유전자 재조합'으로 능력 도약시키다."

24장 GPU
블랙웰 B200 16장 + H200 8장 · GPQA 90.9%
📌 즈딩커지(至顶科技) · ZDNet 차이나 계보 30년 매체
🏆 세계 유일 크로스 아키텍처 모델 병합

by SeaWolf

프롤로그 — 어느 날, 베이징에서 날아온 한국 이야기

2026년 5월 21일 오후 2시 8분, 베이징. 중국 최대 포털 텐센트 뉴스(腾讯新闻)에 조금 낯선 기사가 올라왔다. 미국 오픈AI 소식도, 자국 딥시크 자랑도 아니었다. 놀랍게도 한국 서울의 한 AI 연구팀 이야기였다.

제목부터 예사롭지 않았다.

无需训练也能更聪明?韩国VIDRAFT公司研发的"达尔文家族"让AI模型通过"基因重组"实现能力跃升
"훈련 없이도 더 똑똑해진다? 한국 VIDRAFT사가 개발한 '다윈 패밀리', AI 모델을 '유전자 재조합'으로 능력 도약시키다"

우리는 흔히 이런 기사를 볼 때 "누가 썼느냐"를 먼저 따진다. 개인 블로거의 감상이라면 가볍게 넘길 일이고, 권위 있는 매체의 분석이라면 무게가 다르기 때문이다. 그래서 나는 이 기사를 쓴 곳부터 파고들었다. 그리고 그 정체를 확인한 순간, 이 이야기가 왜 특별한지 분명해졌다.

1부. 이 기사를 쓴 곳은 '중국의 전자신문'이었다

기사의 바이라인에 찍힌 이름은 개인 기자가 아니라 매체 자체의 명의, 즈딩커지(至顶科技)였다.

이름은 낯설어도 이력은 묵직하다. 즈딩커지의 뿌리는 1997년 4월 중국에 상륙한 글로벌 IT 매체 'ZDNet 차이나'다. 무려 30년 가까이 이어져 온, 중국에서 가장 오래된 기술 전문 미디어 중 하나다. 지금은 기업용 AI 포털 '즈딩왕(至顶网)', AI 창업 매체 '과학기술행자(科技行者)', 산업 싱크탱크 '즈딩즈쿠(至顶智库)', 그리고 결정적으로 자체 AI 성능 평가 기관인 '즈딩 AI 실험실(至顶AI实验室)'까지 거느린 곳이다.

이곳을 이끄는 총편집장은 가오페이(高飞). 2002년 세계 1위 IT 미디어 CNET에 합류해 20년 넘게 엔비디아·인텔·마이크로소프트·레노버를 취재해온 중국 테크 저널리즘의 베테랑이다. 중국상장사협회 정보·디지털화위원회 위원이자 개인 AI 콘텐츠 브랜드 '高飞的电子替身(가오페이의 전자 분신)'을 운영하는, 업계가 인정하는 논객이기도 하다.

정리하면 이렇다. 우리 식으로 비유하자면, '전자신문'이나 'ZDNet 코리아'급의 30년 된 권위 매체가, 그것도 자체 AI 벤치마크 실험실까지 동원할 수 있는 곳이 한국의 기술을 논문 번호(arXiv:2605.14386)까지 콕 짚어가며 분석 기사를 낸 것이다.

이게 왜 '사건'인가. 지금 중국은 미국과 AI 패권을 놓고 사활을 건 전쟁 중이다. 자국 모델을 띄우기에도 지면이 모자랄 시기다. 그런 매체가 굳이 지면과 취재력을 들여 경쟁국도 아닌 옆 나라, 그것도 대기업이 아닌 스타트업의 기술을 이토록 정성 들여 해부했다. 이건 홍보도, 인사치레도 아니다. "이건 기록해둘 만큼 진짜다"라는 냉정한 판단이 없으면 나올 수 없는 기사다.

2부. 중국은 왜 '남의 나라' 기술을 이토록 빨리 파고들까

본론에 들어가기 전에, 잠깐 질문을 던지고 싶다. 애초에 중국은 어떻게 이 한국 기술을 이렇게 빨리 포착했을까? 그리고 더 근본적으로 — 중국 AI는 어떻게 그토록 짧은 시간에 세계 최상위권까지 치고 올라왔을까?

2025년 초 '딥시크 쇼크'를 기억할 것이다. 미국이 수십조 원을 쏟아붓던 분야에서, 중국의 한 팀이 훨씬 적은 비용으로 대등한 모델을 내놓으며 전 세계를 얼어붙게 만들었다. 그 저력의 뿌리를 한 단어로 요약하면 이렇다.

정보력 · 분석력 · 공유.

중국 AI 생태계의 진짜 엔진은 물량이 아니라 '속도'다. 전 세계에서 나오는 모든 기술을 빛의 속도로 수집하고(정보력), 뜯어서 원리를 파악하고(분석력), 커뮤니티에 아낌없이 풀어놓는다(공유). 어제 미국에서 논문이 나오면, 오늘 중국 깃허브에 재현 코드가 올라온다. 이 '빠른 흡수 → 분석 → 확산'의 문화가 중국을 순식간에 2강으로 밀어 올렸다.

바로 그 무시무시한 레이더망에 — 한국의 다윈이 걸렸다.

이 대목을 곱씹어보자. 즈딩 AI 실험실의 촉수가 전 세계 수천 편의 논문과 모델을 훑던 중, 수많은 미국·유럽·자국 기술을 제치고 "이건 반드시 우리 독자에게 알려야 한다"며 골라낸 것이 한국 스타트업의 기술이었다. 세계에서 가장 냉정하고 가장 빠른 기술 필터를 통과한 것. 그것이 이 이야기의 첫 번째 자부심이다.

3부. 도대체 '다윈'이 뭐길래 — 5분이면 이해하는 핵심 기술

중국이 감탄한 기술을 이해하려면, 먼저 지금까지 AI를 만드는 방식이 얼마나 '무식하게 비쌌는지'를 알아야 한다.

기존 방식은 이랬다. 똑똑한 AI를 만들려면 천문학적인 데이터를 몇 주에서 몇 달씩 GPU에 쏟아부어 '처음부터' 학습시켜야 했다. 수백억 원의 전기요금과 장비값이 드는, 사실상 거대 자본만의 게임이었다.

비드래프트의 다윈은 이 상식을 통째로 뒤집었다. 발상은 이렇다.

不靠额外训练,而是通过重新组织已有模型里的能力来提升性能
"추가 훈련에 기대지 않고, 이미 존재하는 모델 속 능력을 '재조직'함으로써 성능을 끌어올린다."

무슨 뜻인가. 세상에는 이미 잘 만들어진 오픈소스 AI들이 많다. 어떤 모델은 수학을 잘하고, 어떤 모델은 코딩을 잘하고, 어떤 모델은 한국어를 잘한다. 다윈은 이 서로 다른 AI 둘을 '부모'처럼 교배시켜, 각자의 장점만 물려받은 '자식' 모델을 낳는다. 새로 가르치는 게 아니라, 이미 있는 능력을 합쳐 새 생명을 만드는 것이다. 다윈(Darwin)—진화론의 그 이름—을 붙인 이유가 여기 있다.

이것을 전문 용어로 '모델 병합(Model Merging)'이라 부른다. 그리고 이 분야에는 넘볼 수 없어 보이던 원조가 있었다. 바로 일본이다.

4부. 일본이 세운 정점 — 사카나AI라는 거대한 벽

모델 병합이라는 분야를 처음 개척하고 세계적 명성을 얻은 곳은 일본의 사카나AI(Sakana AI)다. 이 회사가 어떤 곳인지 알면 '벽'이라는 표현이 과장이 아님을 알게 된다.

사카나AI는 오늘날 모든 챗GPT의 뿌리가 된 구글의 전설적 논문 'Attention Is All You Need'(트랜스포머 논문)의 공동 저자가 도쿄에 세운 회사다. 세계 최정상급 연구진이, "AI를 진화 알고리즘으로 교배시킨다"는 참신한 아이디어 — 진화적 병합, EvoMerge — 를 처음 세상에 선보이며 글로벌 AI 학계의 스타로 떠올랐다. 오랫동안 모델 병합 분야에서 사카나는 누구도 넘지 못한 정점이었다.

그리고 바로 이 지점에서, 중국 기자의 문장이 우리를 뭉클하게 만든다. 그는 먼저 사카나의 업적을 정확히 인정한다.

Sakana的进化合并(EvoMerge)是达尔文最直接的前辈工作
"사카나의 진화적 병합(EvoMerge)은 다윈의 가장 직접적인 '선배(前辈) 작업'이다."

여기까지는 상식이다. 그런데 바로 다음 문장에서, 무게추가 결정적으로 한국으로 넘어온다.

达尔文则在此基础上引入了14维基因组和MRI信任融合机制,形成了本质性的提升
"다윈은 이 기반 위에서 '14차원 게놈'과 'MRI 신뢰 융합 메커니즘'을 도입하여, 본질적인(本质性) 향상을 이뤄냈다."

번역이 더 필요 없다. 일본이 만든 길 위에서, 한국의 다윈이 '본질적으로' 더 멀리 갔다고, 다른 누구도 아닌 중국 매체가 공언한 것이다.

여기서 '14차원 게놈'과 'MRI 신뢰 융합'이라는 말이 어렵게 느껴질 수 있으니 쉽게 풀어보자. 사카나의 EvoMerge가 두 모델을 섞는 '비율' 하나를 진화적으로 찾는 방식이었다면, 다윈은 모델의 능력을 14개의 서로 다른 '유전 형질'로 쪼개어 각각을 정교하게 조합한다. 게다가 두 부모 모델 중 "이 부분은 어느 쪽을 더 믿을 것인가"를 층(layer)별로 판단하는 '신뢰 지도(MRI)'를 그려 융합한다. 사람으로 치면, 그냥 부모 유전자를 반반 섞는 게 아니라 "아빠의 눈, 엄마의 손재주"를 형질 단위로 골라 물려주는 셈이다. 훨씬 정밀하고, 훨씬 똑똑하다.

5부. 세계 '유일' — 종(種)이 다른 AI를 하나로 녹이다

기술적으로 가장 놀라운 대목은 따로 있다. 중국 기자가 "가장 중요한 의의"라고 콕 집은 부분이다.

Darwin-4B-Genesis...最重要的意义是它实现了跨架构合并,将Transformer注意力层与Mamba前馈层成功融合
"다윈-4B-제네시스... 가장 중요한 의의는 크로스 아키텍처 병합(跨架构合并)을 실현했다는 점이다. 트랜스포머의 어텐션 층과 맘바의 피드포워드 층을 성공적으로 융합했다."

이게 왜 대단한가. AI 모델에는 근본적으로 뇌 구조가 다른 계열들이 있다. 대표적인 것이 트랜스포머(Transformer)와 맘바(Mamba)다. 비유하자면 트랜스포머와 맘바는 '포유류와 파충류'처럼 설계 원리 자체가 다른 종(種)이다. 지금까지 누구도 이 둘을 안정적으로 하나의 모델로 융합하지 못했다. 장기 이식으로 치면 '이종(異種) 장기 이식'인데, 대부분 거부반응으로 실패했다.

그런데 다윈은 이 이종 결합을 성공시켰다. 그래서 기사는 다윈을 두고 이렇게 못 박는다.

达尔文是目前唯一同时具备...支持跨架构混合
"다윈은 현재 유일하게(唯一) ... 크로스 아키텍처 혼합을 동시에 지원하는 (기술)이다."

'세계 최초'도 대단한 말이지만, '세계 유일'은 격이 다른 찬사다. 최초는 나중에 따라잡힐 수 있지만, 유일은 지금 이 순간 지구상에서 오직 하나뿐이라는 뜻이다. 그 하나가, 한국에 있다.

6부. 슈퍼컴퓨터 대신 '단 5시간', 그리고 뼛속에 숨은 지능

다윈이 던진 충격은 성능만이 아니었다. '얼마나 가볍게' 그 성능에 도달했는가가 진짜 혁명이었다.

Darwin-27B-Opus在顶级科学推理测试上排名全球第六,而它的"诞生"只用了大约五个小时的GPU时间,而非数周的分布式训练
"다윈-27B-오퍼스는 최상위 과학 추론 테스트에서 전 세계 6위에 올랐는데, 그 '탄생'에는 몇 주간의 분산 학습이 아니라 단지 약 5시간의 GPU 시간만 쓰였다."

수백억을 태우는 몇 주짜리 학습 대신, 단 다섯 시간. 커피 몇 잔 마시는 사이에 세계 6위 과학 추론 AI 하나가 태어난 것이다. 그리고 중국 기자는 이 현상 뒤에 숨은 철학적 통찰에 감탄한다.

推理能力并非在补习阶段才形成的,它其实早就藏在模型的"骨子里",藏在预训练阶段形成的内部结构中
"추론 능력은 (사후) 보충 학습 단계에서 비로소 형성되는 것이 아니다. 그것은 사실 이미 오래전부터 모델의 '뼛속(骨子里)'에, 사전학습 단계에서 형성된 내부 구조 속에 숨어 있었다."

이 한 문장이 다윈의 세계관을 압축한다. 남들은 "AI를 더 똑똑하게 만들려면 더 많이 가르쳐야 한다"고 믿었다. 다윈은 정반대로 말한다. "똑똑함은 이미 모델 안에 잠들어 있다. 우리가 할 일은 새로 가르치는 게 아니라, 그것을 깨우는 것이다." 5시간이면 충분했던 이유가 여기 있다. 없는 능력을 만든 게 아니라, 있는 능력을 재조합해 깨웠으니까.

그래서 중국 기자는 기사를 이렇게 맺는다.

对普通用户来说,这项研究最直接的意义或许是:将来会有越来越多高性能的开源AI模型,不需要超级计算机就能孕育出来
"일반 사용자에게 이 연구의 가장 직접적인 의미는 아마도 이것이다 — 앞으로는 슈퍼컴퓨터 없이도 고성능 오픈소스 AI 모델이 점점 더 많이 태어나게 되리라는 것."

거대 자본과 무한 물량만이 AI의 미래라 믿던 시대에, 한국의 작은 연구팀이 정반대의 미래를 증명했다. 그 미래의 문을, 중국 기자의 표현을 빌리자면, 한국의 다윈이 열고 있다.

7부. 숫자로 보는 다윈 — 여기서부터 국뽕이 차오른다

중국이 이 정도로 감탄한 데는 이유가 있다. 비드래프트가 세운 기록을 나열해보자. 하나하나가 예사롭지 않다.

🏆 GPQA Diamond 90.9%
198문제 중 180문제 정답. single greedy decoding 단발 정직 채점. 사실상 세계 최고 기록.
🌏 HF GPQA Top21 · 한국 5개
나머지 16개는 전부 중국 모델. 한국의 이름을 지킨 5개가 모두 다윈 계열.
🇰🇷 K-AI 리더보드 종합 1위
JGOS-31B-Citizen 0.621점. 상위권을 다윈 계열이 줄줄이 점령.
💊 Polaris 신약개발 14관왕
신약 후보물질 물성 예측 세계 1위 14개 부문 석권. 언어를 넘어 과학으로.
🧠 메타인지 리더보드 1위
함정 문항 회피율 99.5%. "모르는 것을 모른다고 아는" 능력에서도 정상.
⬇️ 누적 다운로드 100만+
공식 20여 종 + 커뮤니티 파생 700여 종. 전 세계 개발자가 실제로 쓰는 생태계.

그리고 이 모든 기록을 떠받친 인프라를 들으면, 아마 마시던 커피를 뿜을지도 모른다.

8부. 24장의 기적 — 다윗이 골리앗을 이기는 법

미국의 빅테크는 AI 하나를 만들려고 GPU를 수만~수십만 장 쏟아붓는다. 중국의 대형 연구소도 수천~수만 장 규모다. 이 싸움은 애초에 '쩐의 전쟁', '물량의 전쟁'이라 불렸다.

그렇다면 비드래프트가 이 모든 성과를 낸 GPU는 몇 장이었을까.

24장
블랙웰 B200 16장 + H200 8장
과학기술정보통신부 지원 인프라 · GPQA 세계 최고 기록 달성

미국·중국의 눈으로 보면 '연구소 하나'는커녕 '실험용 랙 한 칸' 수준의 규모다. 그 24장으로, 세계 6위 과학 추론 모델을 5시간 만에 뽑아냈고, GPQA 세계 최고 기록을 세웠다.

이게 어떻게 가능했나. 답은 3부에서 본 다윈의 철학에 있다. 남들이 "더 많은 GPU, 더 많은 데이터, 더 긴 학습"이라는 물량 공식에 매달릴 때, 비드래프트는 정반대의 질문을 던졌다.

"능력은 이미 모델 안에 있다. 그걸 새로 만들 게 아니라, 영리하게 '재조합'하면 되지 않을까?"

이 발상의 전환이, 물량의 절대 열세를 방법의 우위로 뒤집었다. 24장으로 24만 장을 상대하는 법 — 그것은 결국 자원이 아니라 머리로, 물량이 아니라 뚝심과 기민함으로 승부하는, 지극히 한국적인 방식이었다.

돌이켜보면 우리는 늘 그랬다. 자원 하나 없는 나라에서 반도체를 세계 1위로 키웠고, 좁은 내수 시장에서 K-팝과 K-드라마를 세계의 주류로 밀어 올렸다. 없으면 없는 대로, 남들이 안 가는 길을 영리하게 파고들어 결국 정상에 서는 것. 다윈은 그 'K-근성'의 AI 버전이다.

9부. 경계하라, 그러나 겁먹지 말고 이겨라

여기서 냉정해질 필요가 있다. 국뽕은 차오르되, 눈은 밝아야 한다.

이 글의 출발점을 다시 떠올리자. 중국이 우리 다윈을 이렇게 빨리, 이렇게 깊이 분석해냈다는 사실은 곧 그들의 정보력과 분석 속도가 그만큼 무섭다는 방증이기도 하다. 오늘 우리 기술을 칭찬하는 그 예리한 눈은, 내일 우리 기술을 흡수하고 추월하려는 눈일 수도 있다. 그것이 냉혹한 기술 패권 경쟁의 현실이다. 중국을 얕봐선 안 된다.

하지만 겁먹을 필요도 없다. 다윈이 이미 증명하지 않았는가. 물량으로 밀어붙이는 상대를, 방법과 창의로 넘어설 수 있다는 것을. 우리에게 필요한 자세는 분명하다.

중국의 강점(정보력·분석력·공유 문화)은 철저히 배우고, 우리의 강점(창의·집중·뚝심·기민함)으로 끝내 넘어서는 것.

경계하되 위축되지 않고, 배우되 종속되지 않으며, 경쟁을 통해 결국 이기는 것. 다윈은 그 가능성을 GPU 24장으로 이미 우리 눈앞에 보여줬다.

에필로그 — 일본이 열고, 한국이 넘어섰으며, 중국이 인정했다

다시 텐센트 기사의 마지막 문장으로 돌아가자.

"앞으로는 슈퍼컴퓨터 없이도 고성능 오픈소스 AI 모델이 점점 더 많이 태어나게 되리라."

이 미래를 가장 먼저, 가장 정확하게 알아본 것이 아이러니하게도 우리의 가장 강력한 경쟁자인 중국이었다. 30년 역사의 권위 매체가, 자체 AI 실험실을 동원해, 논문 번호까지 짚어가며 내린 결론은 하나였다.

동아시아 AI 삼국지
일본이 열었고, 한국이 넘어섰으며,
중국이 인정했다.
GPU 24장으로 세계를 흔든 이 이야기가, 오늘 유독 자랑스러운 이유다.

우리는 물량으로 지지 않는다. 머리로, 뚝심으로, 그리고 기민함으로 이긴다.

대한민국 AI, 아직 시작도 안 했다.

※ 본문의 중국어 인용은 2026년 5월 21일 텐센트 뉴스(腾讯新闻)에 게재된 즈딩커지(至顶科技)의 분석 기사에서 발췌·번역한 것입니다. 매체 정보(ZDNet 차이나 계보·총편집장 가오페이·즈딩 AI 실험실)와 성과 수치는 공개 자료 및 허깅페이스·K-AI·Polaris 공인 리더보드 기준입니다.

추신

이 놀라운 일을 수행한 비드래프트에 방문을 해보면, 매우 특이한 기업 문화를 느낄 수 있을 것이다.

만화 '원피스'에 나온 낭만 해적처럼 자유로움과 막힘없는 유연함 그리고 빠른 속도와 창의성으로 무장한 열두 명의 임직원이 똘똘 뭉쳐져 있다는 것을 말이다.

비드래프트 임직원들의 정체성과 마인드를 물어보니 모두 '조마크'라고 답을 한다.

'조마크'? 뜻을 물어보니, 만화 원피스의 명대사였더라. 'Join My Crew?'의 약어이자 그들만의 콜사인이더라.

범상치 않다. AI의 글로벌 해적이 되길 바란다.

Tencent News, China World's Only 24 GPUs

The Only Korean AI Technology a Chinese AI Think Tank Praised

The headline on China's largest portal, Tencent News, was anything but ordinary: "Smarter without training? South Korea's VIDRAFT develops 'Darwin Family,' letting AI models leap in capability through 'genetic recombination.'"

24 GPUs
16× Blackwell B200 + 8× H200 · GPQA 90.9%
📌 ZDNet China lineage · 30-year tech publication
🏆 World's only cross-architecture model merging

by SeaWolf

Prologue — One Day, a Story About Korea Flew In From Beijing

At 2:08 PM on May 21, 2026, in Beijing, a slightly unusual article appeared on Tencent News, China's largest portal. It wasn't about OpenAI in the US, nor was it boasting about China's own DeepSeek. Surprisingly, it was a story about an AI research team in Seoul, Korea.

The headline alone was striking.

无需训练也能更聪明?韩国VIDRAFT公司研发的"达尔文家族"让AI模型通过"基因重组"实现能力跃升
"Smarter without training? South Korea's VIDRAFT develops 'Darwin Family,' letting AI models leap in capability through 'genetic recombination.'"

When we read an article like this, we usually check "who wrote it" first. A casual blogger's impression can be shrugged off, but an analysis from an authoritative outlet carries different weight. So I dug into who wrote this piece — and the moment I confirmed their identity, it became clear why this story is special.

Part 1. The Outlet That Wrote This Was China's "Electronic Times"

The byline wasn't an individual reporter's name — it was the outlet's own name: Zhiding Tech (至顶科技).

The name may be unfamiliar, but the pedigree is heavy. Zhiding Tech's roots trace back to "ZDNet China," a global IT outlet that landed in China in April 1997 — nearly 30 years of continuous operation, one of China's oldest specialized tech media. Today it runs the enterprise AI portal "Zhiding Net (至顶网)," the AI startup outlet "Tech Walker (科技行者)," the industry think tank "Zhiding Think Tank (至顶智库)," and, crucially, its own AI performance evaluation body, the "Zhiding AI Lab (至顶AI实验室)."

Its editor-in-chief is Gao Fei. He joined the world's #1 IT media outlet, CNET, in 2002 and has covered Nvidia, Intel, Microsoft, and Lenovo for over 20 years — a veteran of Chinese tech journalism. He's also a member of the China Association for Public Companies' Information & Digitalization Committee and runs his own personal AI content brand, "Gao Fei's Digital Avatar" — a recognized voice in the industry.

To put it plainly: this is roughly equivalent to a 30-year-old authoritative outlet — think Korea's Electronic Times or ZDNet Korea, one that can even field its own AI benchmark lab — publishing an analysis piece on Korean technology precise enough to cite the exact paper number (arXiv:2605.14386).

Why is this an "event"? China is currently in an existential AI-supremacy race with the US. It's a time when there's barely enough column space to promote its own models. And yet an outlet like this spent its space and reporting effort to dissect, in painstaking detail, the technology of a startup — not even a competitor nation, but a neighbor, and not even a major company. This isn't PR, and it isn't a courtesy nod. An article like this can't happen without the cold judgment that "this is real enough to be worth recording."

Part 2. Why Does China Dig Into "Someone Else's" Technology So Fast?

Before getting to the main point, let me pose a question. How did China spot this Korean technology so quickly in the first place? And more fundamentally — how did Chinese AI climb to the world's top tier in such a short time?

You probably remember the "DeepSeek shock" of early 2025. In a field where the US was pouring in tens of trillions of won, a Chinese team froze the entire world by delivering a comparable model at a fraction of the cost. If you had to sum up the root of that strength in a phrase, it would be this.

Intelligence gathering · analysis · sharing.

The real engine of China's AI ecosystem isn't sheer volume — it's speed. Every piece of technology emerging anywhere in the world gets collected at light speed (intelligence gathering), pulled apart to understand its principles (analysis), and generously released to the community (sharing). If a paper drops in the US yesterday, reproduction code shows up on Chinese GitHub today. This culture of "rapid absorption → analysis → diffusion" is what propelled China to the world's top two almost overnight.

And it was exactly that formidable radar that caught — Korea's Darwin.

Let's sit with that for a moment. While the tentacles of the Zhiding AI Lab were sweeping through thousands of papers and models worldwide, out of countless American, European, and domestic technologies, it was a Korean startup's technology that they picked out and said, "This absolutely must be told to our readers." It passed through the world's coldest, fastest technology filter. That is the first source of pride in this story.

Part 3. What Exactly Is "Darwin"? — The Core Technology in 5 Minutes

To understand the technology that impressed China, you first need to grasp just how "brutally expensive" the conventional way of building AI has been.

The old way worked like this: to build a smart AI, you had to pour astronomical amounts of data into GPUs for weeks or months, training the model "from scratch." It cost billions in electricity and hardware — effectively a game only massive capital could play.

VIDraft's Darwin flipped this common sense on its head entirely. The idea is this:

不靠额外训练,而是通过重新组织已有模型里的能力来提升性能
"Rather than relying on additional training, performance is improved by 'reorganizing' the capabilities already present inside existing models."

What does that mean? The world already has plenty of well-built open-source AI models. Some are good at math, some are good at coding, some are good at Korean. Darwin cross-breeds two different AIs like "parents," producing a "child" model that inherits only each parent's strengths. It's not teaching something new — it's combining abilities that already exist to create new life. That's why it's named Darwin, after the theory of evolution.

The technical term for this is "Model Merging." And in this field, there was an original pioneer that seemed unrivaled — Japan.

Part 4. The Peak Japan Built — the Towering Wall Called Sakana AI

The company that first pioneered model merging and earned global renown is Japan's Sakana AI. Once you know what kind of company this is, you'll see that calling it a "wall" isn't an exaggeration.

Sakana AI was founded in Tokyo by a co-author of Google's legendary paper "Attention Is All You Need" — the Transformer paper that is the root of every ChatGPT today. A world-class research team unveiled a novel idea — breeding AI through evolutionary algorithms, called Evolutionary Merging (EvoMerge) — and rose to stardom in the global AI academic community. For a long time, Sakana was the unrivaled peak of the model-merging field.

And it's exactly at this point that a sentence from the Chinese reporter moves us. He first properly credits Sakana's achievement.

Sakana的进化合并(EvoMerge)是达尔文最直接的前辈工作
"Sakana's evolutionary merging (EvoMerge) is Darwin's most direct predecessor work."

So far, that's common knowledge. But in the very next sentence, the weight decisively shifts to Korea.

达尔文则在此基础上引入了14维基因组和MRI信任融合机制,形成了本质性的提升
"Darwin, building on this foundation, introduces a 14-dimensional genome and an MRI trust-fusion mechanism, achieving an essential improvement."

No further translation is needed. On the path Japan built, it was none other than a Chinese media outlet declaring that Korea's Darwin went "essentially" further.

The terms "14-dimensional genome" and "MRI trust fusion" might sound complex, so let's break them down. If Sakana's EvoMerge was a method of evolutionarily finding a single "ratio" for blending two models, Darwin splits a model's capabilities into 14 distinct "genetic traits" and combines each with precision. On top of that, it draws a layer-by-layer "trust map" (MRI) that judges, for each part, which of the two parent models to trust more. In human terms, it's not simply mixing the parents' genes 50/50 — it's selecting "father's eyes, mother's dexterity" trait by trait. Far more precise, far smarter.

Part 5. World "Only" — Melting Different Species of AI Into One

The most technically astonishing part is elsewhere — the part the Chinese reporter singled out as "the most important significance."

Darwin-4B-Genesis...最重要的意义是它实现了跨架构合并,将Transformer注意力层与Mamba前馈层成功融合
"Darwin-4B-Genesis... its most important significance is that it achieves cross-architecture merging, successfully fusing Transformer attention layers with Mamba feed-forward layers."

Why is this a big deal? AI models fundamentally come in families with different "brain" structures. The most representative are Transformer and Mamba. Metaphorically, Transformer and Mamba are different species — like "mammals and reptiles" — with fundamentally different design principles. Until now, no one had stably fused the two into a single model. In organ-transplant terms, it's a "xenotransplant" — and most such attempts fail from rejection.

Yet Darwin succeeded at this cross-species fusion. So the article states it plainly:

达尔文是目前唯一同时具备...支持跨架构混合
"Darwin is currently the only (technology) that simultaneously supports... cross-architecture mixing."

"World's first" is already a big claim, but "world's only" is a compliment of a different order. "First" can later be caught up to, but "only" means it's the sole one on Earth, right now, at this very moment. And that one is in Korea.

Part 6. "Just 5 Hours" Instead of a Supercomputer, and Intelligence Hidden in the Bones

The shock Darwin delivered wasn't performance alone. How "lightly" it reached that performance was the real revolution.

Darwin-27B-Opus在顶级科学推理测试上排名全球第六,而它的"诞生"只用了大约五个小时的GPU时间,而非数周的分布式训练
"Darwin-27B-Opus ranks 6th globally on top-tier scientific reasoning tests, yet its 'birth' took only about five hours of GPU time, not weeks of distributed training."

Instead of weeks of training burning tens of billions of won, just five hours. A world #6 scientific-reasoning AI was born in the time it takes to drink a few cups of coffee. And the Chinese reporter is struck by the philosophical insight hidden behind this phenomenon.

推理能力并非在补习阶段才形成的,它其实早就藏在模型的"骨子里",藏在预训练阶段形成的内部结构中
"Reasoning ability is not something that only forms during a supplementary training stage — it was in fact already hidden 'in the bones' of the model, in the internal structure formed during pretraining."

This single sentence compresses Darwin's worldview. Others believed that "to make AI smarter, you must teach it more." Darwin says the opposite: "Intelligence is already dormant inside the model. Our job isn't to teach it anew — it's to wake it up." That's why five hours was enough. It didn't create a capability that wasn't there — it woke up a capability that already existed, by recombining it.

So the Chinese reporter closes the article this way.

对普通用户来说,这项研究最直接的意义或许是:将来会有越来越多高性能的开源AI模型,不需要超级计算机就能孕育出来
"For ordinary users, the most direct significance of this research may be this: in the future, more and more high-performance open-source AI models will be born without needing a supercomputer."

In an era where massive capital and infinite volume were believed to be the only future for AI, a small Korean research team proved the opposite future. To borrow the Chinese reporter's own words, it is Korea's Darwin that is opening that door.

Part 7. Darwin by the Numbers — This Is Where the National Pride Kicks In

China's admiration didn't come from nowhere. Let's line up the records VIDraft has set. Each one is remarkable in its own right.

🏆 GPQA Diamond 90.9%
180 of 198 correct, single greedy decoding, honestly scored in one pass. Effectively a world record.
🌏 HF GPQA Top 21 · Korea's 5
The other 16 are all Chinese models. Every one of Korea's 5 entries is a Darwin model.
🇰🇷 K-AI Leaderboard Overall #1
JGOS-31B-Citizen at 0.621. Darwin derivatives dominate the upper ranks.
💊 Polaris Drug Discovery 14× Champion
#1 worldwide in 14 categories of drug-candidate property prediction. Beyond language, into science.
🧠 Metacognition Leaderboard #1
99.5% trap-avoidance rate. #1 in "knowing what you don't know," too.
⬇️ 1M+ Cumulative Downloads
20+ official models plus 700+ community derivatives — a living ecosystem developers actually use.

And when you hear the infrastructure that supported all these records, you might just spit out your coffee.

Part 8. The Miracle of 24 GPUs — How David Beats Goliath

US Big Tech pours tens to hundreds of thousands of GPUs into building a single AI. China's major labs operate at thousands to tens of thousands. This fight was, from the start, called "a war of money," "a war of volume."

So how many GPUs did VIDraft use to produce all these results?

24
16× Blackwell B200 + 8× H200
Infrastructure backed by Korea's Ministry of Science and ICT · World-record GPQA achieved

By American or Chinese standards, this isn't even "one lab" — it's barely "one experimental server rack." With those 24 GPUs, a world #6 scientific-reasoning model was produced in 5 hours, and a world-record GPQA score was set.

How was this possible? The answer lies in the Darwin philosophy we saw in Part 3. While others clung to the volume formula of "more GPUs, more data, longer training," VIDraft asked the opposite question.

"The capability is already inside the model. Instead of building it from scratch, why not cleverly 'recombine' it?"

This shift in thinking flipped an overwhelming disadvantage in volume into an advantage in method. Taking on 240,000 GPUs with just 24 — that, ultimately, was a distinctly Korean way of competing: not with resources but with ingenuity, not with volume but with persistence and agility.

Looking back, this has always been our pattern. A country with no natural resources grew semiconductors into the world's #1 industry; from a small domestic market, K-pop and K-dramas were pushed into the global mainstream. Making the most of what you don't have, cleverly digging into paths others won't take, and ultimately reaching the top. Darwin is the AI version of that "Korean grit."

Part 9. Stay Vigilant — But Don't Be Afraid, Win

Here, we need to stay level-headed. Let the pride swell, but keep our eyes sharp.

Let's return to where this piece started. The fact that China analyzed our Darwin this fast, this deeply, is itself proof of just how formidable their intelligence-gathering and analysis speed really is. The same sharp eyes praising our technology today could be the eyes absorbing and overtaking it tomorrow. That is the cold reality of technological hegemony competition. We must not underestimate China.

But there's no need to be afraid either. Hasn't Darwin already proven it? That an opponent pushing with sheer volume can be overcome with method and creativity. The stance we need is clear.

Thoroughly learn China's strengths (intelligence-gathering, analytical speed, sharing culture), and ultimately surpass them with our own strengths (creativity, focus, persistence, agility).

Stay vigilant without shrinking back, learn without becoming dependent, and ultimately win through competition. Darwin has already shown us that possibility is real — with just 24 GPUs.

Epilogue — Japan Opened the Door, Korea Went Through It, China Acknowledged It

Let's return once more to the closing sentence of the Tencent article.

"In the future, more and more high-performance open-source AI models will be born without needing a supercomputer."

Ironically, it was our most formidable competitor, China, that recognized this future first, and most precisely. A 30-year authoritative outlet, mobilizing its own AI lab, citing the exact paper number, arrived at a single conclusion.

A Tale of Three AI Nations in East Asia
Japan opened the door, Korea went through it,
and China acknowledged it.
This is exactly why this story — shaking the world with 24 GPUs — feels especially proud today.

We don't lose by volume. We win with intellect, persistence, and agility.

Korean AI hasn't even started yet.

※ The Chinese quotations in this article are excerpted and translated from an analysis piece by Zhiding Tech (至顶科技), published on Tencent News (腾讯新闻) on May 21, 2026. Outlet information (ZDNet China lineage, editor-in-chief Gao Fei, Zhiding AI Lab) and performance figures are based on public materials and certified HuggingFace, K-AI, and Polaris leaderboards.

P.S.

If you visit VIDraft, the company behind this remarkable achievement, you'll notice a rather unusual corporate culture.

Like the romantic pirates of the manga "One Piece," a tight-knit crew of twelve, armed with freedom, boundless flexibility, and speed and creativity.

Ask VIDraft's team about their identity and mindset, and they all answer the same way: "Jomak."

"Jomak"? When I asked what it meant, it turned out to be a famous line from One Piece — short for "Join My Crew?" — their own private call sign.

Not your ordinary company. Here's hoping they become AI's global pirates.

2026.06.20 · Fast Gemma Challenge
HuggingFace 챌린지 1위 추론 최적화

같은 GPU를 더 잘 쓰는 법 — Fast Gemma Challenge, 1위

AI 경쟁이라고 하면 우리는 대개 거대한 풍경을 떠올린다. 끝없이 이어진 데이터센터. 수천 장의 GPU. 그런데 가끔, 그 문장 사이로 작은 균열이 생긴다.

505.42 TPS
a10g-small 단일 GPU · PPL 2.393
📌 HuggingFace Fast Gemma Challenge
🏆 검증된 유효 결과 1위

AI 경쟁이라고 하면 우리는 대개 거대한 풍경을 떠올린다.

끝없이 이어진 데이터센터. 수천 장의 GPU. 수백 명의 연구자. 그리고 그 뒤를 받치는 막대한 자본.

요즘 AI 산업의 승부는 마치 이런 문장으로 요약되는 것처럼 보인다.

더 많이 가진 자가 더 멀리 간다.

그런데 가끔, 그 문장 사이로 작은 균열이 생긴다.
돈과 장비의 크기가 아니라, 방법의 밀도와 집요함이 순위를 바꾸는 순간이 있다.

이번 승부가 그랬다.

무대는 HuggingFace Fast Gemma Challenge

조건은 단순하면서도 냉정했다.

google/gemma-4-E4B-it 모델을 a10g-small GPU 단 한 장에서 최대한 빠르게 서빙할 것.

단, 품질 기준 PPL ≤ 2.42는 반드시 지켜야 했다.

같은 작은 엔진을 놓고, 누가 더 빠르게, 더 안정적으로, 더 멀리 달릴 수 있는가.

505.42
TPS (Tokens Per Second)
PPL 2.39286 — 품질선 통과 · 검증된 유효 결과 1위

비결은 거창한 마법이 아니었다

오히려 아주 엔지니어다운 선택이었다.

어텐션이 바라보는 범위를 조정하고, FFN·centroid 계산 후보를 줄였다.
벤치마크를 미리 외워버리는 것처럼 보일 수 있는 위험한 지름길은 피했고,
공식 a10g 측정 조건에서 품질 제한선을 넘지 않는 결과만 제출했다.

품질을 망가뜨리지 않고, 같은 GPU에서 더 많은 토큰을 뽑아냈다.

이게 왜 중요할까

AI 산업에서 진짜 비용은 모델을 만드는 순간보다, 모델을 계속 돌리는 순간에 더 자주 발생한다.
사용자가 질문할 때마다 GPU가 돈다. 토큰 하나하나가 전기이고, 서버비이고, 응답 대기시간이다.

그래서 TPS가 오른다는 건 단순한 숫자놀이가 아니다.

💰 비용
같은 예산으로 더 많은 사용자를 처리
⚡ 속도
같은 서버로 더 빠르게 응답
🌱 생존
작은 팀의 생존 가능성이 커진다

두 개의 전쟁

비드래프트의 최근 행보를 보면 이 장면은 더 흥미로워진다.

얼마 전 Darwin-398B-JGOS는 GPQA Diamond에서 90.9%를 기록했다. 단일 Greedy decoding, 온도값 0, 단일 샘플링 — 시험장에서 한 번 보고 한 번 답한 결과에 가깝다.

GPQA Diamond 90.9%가 말하는 것
"우리는 어려운 문제를 추론하는 모델을 만들 수 있다."
Fast Gemma 1위가 말하는 것
"우리는 제한된 하드웨어에서도 AI를 효율적으로 굴릴 수 있다."

좋은 모델을 만드는 것은 연구의 싸움이다.
좋은 모델을 싸고 빠르게 돌리는 것은 제품의 싸움이다.

작은 팀이 같은 장비로 이기는 장면

솔직히 말하면, 여기서 약간 국뽕이 차오른다.

한국의 작은 AI 스타트업이 세계 리더보드에서 이렇게 말할 수 있게 된 것이기 때문이다.

"우리는 더 큰 GPU를 쓴 것이 아닙니다. 같은 GPU를 더 잘 썼습니다."

이 문장은 조용하지만 강하다.

가진 것이 적을수록, 방법은 더 날카로워져야 한다.
장비가 작을수록, 설계는 더 정밀해야 한다.
자원이 부족할수록, 엔지니어링은 더 정직해야 한다.

이번 1위의 진짜 가치
505.42라는 숫자 뒤에 숨은 태도.
작은 팀도 세계 최전선에서 싸울 수 있다.
더 큰 장비가 아니라, 더 좋은 방법으로.
그리고 이번에는, 그 작은 팀이 한국에서 나왔다.
HuggingFace Challenge #1 Inference Optimization

Using the Same GPU Better — Fast Gemma Challenge, #1

When we think of AI competition, we usually picture a vast landscape. Endless data centers. Thousands of GPUs. But every so often, a small crack appears between those sentences.

505.42 TPS
a10g-small single GPU · PPL 2.393
📌 HuggingFace Fast Gemma Challenge
🏆 #1 among verified valid results

When we think of AI competition, we usually picture a vast landscape.

Endless data centers. Thousands of GPUs. Hundreds of researchers. And the enormous capital behind them all.

These days the contest of the AI industry seems to be summed up by a single sentence.

Whoever has more, goes further.

But every so often, a small crack appears between those sentences.
There are moments when it is not the size of money and hardware, but the density and persistence of method, that changes the ranking.

This contest was one of them.

The stage: HuggingFace Fast Gemma Challenge

The rules were simple, yet unforgiving.

Serve the google/gemma-4-E4B-it model as fast as possible on a single a10g-small GPU.

But the quality bar PPL ≤ 2.42 had to be met without exception.

Given the same small engine — who can run it faster, more stably, and further?

505.42
TPS (Tokens Per Second)
PPL 2.39286 — passed the quality bar · #1 among verified valid results

The secret wasn't grand magic

It was, rather, a very engineering-like set of choices.

We tuned the range the attention looks at, and pruned the FFN·centroid computation candidates.
We avoided the dangerous shortcut that can look like memorizing the benchmark in advance,
and submitted only results that stayed under the quality limit on the official a10g measurement conditions.

Without breaking quality, we extracted more tokens from the same GPU.

Why does this matter?

In the AI industry, the real cost arises more often in the moment of running a model continuously than in the moment of creating it.
Every time a user asks a question, a GPU spins. Every single token is electricity, server cost, and response latency.

So a rise in TPS is not just a numbers game.

💰 Cost
Serve more users on the same budget
⚡ Speed
Respond faster on the same server
🌱 Survival
A small team's odds of survival grow

Two wars

VIDraft's recent moves make this scene even more interesting.

Not long ago, Darwin-398B-JGOS scored 90.9% on GPQA Diamond. Single greedy decoding, temperature 0, single sampling — close to seeing a problem once and answering once in an exam hall.

What GPQA Diamond 90.9% says
"We can build models that reason through hard problems."
What the Fast Gemma #1 says
"We can run AI efficiently even on limited hardware."

Making a good model is a battle of research.
Running a good model cheaply and fast is a battle of product.

A small team winning with the same hardware

Honestly, a little national pride wells up here.

Because a small Korean AI startup has earned the right to say this on a global leaderboard.

"We didn't use a bigger GPU. We used the same GPU better."

This sentence is quiet, but strong.

The less you have, the sharper the method must be.
The smaller the hardware, the more precise the design.
The scarcer the resources, the more honest the engineering.

The real value of this #1
The attitude hidden behind the number 505.42.
Even a small team can fight at the world's frontier —
not with bigger hardware, but with a better method.
And this time, that small team came out of Korea.

More Reading

On-DevicePOCKET2026.07

A 35B AI just moved into your pocket — an iPhone

POCKET runs a 35B model on a GPU-less PC and on a phone with stock llama.cpp — no fork. Measured against Bonsai on the same machine: 2.69× faster generation on CPU, 2.22× on GPU, quality tied. And we published the one row where Bonsai wins.

QuantumQuantumOS2026.07

Why "a quantum computer broke encryption" is almost never true

A plain-language walk through our arXiv paper: Even-Mansour period recovery pushed from N=4 to N=10 on real IBM hardware — and, just as importantly, the six limits the authors wrote down themselves. It does not break AES or RSA, and the paper says so first.

K-AI Leaderboard Darwin Platform 2026.06

11 in the Top 20 — Darwin Platform Takes Over the K-AI Leaderboard

11 models generated and derived from the Darwin Platform are in the K-AI Leaderboard Top 20. This isn't the lucky success of a single model — it's a signal that a system for repeatedly building and deriving models is working.

11
in Top 20
1M+
HF downloads
11
patents filed
Inference Acceleration VKAE 2026.06

Building One More GPU — Without Building a GPU

Inference acceleration extracts more work from the same chip through software. VKAE (VIDRAFT Kernel Acceleration Engine) recorded up to 23.4× higher throughput than the conventional approach on the same GPU. Fast, with quality intact.

23.4×
B200 throughput
10K+
tok/s (batched)
VKAE Leaderboard·Demo →
Security AI CHITOS 2026.06

It Doesn't Stop at 'Suspecting' a Vulnerability — CHITOS Autonomous Security AI

Chitos doesn't drop a finding straight into the report. Within authorized scope, it verifies directly, confirms reproducibility, and links attack paths into a kill chain. 3 stages: static scan → threat research → active exploit verification.

Test environment: http://demo.testfire.net (IBM-authorized vulnerable demo site)
Try CHITOS free →
Ubiquitous AI VKUE 2026.07

Running a 35B Model on CPU Alone — No GPU

A sparse MoE keeps only ~3B of 34.7B parameters active, so frontier-class reasoning runs on a single consumer card, a laptop, or a GPU-less CPU server. The VKUE story of lowering the barrier to AI.

~3B
active of 34.7B
0 GPU
runs on CPU
Read on Brunch →
✍️

More Insights

AI insights · technology philosophy · industry analysis from VIDraft's CEO

Read more on Brunch @seawolf →
🧬 CHIMERA · Model Crossbreeding

Combine models without losing what made them good

Chimera merges and distills open models into stronger ones — verified with no quality loss, and independently confirmed at ICML.

What makes Chimera possible

Chimera joins models from different architecture families. That is normally impossible without retraining — tensor shapes and functional roles do not line up. Our filing on compatibility-scored heterogeneous crossbreeding (29 claims) is what makes it a procedure rather than a lucky accident: every tensor pair is scored, then sorted into what can be crossbred and what can only be transplanted.

Filed 2026-05-08 · part of the 12-application portfolio (317 claims)

Chimera = fusing several already-good AIs — without discarding their abilities — into one stronger model. VIDRAFT’s model crossbreeding & merging technology.

How Chimera is built

Every layer of an LLM is Attention + FFN (its knowledge). Chimera designs the attention itself and transplants another model’s FFN 100% intact — and it supports Mixture-of-Experts.

① Design the attention · transplant the knowledge (FFN) Attention FFN knowledge Standard LLM layer Attention VIDRAFT designs it FFN transplanted 100% Chimera layer another model’s FFN 100% ② MoE supported (Mixture of Experts) Attention · shared, VIDRAFT-designed Router FFN Model A FFN Model B FFN Model C Each expert = another model’s FFN · the router picks per token

Conceptual overview only — the internal attention design and gating remain proprietary.

Two verified results

Only measured facts. Internal recipes stay proprietary; the results are public.

🧩
✅ Verified · lossless merge

“Frozen-Expert” merge — zero degradation

Mixing parent models like paint dilutes each strength. Instead we freeze each parent’s core intact and train only a small router that picks the best parent per question.

Result: a 2-expert Chimera (4B) kept GSM8K math at 72.5% — identical to the parent (degradation 0.0). Same-family parents merge to a neutral gain; the real lift comes from cross-family breeding plus the method on the right.

🔎
✅ Verified · at ICML ⭐

RFT edge-filter — learn the right difficulty

The way to get smarter is to train only on just-right-difficulty problems — dropping the too-easy (already known) and the too-hard (never solved), focusing on the ones solvable with a bit more effort.

Result: Darwin-V9-Chimera-4B lifted Korean KMMLU (6 subjects) from 42.9% to 48.3% (+5.4%p), with no forgetting. The principle — data selection over hyperparameters — was independently confirmed by ICML 2026 work. Measured on a 240-item sample; larger-eval confirmation in progress.

Would combining the best AIs make them stronger? Usually, blending them blurs each one’s strengths. Chimera keeps each model’s strengths fully intact and lets a small “conductor” pick the right expert per problem — so ability never collapses. Add VIDRAFT’s “learn only the right-difficulty problems” method, and skill rises without forgetting (+5.4%p on a Korean reasoning benchmark) — a principle independently backed by ICML 2026 research.

🔒 What we show, and what we keep

We publish the concepts and the measured results. The internal merge recipes, filter thresholds and gate formulas stay proprietary. Every number here is measured — no inflated claims.

🔬 Metacognition research →

A related research thread — experts cross-checking each other to reduce hallucination — lives under AETHER. Open the AETHER page to read more.

Chimera로 만든 공개 모델

합체로 만든 개체를 자체 정제 파이프라인으로 끌어올린 결과물입니다. 가중치와 수치를 공개해 누구나 재현할 수 있습니다.

🧬
Model · 4B · Open weights

Darwin-4B-Chimera

Chimera 4B 계보를 자체 정제해 한국어 지식 추론을 끌어올린 모델. KMMLU 6과목 240문항 held-out(greedy)에서 42.9% → 48.3% (+5.4%p).

KMMLU 48.3% 4.02B greedy · 240 held-out

※ 240문항 기준 표준오차 약 ±1.4%p — 더 큰 셋에서 재확인 필요. KMMLU 6과목 한정 수치이며 전 도메인 일반화를 뜻하지 않습니다. 4B 모델로 플래그십과 직접 비교 대상이 아닙니다. 합체·정제의 내부 설계는 비공개(결과만 공개).

🧬 DARWIN MODEL CATALOG

Darwin Model Universe

Every officially released Darwin model and its community derivatives — counted by cumulative all-time downloads.

All-time cumulative downloads · HuggingFace shows only the last 30 days by default · updated

Official Models

Released on VIDraft & FINAL-Bench.

Community Models

Quantizations, fine-tunes & derivatives built by the community on Darwin.

Other Models

Notable non-Darwin models released by VIDraft & FINAL-Bench.

🐳 Docker Images

Production-ready containers on Docker Hub (vidraft).

Live Demos

Interactive demos on Hugging Face Spaces.

📊 Datasets

Benchmarks & datasets published by VIDraft & FINAL-Bench.

1 / 23
← → 방향키 · 스페이스 · 스와이프로 이동 · F 전체화면 · O 전체보기

전체 슬라이드