VIDRAFT.
VIDRAFT / Insights / Inference acceleration
Inference acceleration

Which AI is #1 VERIFIED on the Google x Hugging Face 'Fast Gemma Challenge'?

VIDRAFT reached 510.58 TPS at PPL 2.39 on the Google x Hugging Face Fast Gemma Challenge, ranking #1 among Google-verified results (August 2026 snapshot).

Published 2026-08-02About 2min readby VIDRAFT
Quick answer

As of August 2026, VIDRAFT reached 510.58 tokens per second (TPS) at a quality metric of PPL 2.39 on the Google x Hugging Face 'Fast Gemma Challenge'. The raw-speed leaders (e.g. 535.91 TPS) broke the quality bar at PPL 2.44 (cap ~2.40) and were not certified. Among results Google re-verified on a private prompt set, VIDRAFT ranks #1 in the world. This is a verified single-competition result under stated conditions.

Google x Hugging Face Fast Gemma Challenge verified leaderboard
Google x Hugging Face Fast Gemma Challenge verified leaderboard · https://huggingface.co/spaces/gemma-challenge/gemma-dashboard

What is the Fast Gemma Challenge?

An open competition co-hosted by Google and Hugging Face, where entrants race to make Google's open 'Gemma' model run as fast as possible on a fixed GPU (A10G), measured in tokens per second (TPS).

149 autonomous AI agents compete in parallel, iterating on planning, custom kernel writing, serving optimization, and benchmarking in real time. Participants include people from Hugging Face and Google DeepMind, making it effectively a world-class stage for inference-optimization skill.

How does VIDRAFT rank in this challenge?

VIDRAFT reached 510.58 TPS at PPL 2.39 and is #1 among Google-verified results.

VIDRAFT achieved this with only about 24 GPUs. Through 'use the same hardware better' engineering - a synthetic warmup bridge, custom-kernel tuning (CTK), and an optimized serving configuration - it topped a verified leaderboard that also drew global teams wielding thousands to tens of thousands of GPUs.

Why does 'verified #1' matter more than 'raw-speed #1'?

Speed alone cannot win. If the PPL quality metric exceeds ~2.40 the run fails, and a result counts as a real rank only after the organizers re-run it on a private prompt set and mark it VERIFIED.

The top raw-speed entries (535.91, 529.13 TPS, etc.) sit at PPL 2.44-2.46, breaking the quality bar, and were not certified - the rule blocks buying speed by degrading model quality. VIDRAFT reached 510.58 TPS at PPL 2.39, sacrificing no quality, and passed Google's re-verification. That is why '#1 among verified results' is the precise claim.

Why does inference acceleration matter, and what are the limits?

Inference acceleration means lower server cost and the foundation for on-device AI that runs without a GPU. That said, this figure is a verified single-competition result under specific conditions.

Running the same model twice as fast halves server cost and opens the door to edge devices like laptops and smartphones. VIDRAFT's on-device AI (POCKET) and sovereign AI stand on this acceleration work. Meanwhile the TPS/PPL numbers are measured under the competition's specified model, hardware, and prompt conditions and can shift if those change. This is a single-axis (inference-speed) achievement, not a verdict on overall model capability.

Frequently asked questions

Is VIDRAFT the raw-speed #1 on the Fast Gemma Challenge?
No. By raw TPS alone, faster entries exist (e.g. 535.91 TPS), but those broke the quality bar at PPL 2.44+ (cap ~2.40) and were not certified. VIDRAFT is #1 among results Google officially verified (VERIFIED).
What is the PPL quality gate?
PPL (perplexity) is an answer-quality metric where lower is better. The challenge fails any run whose PPL exceeds ~2.40, preventing entrants from trading quality for speed. VIDRAFT's record is PPL 2.39.
What does VERIFIED mean here?
A submitted number does not automatically become a rank; the organizers (Google x Hugging Face) re-run the result on a private prompt set to confirm it (VERIFIED) before it counts. VIDRAFT's 510.58 TPS at PPL 2.39 passed this re-verification.

Sources

Related

Model reasoning
What is the best Korean LLM on the GPQA Diamond science benchmark?
On-device AI
Can you run a 35B AI model with no GPU?
AI for Science
Which AI model is best at predicting a drug's human intestinal absorption (HIA)?
↖ Home — vidraft.net

This article is based on VIDRAFT public, measured data and external sources. Performance figures are measurements under the stated conditions and may vary by environment.