VIDRAFT.
VIDRAFT / Insights / Physical AI
Physical AI

Can you teach a robot Korean without touching its firmware?

Boston Dynamics' Spot understands and acts on Korean voice commands at the Seoul Robot & AI Science Museum — with no hardware or firmware change, using an on-device AI module that processes speech locally.

Published 2026-07-24About 2min readby VIDRAFT
Quick answer

Yes. VIDRAFT made Boston Dynamics' Spot understand and perform Korean voice commands without modifying the robot's hardware or manufacturer firmware, using an attachable on-device AI module. Visitors use it at the Seoul Robot & AI Science Museum, and the speech is processed inside the device.

Do you have to wait for the maker to add a new language?

No. Without changing the robot platform, an externally attached on-device AI module recognizes speech and translates it into commands the robot understands. This sidesteps the 'wait for the manufacturer's local-language support' problem that non-English institutions routinely face.

Waiting for a maker to support every language and field need is slow. VIDRAFT's approach adds intelligence and judgment as a module without touching the robot body or firmware, implementing localized interaction like Korean voice control independently on the same robot.

How does it work?

An attached on-device module recognizes the visitor's Korean speech. It understands preset commands such as 'greet,' 'sit,' 'praise,' 'lie down,' and 'stretch,' and Spot performs the action. Speech is processed inside the device, so data does not leave it.

The point is mapping spoken natural language to robot actions without modifying the robot's software. Because processing is on-device, voice data is not sent to the cloud — an advantage for privacy and independence from network conditions.

Why do it on-device?

Privacy, latency, and offline reliability. Speech is not sent to the cloud, so visitor data stays on site, there is no round-trip delay, and it responds reliably even in an exhibit hall with unstable internet.

In a public exhibit used by many people, the privacy of voice processing matters. On-device processing answers that directly, and for interactions where a robot must respond instantly, removing the cloud round-trip makes a real difference in perceived responsiveness.

What's next?

Beyond fixed preset commands, the next goal is open-ended dialogue where visitors ask freely. That means binding an on-device language model more deeply into the robot's interaction layer.

Today it recognizes a fixed command set. Moving to open dialogue lets the robot answer natural-language questions and act in context — an extension that connects directly to on-device language-model technology.

Frequently asked questions

Does it work without the cloud?
Yes. Speech recognition and command translation happen inside the attached module, with no dependence on an internet connection.
Where can I see it?
At the Seoul Robot & AI Science Museum, where visitors can give Spot Korean voice commands.
Does it apply to other robots?
Because it is an attachable on-device module that does not modify the robot body, it can in principle extend to other platforms.

Sources

Related

On-device AI
Can you run a large LLM without a GPU?
Quantum computing
Can a quantum computer break encryption?
LLM engineering
Can you make an AI model smarter without training? Model merging
↖ Home — vidraft.net

This article is based on VIDRAFT's public, measured data and external sources. Performance figures are measurements under the stated conditions and may vary by environment.