
I will introduce the Physical AI content using Unitree Go2 exhibited at DevelopersIO 2026 Osaka #devio2026
This page has been translated by machine translation. View original
Introduction
I'm Fujii (Da) from the Manufacturing Business Technology Department.
At DevelopersIO 2026 Osaka, held on September 28-29, 2026, we exhibited using a Unitree Go2.
On Day 2, I gave a presentation titled "We tried Physical AI after the Unitree Go2 came to the Osaka office," introducing the background of how the Unitree Go2 came to the Osaka office, the content of the exhibit, its architecture, and the development process. The materials are published in the following article.
This article explains the exhibit content.
Exhibit Overview
When visitors spoke to the Go2, it would understand their words and respond with appropriate movements and conversational replies.
For example, when someone said "greet me" or "hello," it would understand that it had been greeted and respond with "I'll greet you!" while performing the Hello movement.
It could also explain what was in view by saying "This is ~~" based on images captured by its camera when asked "What is this?"
A screen was placed next to the Go2 to display the words it heard, the movement it chose, Go2's response, and what was visible on camera.
Speech recognition, intent determination, image-based response processing, and speech synthesis were all handled by a single DGX Spark placed at the venue. No cloud AI services were used.

What We Wanted to Show With This Exhibit
Movements like Hello and Dance can be operated by a person using the Unitree smartphone app or the included remote controller.
However, simply showing a person operating it would just be an introduction to the Unitree Go2 itself.
So for this exhibit, we decided to demonstrate Physical AI — where Go2 itself, rather than a human, understands the content of spoken words and what appears on camera, and then decides how to move and what to say in response.
We also insisted on making it a local LLM/edge AI where inference is completed entirely within the local DGX Spark, without relying on the cloud.
Exhibit Details
Interacting With Go2
When a visitor approaches Go2, it turns toward that person and greets them with phrases like "Oh, hello" or "Could you look over here?"
If no one speaks to it for a while after calling out, it says something like "Well, that's fine" or "Feel free to talk to me whenever you'd like," and returns to a waiting state.
Since Go2 stays on standby with its microphone on, visitors can speak to it at any time (it sometimes picked up surrounding voices and responded on its own).
When spoken to with content that instructs a specific movement, such as "greet me" or "dance," Go2 understands the meaning and performs the movement.
The movements supported at the time of the exhibit are listed in the table below (the movements themselves are pre-defined).
| Movement | Description |
|---|---|
Hello |
Raises and waves a front leg as a greeting. Also used for "shake hands" |
FingerHeart |
Forms a heart shape with its front legs |
Stretch |
Stretches |
Sit |
Sits down. Stands back up on its own afterward |
Dance1 |
Dances |
Dance2 |
Dances with a different choreography from Dance1.When simply told "dance," it chooses either Dance1 or Dance2 |
Scrape |
Raises its upper body and swipes the air with a front leg |
Euler |
Tilts its body left-right, then front-back in sequence |
While performing movements, it announces the chosen movement verbally with phrases like "I'll make a heart" or "I'll sit down."
For words that don't lead to a movement, it tilts its body once as if cocking its head, then generates an appropriate response on the spot.
When it cannot generate a response, it returns a fixed phrase like "Hmm, I didn't understand."
When speech is cut off mid-sentence or too short, it asks again with "Could you please repeat that?"
When a visitor holds an object up to the camera and asks something like "What is this?", Go2 understands what it sees and answers.
When it cannot answer from what it sees, it returns a response like "I couldn't see it clearly."
What Was Displayed on Screen

In the center of the screen, three items were arranged from top to bottom: "Heard Text," "Action," and "Response."
- "Heard Text": The words picked up by the microphone
- "Action": The movement Go2 chose. When words don't lead to a movement, it displays "None of the above"; when asked a question while being shown an object, it displays "Answer by looking"
- "Response": The words Go2 spoke. When asked a question while being shown an object, the camera image used for the answer is also displayed
When spoken to, "Heard Text" appears first, followed shortly by "Action."
Since the display changes in this order, it is clear that Go2 listens first and then makes a decision.
On the right side of the screen, the Go2's camera feed and a 3D model of Go2 were displayed.
The 3D model overlays the temperature of each part using colors — blue for low and red for high.
This was repurposed from something I had previously written about on the blog.
At the top and bottom of the screen, Go2's status, remaining battery level, and the number of times it moved that day were displayed.
Equipment and Jigs Mounted on Go2
To realize the exhibit content described above, additional equipment was mounted on Go2's back.
Although Go2 has a built-in microphone and speaker, I chose to use external ones to capture the direction of sound and to prevent the built-in microphone from picking up Go2's own voice while speaking.
However, as development progressed, things didn't go as originally planned — I ended up detecting the person's location using only the camera rather than sound direction, and switched to half-duplex because the microphone kept picking up Go2's own voice during speech.

| Equipment | Role |
|---|---|
| RealSense D435i | Captures person's position and distance. Also used for images when a visitor shows an object and asks a question. Mounted on the head |
| ReSpeaker XVF3800 | Microphone array. Picks up visitors' voices. Can also detect sound direction |
| Speaker | Outputs Go2's voice |
| Mobile battery | Powers the equipment and USB hub |
| USB hub | Connects equipment to the expansion dock |
| Wireless adapter | Connects the expansion dock to Wi-Fi |
The equipment on the back is secured to the expansion dock's rail using jigs made with a 3D printer.
Initially, the microphone array was placed directly above the expansion dock, but since it picked up the dock's fan noise and couldn't capture voices properly, a jig was used to raise the microphone array to a higher position.
Closing
Thank you to everyone who visited.
Although I announced there would be an exhibit, the power had to be turned off during the presentation due to venue constraints, so I think most people only had a chance to interact with it during the Day 1 social gathering.
I hope this article conveys what the exhibit was like even to those who weren't able to experience it.
The architecture built for the exhibit is covered in the presentation materials, so please also check out the following article.




