I will introduce the Physical AI content using Unitree Go2 exhibited at DevelopersIO 2026 Osaka #devio2026

I will introduce the Physical AI content using Unitree Go2 exhibited at DevelopersIO 2026 Osaka #devio2026

Introducing the Physical AI exhibit using Unitree Go2 shown at DevelopersIO 2026 Osaka. Go2 autonomously decides movements and responses from speech and camera input, displaying its decision process.
2026.10.07

This page has been translated by machine translation. View original

Introduction

I'm Fujii (Da) from the Manufacturing Business Technology Department.

At DevelopersIO 2026 Osaka, held on September 28-29, 2026, we exhibited using a Unitree Go2.

On Day 2, I gave a presentation titled "We tried Physical AI after the Unitree Go2 came to the Osaka office," introducing the background of how the Unitree Go2 came to the Osaka office, the content of the exhibit, its architecture, and the development process. The materials are published in the following article.

https://dev.classmethod.jp/articles/osaka-office-unitree-go2-joined-physical-ai-developersio-2026-osaka-day2-devio2026/

This article explains the exhibit content.

Exhibit Overview

When visitors spoke to the Go2, it would understand their words and respond with appropriate movements and conversational replies.

For example, when someone said "greet me" or "hello," it would understand that it had been greeted and respond with "I'll greet you!" while performing the Hello movement.

It could also explain what was in view by saying "This is ~~" based on images captured by its camera when asked "What is this?"

A screen was placed next to the Go2 to display the words it heard, the movement it chose, Go2's response, and what was visible on camera.

Speech recognition, intent determination, image-based response processing, and speech synthesis were all handled by a single DGX Spark placed at the venue. No cloud AI services were used.

Exhibit scene

What We Wanted to Show With This Exhibit

Movements like Hello and Dance can be operated by a person using the Unitree smartphone app or the included remote controller.
However, simply showing a person operating it would just be an introduction to the Unitree Go2 itself.

So for this exhibit, we decided to demonstrate Physical AI — where Go2 itself, rather than a human, understands the content of spoken words and what appears on camera, and then decides how to move and what to say in response.

We also insisted on making it a local LLM/edge AI where inference is completed entirely within the local DGX Spark, without relying on the cloud.

Exhibit Details

Interacting With Go2

When a visitor approaches Go2, it turns toward that person and greets them with phrases like "Oh, hello" or "Could you look over here?"
If no one speaks to it for a while after calling out, it says something like "Well, that's fine" or "Feel free to talk to me whenever you'd like," and returns to a waiting state.

Since Go2 stays on standby with its microphone on, visitors can speak to it at any time (it sometimes picked up surrounding voices and responded on its own).
When spoken to with content that instructs a specific movement, such as "greet me" or "dance," Go2 understands the meaning and performs the movement.
The movements supported at the time of the exhibit are listed in the table below (the movements themselves are pre-defined).

Movement Description
Hello Raises and waves a front leg as a greeting. Also used for "shake hands"
FingerHeart Forms a heart shape with its front legs
Stretch Stretches
Sit Sits down. Stands back up on its own afterward
Dance1 Dances
Dance2 Dances with a different choreography from Dance1.
When simply told "dance," it chooses either Dance1 or Dance2
Scrape Raises its upper body and swipes the air with a front leg
Euler Tilts its body left-right, then front-back in sequence

While performing movements, it announces the chosen movement verbally with phrases like "I'll make a heart" or "I'll sit down."

For words that don't lead to a movement, it tilts its body once as if cocking its head, then generates an appropriate response on the spot.
When it cannot generate a response, it returns a fixed phrase like "Hmm, I didn't understand."
When speech is cut off mid-sentence or too short, it asks again with "Could you please repeat that?"

When a visitor holds an object up to the camera and asks something like "What is this?", Go2 understands what it sees and answers.
When it cannot answer from what it sees, it returns a response like "I couldn't see it clearly."

What Was Displayed on Screen

Screen when spoken to

In the center of the screen, three items were arranged from top to bottom: "Heard Text," "Action," and "Response."

  • "Heard Text": The words picked up by the microphone
  • "Action": The movement Go2 chose. When words don't lead to a movement, it displays "None of the above"; when asked a question while being shown an object, it displays "Answer by looking"
  • "Response": The words Go2 spoke. When asked a question while being shown an object, the camera image used for the answer is also displayed

When spoken to, "Heard Text" appears first, followed shortly by "Action."
Since the display changes in this order, it is clear that Go2 listens first and then makes a decision.

On the right side of the screen, the Go2's camera feed and a 3D model of Go2 were displayed.
The 3D model overlays the temperature of each part using colors — blue for low and red for high.
This was repurposed from something I had previously written about on the blog.

https://dev.classmethod.jp/articles/dafujii-robot-urdf-motor-temperature-mapping/

At the top and bottom of the screen, Go2's status, remaining battery level, and the number of times it moved that day were displayed.

Equipment and Jigs Mounted on Go2

To realize the exhibit content described above, additional equipment was mounted on Go2's back.
Although Go2 has a built-in microphone and speaker, I chose to use external ones to capture the direction of sound and to prevent the built-in microphone from picking up Go2's own voice while speaking.

However, as development progressed, things didn't go as originally planned — I ended up detecting the person's location using only the camera rather than sound direction, and switched to half-duplex because the microphone kept picking up Go2's own voice during speech.

Equipment on the back

Equipment Role
RealSense D435i Captures person's position and distance. Also used for images when a visitor shows an object and asks a question. Mounted on the head
ReSpeaker XVF3800 Microphone array. Picks up visitors' voices. Can also detect sound direction
Speaker Outputs Go2's voice
Mobile battery Powers the equipment and USB hub
USB hub Connects equipment to the expansion dock
Wireless adapter Connects the expansion dock to Wi-Fi

The equipment on the back is secured to the expansion dock's rail using jigs made with a 3D printer.
Initially, the microphone array was placed directly above the expansion dock, but since it picked up the dock's fan noise and couldn't capture voices properly, a jig was used to raise the microphone array to a higher position.

Closing

Thank you to everyone who visited.

Although I announced there would be an exhibit, the power had to be turned off during the presentation due to venue constraints, so I think most people only had a chance to interact with it during the Day 1 social gathering.
I hope this article conveys what the exhibit was like even to those who weren't able to experience it.

The architecture built for the exhibit is covered in the presentation materials, so please also check out the following article.

https://dev.classmethod.jp/articles/osaka-office-unitree-go2-joined-physical-ai-developersio-2026-osaka-day2-devio2026/

Share this article

カジュアル面談受付中

Related articles