Field notes · Hugging Face SO-101 · LeRobot
In a few evenings I took an SO-101 robot arm kit from a box of parts to an arm that dances, reacts to a thumbs-up, and pushes a tomato into a hole. This post walks through the steps in plain terms, for anyone new to robotics.
The SO-101 is an open-source, 3D-printed arm from Hugging Face’s LeRobot project. It comes as a pair.
| Part | What it is |
|---|---|
| Follower arm | The robot that does the work. Six motors: base, shoulder, elbow, wrist flex, wrist roll, gripper. |
| Leader arm | A copy you move by hand. The follower mirrors it, which is how you control the robot live. |
| Servo | A motor that knows its own position and can hold any angle you send it. |
| Motor ID | Each servo’s address on a shared cable, 1 to 6, so the computer can talk to one at a time. |
| Controller board | Connects the servo cable to your laptop by USB. |
| Calibration | Teaching the software where each joint’s middle and limits are. |
uv.
Kinesthetic teaching means showing the robot a task by moving it with your own hands. I switched the follower’s motors off, guided it through the motion, and a script recorded every joint about 27 times a second. Then the arm played the recording back by itself.
It’s simple and needs no AI, but the arm repeats exactly one motion. It doesn’t see the object, so the object has to be in the same place every time.
My task was pushing a tomato into the cable hole in my desk. It took six recordings. The first three missed because the elbow stopped short: my calibration had only swept it halfway, so the servo refused to go further on replay. After recalibrating, the winner was one short, smooth push with the tomato always on the same mark, because a replay can’t adapt to a moved tomato.
A pre-scripted motion is written in code instead of shown by hand. I listed a few poses for the arm to hit, such as arm up, sway left, sway right and nod, and a script moved it between them.
Compared with kinesthetic teaching, nobody has to demonstrate anything and every run is identical. But it’s still blind: the arm follows the script whatever is in front of it.
The one real AI model in the project is a camera trigger: give a thumbs-up and the arm grips three times. I didn’t film this one, but in a 90-second test it reacted to all six of my thumbs-ups.
This is inference: running a model someone else already trained. Google’s MediaPipe gesture recognizer reads each camera frame on the laptop’s CPU, about 30 frames a second, and labels the hand gesture. My code fires the grip after three thumbs-up frames in a row, then waits for the hand to drop before counting again.
The same idea scales up. Swap the gesture model for a vision-language model that looks at the table and decides what to pick up, and you have a robot that follows spoken instructions, still without training anything yourself.