#
Building a Robot Teleoperation System for Data Collection
If you're training robot policies with imitation learning, the quality of your dataset almost entirely depends on how good your teleoperation system is. A clunky, high-latency teleop rig produces jerky, inconsistent demonstrations — and your policy will happily learn those bad habits. This tutorial walks through designing and building a teleoperation system purpose-built for collecting clean, high-frequency demonstration data.
#
What "good" teleoperation looks like for data collection
Teleoperation for entertainment or remote inspection has different priorities than teleoperation for data collection. For ML data collection, you specifically need:
Low, consistent latency — variable lag introduces noise that looks like intentional motion to a learning algorithm. #
High-frequency, synchronized logging — every joint command needs a timestamp that lines up with camera frames and sensor readings. #
Smooth, continuous control — discrete or bang-bang inputs (keyboard-style) produce trajectories that don't resemble natural human motion. #
Repeatability — the same operator should be able to produce similar trajectories across many trials, which matters for later behavior cloning.
#
System architecture
A typical teleoperation-for-data-collection stack has four layers: #
Input device layer — leader arm, joystick, VR controller, or a haptic device that produces a continuous control signal. #
Mapping layer — converts the input device's pose/state into a target pose or joint command for the follower robot (this is where retargeting happens if the input and robot have different kinematics). #
Control layer — a low-level controller (impedance, PD, or inverse kinematics solver) that turns target commands into actuator signals safely. #
Logging layer — records synchronized streams: joint states, end-effector pose, camera frames, gripper state, and the raw input signal.
#
Choosing your input device
Common choices, roughly in order of setup complexity:
Leader-follower arm pairs (e.g., a low-cost duplicate arm you move by hand) — gives the most natural kinesthetic feel and near 1:1 kinematic mapping. #
3D mouse / SpaceMouse — cheap, precise for 6-DOF end-effector control, but less intuitive for beginners. #
Gamepad — good for mobile base + simple arm tasks, poor for fine manipulation. #
VR controllers — excellent for full 6-DOF pose control with intuitive hand tracking (covered in depth in a separate tutorial in this series).
For manipulation tasks specifically, leader-follower arms and VR controllers tend to produce the cleanest demonstrations because the operator's hand motion maps almost directly onto the desired end-effector motion.
#
Implementing the control loop
Here's a minimal Python control loop structure using an end-effector pose target and a simple IK-based controller:
A few practical notes on this loop:
- Run it at a fixed control rate (20–50 Hz is typical) and log the actual elapsed time per iteration, not just the nominal timestamp — drift matters for imitation learning.
- Apply a low-pass filter to the raw input signal before mapping it to a target pose. Human hand tremor and sensor noise both show up as high-frequency jitter that hurts downstream policy training.
- Always clamp the target pose to safe workspace bounds before sending it to the IK solver, not after — this avoids solver instability near joint limits.
#
Synchronizing multiple data streams
The hardest part of a teleop data collection system usually isn't the control loop — it's keeping the camera stream, joint state stream, and input stream synchronized. A few approaches:
Single-threaded polling loop (shown above) — simplest, works well if your camera and robot APIs are fast enough to poll synchronously. #
Multi-threaded with timestamp alignment — each sensor runs on its own thread/process and pushes timestamped samples into a shared buffer; a separate aligner thread matches samples by nearest timestamp. Necessary once you add multiple cameras or higher-rate sensors like force-torque. #
Hardware trigger sync — for research-grade setups, a hardware trigger line pulses all sensors simultaneously. Overkill for most hobbyist or startup projects, but worth knowing about if you're chasing sub-millisecond alignment.
For most imitation learning use cases, multi-threaded polling with timestamp-based nearest-neighbor alignment is a good default — it's robust and doesn't require special hardware.
#
Safety layers you shouldn't skip
Because a human is now directly driving a robot arm in real time, add these guardrails regardless of how "just for data collection" the setup feels:
- Workspace bounding box enforced in software, independent of the IK solver.
- A velocity limiter on the mapped target pose to prevent sudden jumps if the input device glitches.
- A dead-man's switch (physical button or trigger) that must be held to enable motion — releasing it should freeze or gently stop the robot.
- An emergency stop that cuts power at the hardware level, not just a software flag.
#
Putting it together
Once your loop is running reliably, wrap each demonstration collection session with clear episode boundaries — a "start recording" and "end recording" signal (often a button press) so each demonstration becomes a discrete, labeled trajectory rather than one continuous, ambiguous stream. This episodic structure is exactly what you'll need in the next step of the pipeline: capturing human demonstrations for imitation learning.
#
Useful Links
Website: [www.v-modal.com](https://www.v-modal.com)
SDK Flutter: [v-modal/vmodal_sdk_flutter](https://github.com/v-modal/vmodal_sdk_flutter)
SDK Android: [v-modal/vmodal_sdk_android](https://github.com/v-modal/vmodal_sdk_android)
Discord: [https://discord.gg/K72z28KUx](https://discord.gg/K72z28KUx)