Embodied AI · Learning from Demonstration

Robot policies learned from VR demonstration, not months of teleop.

SIBNIK turns VR teleoperation of real hardware into structured training data for robot learning — collecting diverse demonstrations at near-zero marginal cost per scene variation.

Early-stage — building and validating the pipeline on a Franka Panda arm
Why

Demonstration data is the bottleneck in robot learning

Modern imitation-learning policies (ACT, Diffusion Policy, VLAs) are only as good as the demonstrations they're trained on — and collecting those demonstrations on real hardware is slow, expensive, and hard to scale across scene variation.

Real hardware, VR interface

A human teleoperates a real robot arm through a VR headset and controllers, giving policies demonstrations grounded in true physical dynamics — not sim-only approximations.

A dedicated interface layer

A software layer handles interaction, kinematics, and scene variation; the robot handles reality. That split is what makes scaling demonstration collection tractable.

Standard, reusable datasets

Every session is captured in LeRobot format — multi-camera RGB, joint states, and actions — ready to drop into ACT, SmolVLA, or other imitation-learning pipelines.

Safety-first on real hardware

Joint limits, slew-rate limiting, a deadman switch, and a watchdog sit between the VR input and the arm at all times.

How it works

From VR demonstration to a trained policy

The current pipeline runs end-to-end on a Franka Panda 7-DOF arm.

Teleoperate in VR

A custom VR front end drives the Franka Panda with full kinematics, inverse kinematics, and an intuitive manipulation gizmo for precise control.

Bridge to the real arm

A custom MuJoCo bridge over ZeroMQ streams commands to the physical robot in real time.

Capture multi-camera observations

Two static RealSense cameras record synchronized RGB views alongside joint state and action data.

Validate with open-loop replay

Every recorded episode is replayed open-loop as a mandatory gate before it's counted toward a training set.

Fine-tune the policy

Episodes are assembled into LeRobot-format datasets and used to fine-tune SmolVLA / ACT policies, starting from a target of 50 demonstrations per task.

MuJoCo ZeroMQ Franka Panda RealSense D435 LeRobot SmolVLA ACT
Where this goes

Toward humanoid embodiment

The near-zero marginal cost of generating a new scene variation in software is the whole point: it's what lets demonstration collection scale the way robot learning needs it to. The Franka Panda is the proving ground — humanoid robotics is the target embodiment.

Team

Founder

SR

Salar Rezayani

Founder, SIBNIK

PhD student in robotics at the University of Waterloo, working out of RoboHub. Roughly three years in industry as a VR/XR developer, spanning VR, multiplayer, ML integration, and advanced rendering. SIBNIK is built independently — outside of lab time and lab equipment.

Get in touch

Interested in the approach, collaborating, or just want to talk robot learning.

[email protected]