Papers
Topics
Authors
Recent
Search
2000 character limit reached

Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots

Published 15 Feb 2024 in cs.RO | (2402.10329v3)

Abstract: We present Universal Manipulation Interface (UMI) -- a data collection and policy learning framework that allows direct skill transfer from in-the-wild human demonstrations to deployable robot policies. UMI employs hand-held grippers coupled with careful interface design to enable portable, low-cost, and information-rich data collection for challenging bimanual and dynamic manipulation demonstrations. To facilitate deployable policy learning, UMI incorporates a carefully designed policy interface with inference-time latency matching and a relative-trajectory action representation. The resulting learned policies are hardware-agnostic and deployable across multiple robot platforms. Equipped with these features, UMI framework unlocks new robot manipulation capabilities, allowing zero-shot generalizable dynamic, bimanual, precise, and long-horizon behaviors, by only changing the training data for each task. We demonstrate UMI's versatility and efficacy with comprehensive real-world experiments, where policies learned via UMI zero-shot generalize to novel environments and objects when trained on diverse human demonstrations. UMI's hardware and software system is open-sourced at https://umi-gripper.github.io.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (53)
  1. Human-to-robot imitation in the wild. In Proceedings of Robotics: Science and Systems (RSS), 2022.
  2. Affordances from human videos as a versatile representation for robotics. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13778–13790, 2023.
  3. Rt-1: Robotics transformer for real-world control at scale. In Proceedings of Robotics: Science and Systems (RSS), 2023.
  4. Humanoid robot teleoperation with vibrotactile based balancing feedback. In Haptics: Neuroscience, Devices, Modeling, and Applications: 9th International Conference, EuroHaptics 2014, Versailles, France, June 24-26, 2014, Proceedings, Part II 9, pages 266–275. Springer, 2014.
  5. The ycb object and model set: Towards common benchmarks for manipulation research. In 2015 International Conference on Advanced Robotics (ICAR), pages 510–517, 2015. doi: 10.1109/ICAR.2015.7251504.
  6. Orb-slam3: An accurate open-source library for visual, visual–inertial, and multimap slam. IEEE Transactions on Robotics, 37(6):1874–1890, 2021a.
  7. Orb-slam3: An accurate open-source library for visual, visual–inertial, and multimap slam. IEEE Transactions on Robotics, 37(6):1874–1890, 2021b. doi: 10.1109/TRO.2021.3075644.
  8. Learning generalizable robotic reward functions from “in-the-wild” human videos. In Proceedings of Robotics: Science and Systems (RSS), 2021.
  9. Diffusion policy: Visuomotor policy learning via action diffusion. In Proceedings of Robotics: Science and Systems (RSS), 2023.
  10. On hand-held grippers and the morphological gap in human manipulation demonstration. arXiv preprint arXiv:2311.01832, 2023.
  11. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations, 2021.
  12. Bridge data: Boosting generalization of robotic skills with cross-domain datasets. In Proceedings of Robotics: Science and Systems (RSS), 2022.
  13. Low-cost exoskeletons for learning whole-arm manipulation in the wild. arXiv preprint arXiv:2309.14975, 2023.
  14. Mobile aloha: Learning bimanual mobile manipulation with low-cost whole-body teleoperation. arXiv preprint arXiv:2401.02117, 2024.
  15. Automatic generation and detection of highly reliable fiducial markers under occlusion. Pattern Recognition, 47(6):2280–2292, 2014. ISSN 0031-3203. doi: https://doi.org/10.1016/j.patcog.2014.01.005. URL https://www.sciencedirect.com/science/article/pii/S0031320314000235.
  16. Deep residual learning for image recognition. corr abs/1512.03385 (2015), 2015.
  17. GoPro Inc. Gpmf introuction: Parser for gpmf™ formatted telemetry data used within gopro® cameras. https://gopro.github.io/gpmf-parser/. Accesssed: 2023-01-31.
  18. Bc-z: Zero-shot task generalization with robotic imitation learning. In Conference on Robot Learning (CoRL), volume 164, pages 991–1002. PMLR, 2022.
  19. Giving robots a hand: Broadening generalization via hand-centric human video demonstrations. In Deep Reinforcement Learning Workshop NeurIPS, 2022.
  20. VIP: Towards universal visual reward and representation via value-implicit pre-training. In The Eleventh International Conference on Learning Representations, 2023.
  21. Roboturk: A crowdsourcing platform for robotic skill learning through imitation. In Conference on Robot Learning (CoRL), volume 87, pages 879–893. PMLR, 2018.
  22. R3m: A universal visual representation for robot manipulation. In Proceedings of The 6th Conference on Robot Learning (CoRL), volume 205, pages 892–909. PMLR, 2022.
  23. Tax-pose: Task-specific cross-pose estimation for robot manipulation. In Proceedings of The 6th Conference on Robot Learning (CoRL), volume 205, pages 1783–1792. PMLR, 2023.
  24. The surprising effectiveness of representation learning for visual imitation. In Proceedings of Robotics: Science and Systems (RSS), 2022.
  25. Learning of compliant human–robot interaction using full-body haptic interface. Advanced Robotics, 27(13):1003–1012, 2013.
  26. Characterizing input methods for human-to-robot demonstrations. In 2019 14th ACM/IEEE International Conference on Human-Robot Interaction (HRI), pages 344–353. IEEE, 2019.
  27. Dexmv: Imitation learning for dexterous manipulation from human videos. In European Conference on Computer Vision, pages 570–587. Springer, 2022.
  28. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PMLR, 2021.
  29. Recent advances in robot learning from demonstration. Annual review of control, robotics, and autonomous systems, 3:297–330, 2020.
  30. Latent plans for task-agnostic offline reinforcement learning. In Proceedings of The 6th Conference on Robot Learning (CoRL), volume 205, pages 1838–1849. PMLR, 2023.
  31. Scalable. intuitive human to robot skill transfer with wearable human machine interfaces: On complex, dexterous tasks. In 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 6318–6325. IEEE, 2023.
  32. Learning predictive models from observation and interaction. In European Conference on Computer Vision, pages 708–725. Springer, 2020.
  33. Reinforcement learning with videos: Combining offline observations with interaction. In Proceedings of the 2020 Conference on Robot Learning (CoRL), volume 155, pages 339–354. PMLR, 2021.
  34. Deep imitation learning for humanoid loco-manipulation through human teleoperation. In 2023 IEEE-RAS 22nd International Conference on Humanoid Robots (Humanoids), pages 1–8. IEEE, 2023.
  35. On bringing robots home. arXiv preprint arXiv:2311.16098, 2023.
  36. Concept2robot: Learning manipulation concepts from instructions and human demonstrations. The International Journal of Robotics Research, 40(12-14):1419–1434, 2021.
  37. Videodex: Learning dexterity from internet videos. In Proceedings of The 6th Conference on Robot Learning (CoRL), volume 205, pages 654–665. PMLR, 2023.
  38. Distilled feature fields enable few-shot language-guided manipulation. In Proceedings of The 7th Conference on Robot Learning (CoRL), volume 229, pages 405–424. PMLR, 2023.
  39. Neural descriptor fields: Se (3)-equivariant object representations for manipulation. In 2022 International Conference on Robotics and Automation (ICRA), pages 6394–6400. IEEE, 2022.
  40. Grasping in the wild: Learning 6dof closed-loop grasping from low-cost demonstrations. Robotics and Automation Letters, 2020.
  41. SEED: Series elastic end effectors in 6d for visuotactile tool use. In 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 4684–4691, 2022. doi: 10.1109/IROS47612.2022.9982092.
  42. A force-sensitive exoskeleton for teleoperation: An application in elderly care robotics. In 2023 IEEE International Conference on Robotics and Automation (ICRA), pages 12624–12630. IEEE, 2023.
  43. Mimicplay: Long-horizon imitation learning by watching human play. In Proceedings of The 7th Conference on Robot Learning (CoRL), volume 229, pages 201–221. PMLR, 2023.
  44. Error-aware imitation learning from teleoperation data for mobile manipulation. In Proceedings of the 5th Conference on Robot Learning (CoRL), volume 164, pages 1367–1378. PMLR, 2022.
  45. GELLO: A general, low-cost, and intuitive teleoperation framework for robot manipulators. In Towards Generalist Robots: Learning Paradigms for Scalable Skill Acquisition @ CoRL2023, 2023.
  46. Towards a personal robotics development platform: Rationale and design of an intrinsically safe personal robot. In 2008 IEEE International Conference on Robotics and Automation, pages 2165–2170. IEEE, 2008.
  47. Masked visual pre-training for motor control. arXiv:2203.06173, 2022.
  48. Learning by watching: Physical imitation of manipulation skills from human videos. In 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 7827–7834. IEEE, 2021.
  49. Visual imitation made easy. In Conference on Robot Learning (CoRL), volume 155, pages 1992–2005. PMLR, 2021.
  50. Deep imitation learning for complex manipulation tasks from virtual reality teleoperation. In 2018 IEEE International Conference on Robotics and Automation (ICRA), pages 5628–5635. IEEE, 2018.
  51. Benefit of large field-of-view cameras for visual odometry. In 2016 IEEE International Conference on Robotics and Automation (ICRA), pages 801–808, 2016. doi: 10.1109/ICRA.2016.7487210.
  52. Learning fine-grained bimanual manipulation with low-cost hardware. In Proceedings of Robotics: Science and Systems (RSS), 2023.
  53. Viola: Imitation learning for vision-based manipulation with object proposal priors. In Proceedings of The 6th Conference on Robot Learning (CoRL), volume 205, pages 1199–1210. PMLR, 2023.
Citations (86)

Summary

  • The paper proposes **UMI (Universal Manipulation Interface)**, combining a portable gripper and wide-angle camera/fisheye lens SLAM system/html> for robust, transferable visuomotor robotic control. It achieves 87.5% success in dynamic tasks and 90% cross-robot success.

Problem formulation and contribution

The paper addresses a persistent bottleneck in robot learning from demonstration: teleoperation provides robot-compatible actions but is expensive, embodiment-specific, and difficult to deploy outside laboratory settings, whereas passive human videos are scalable but lack explicit robot actions and exhibit substantial observation and embodiment gaps. “Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots” proposes UMI, a hardware–software framework intended to occupy the intermediate regime: demonstrations are collected by humans in arbitrary environments using portable hand-held grippers, while the resulting visuomotor policies can be deployed on multiple robot platforms without collecting demonstrations on those robots (2402.10329).

The central claim is that transferability depends less on a single learning algorithm than on jointly designing the demonstration interface, state representation, action parameterization, temporal synchronization, and policy architecture. UMI therefore combines a low-cost hand-held gripper, a wrist-mounted wide-angle camera, mirror-based depth cues, visual–inertial SLAM, continuous gripper control, latency compensation, relative end-effector trajectories, and Diffusion Policy (2402.10329). The resulting system is evaluated on single-arm, dynamic, bimanual, deformable-object, and long-horizon tasks, as well as on out-of-distribution environments and objects.

Demonstration interface

The physical interface is a trigger-activated, 3D-printed parallel-jaw gripper weighing approximately 780 g. Its bill of materials is reported as $73, excluding a GoPro camera and accessories costing approximately $298. The gripper uses soft TPU fingers and continuous width tracking, rather than a binary open–close command. The same camera–gripper geometry is reproduced on the deployment robot, reducing the observation discrepancy between demonstrations and execution and avoiding explicit camera-to-robot-world calibration.

The interface is deliberately sensor-minimal: the GoPro supplies RGB video and inertial measurements, while fiducial markers on the gripper enable estimation of finger width. This design is important for portability, but it shifts the burden of state estimation onto visual–inertial tracking. A 155-degree fisheye lens supplies broad scene coverage from the wrist-mounted viewpoint. The authors use the raw fisheye image rather than rectifying it into a pinhole projection. Their argument is that rectification severely expands peripheral regions while reducing the effective resolution of the central task-relevant area.

Figure 1

Figure 1: UMI uses a portable hand-held gripper with a wrist-mounted GoPro, fisheye observation, and physical side mirrors.

The fisheye camera addresses a specific failure mode of wrist-mounted sensing: the manipulated object can dominate the image while relevant context, support surfaces, targets, and obstacles remain outside the field of view. The paper’s ablation supports this design choice. On the cup-arrangement task, rectifying and cropping the image to a 69-degree field of view reduces success from 20/20 to 11/20, or 55%. The degradation persists even when the object remains visible, with the authors reporting jittery and unnecessarily multimodal behavior. This result suggests that the useful contribution of the fisheye lens is not merely object visibility but preservation of contextual cues needed for action disambiguation.

Figure 2

Figure 2: Rectification of the 155-degree fisheye image distorts peripheral context and compresses central task-relevant information.

The interface also places two mirrors in the peripheral camera view. These mirrors generate virtual viewpoints with different optical centers, providing implicit stereo cues without additional cameras or depth sensors. Because reflected objects have reversed orientation, UMI digitally reflects the mirror crops and swaps the left and right views before policy training. The distinction is empirically consequential: directly supplying uncorrected mirror images achieves 17/20 success, compared with 18/20 without mirrors, whereas digitally corrected mirror views achieve 20/20. Thus, the mirrors are not automatically beneficial; their utility depends on representation-level normalization.

Figure 3

Figure 3: Side mirrors provide virtual stereo views, which are digitally reflected to restore consistent object orientation.

Visual–inertial action recovery

UMI recovers six-degree-of-freedom gripper motion using a modified monocular–inertial SLAM pipeline based on ORB-SLAM3 (2402.10329). The GoPro IMU supplies metric scale and maintains short-term tracking through intervals of motion blur or insufficient visual texture. This capability is particularly relevant to dynamic manipulation, for which monocular structure-from-motion can suffer from scale ambiguity and tracking loss.

The system uses a map-then-localize procedure. A scene map is built before demonstrations, and each subsequent video is relocalized to that map. The authors modify the SLAM system so that relocalization serves as initialization while normal SLAM optimization continues, allowing the map to adapt to scene changes. Fiducial markers can additionally be used during mapping to resolve ambiguities caused by distant features, repeated patterns, or outdoor scenes; the markers need not appear during demonstrations.

A motion-capture benchmark containing seven single-gripper and seven bimanual tasks yields a mean absolute trajectory error of 6.1 mm in position and 3.5 degrees in rotation. Relative pose error between two grippers is 10.1 mm and 0.8 degrees, respectively. These measurements establish that the recovered trajectories are sufficiently precise for the demonstrated tasks, but they do not establish uniformly accurate tracking across textureless or highly dynamic environments. The paper explicitly concedes that the method inherits the texture requirements of visual SLAM.

Figure 4

Figure 4: Visual–inertial SLAM provides metric gripper trajectories and shared-map relative pose for bimanual demonstrations.

Continuous gripper control is another important interface decision. The authors argue that release timing in dynamic tossing depends on object width and therefore cannot be represented reliably by a binary action. Soft fingers also provide passive compliance during contact-rich behaviors, including faucet manipulation and cloth pickup. This mechanical compliance is not equivalent to force sensing, however; the system infers grasp-related behavior through finger width and finger deformation rather than directly measuring interaction forces.

Policy interface and temporal alignment

The policy receives synchronized sequences of RGB images, relative end-effector poses, and gripper widths, and predicts sequences of relative end-effector poses and gripper widths. Diffusion Policy is used throughout the experiments to represent multimodal action distributions, including the clockwise and counter-clockwise solutions to cup reorientation (2402.10329).

The action representation is a sequence of SE(3)SE(3) transforms, each defined relative to the same current end-effector pose. This differs from delta actions, which compose successive local increments and accumulate tracking error, and from absolute actions, which require a globally calibrated coordinate frame. Relative trajectories also make the policy less sensitive to camera displacement and robot-base placement.

Figure 5

Figure 5: Relative trajectories reference all predicted poses to the current end-effector pose, avoiding delta-action error accumulation and absolute-frame calibration.

The experimental comparison strongly favors the proposed representation. On cup arrangement, relative trajectories achieve 20/20 success. Delta actions achieve 16/20, or 80%, while absolute actions achieve only 5/20, or 25%. The absolute-action result is particularly revealing: although absolute actions are theoretically expressive, the required calibration between SLAM coordinates and the robot base introduces a severe systematic bias. The implication is that coordinate-frame reliability can dominate the nominal advantages of an action representation.

UMI also represents proprioceptive history as a relative trajectory. With a short observation horizon, this supplies velocity-like information while remaining invariant to the robot base and scene-level coordinate choice. In bimanual settings, relative inter-gripper pose is explicitly provided. Removing this signal reduces cloth-folding success from 14/20, or 70%, to 6/20, or 30%. The observed failure is asynchronous grasping of the sweater hem, indicating that visual overlap alone is insufficient for precise two-arm coordination in this setup.

Temporal alignment is treated as part of the policy interface rather than as an implementation detail. During data collection, camera, IMU, gripper, and pose measurements are synchronized within the recording pipeline. During deployment, camera, proprioception, inference, arm execution, and gripper execution introduce heterogeneous delays. UMI measures these delays independently, interpolates lower-latency streams to the camera timestamp, discards actions that are already outdated, and transmits future action commands early enough to compensate for execution latency.

Figure 6

Figure 6: UMI synchronizes heterogeneous observation streams and advances action commands to compensate for execution delay.

The dynamic tossing experiment isolates the effect of this procedure. With latency matching, the policy successfully tosses 105 of 120 objects, or 87.5%. Disabling latency matching reduces performance to 69/120, or 57.5%, a 30 percentage-point drop. The degradation is attributed to jittery arm motion, mismatched gripper and arm execution, and incorrect release timing. This result supports the paper’s stronger systems-level claim: for high-speed manipulation, temporal calibration can be as consequential as model architecture or data volume.

Capability experiments

UMI is evaluated on four tasks with randomized initial states and 20 evaluation episodes for most narrow-domain experiments. The tasks span different sources of difficulty rather than merely different object categories.

Task Demonstrations Main capability Success
Cup arrangement 305 Prehensile, non-prehensile, multimodal actions 20/20
Dynamic tossing 280 Rapid motion and release timing 105/120 objects
Bimanual cloth folding 250 Deformable-object and two-arm coordination 14/20
Dish washing 258 Long horizon, articulated and deformable objects 14/20

The cup task requires placing an espresso cup upright on a saucer with its handle within ±15\pm 15 degrees of the desired orientation. It combines grasping, pushing, relative-depth estimation, and multimodal reorientation. The policy achieves 100% success on the primary robot and 90% on a Franka FR2, with both FR2 failures attributed to joint-limit violations. This cross-robot result is evidence for hardware-agnostic deployment, although the failures also demonstrate that the policy is not intrinsically embodiment-aware. UMI instead relies on kinematic filtering and suitable robot placement.

Figure 7

Figure 7: Narrow-domain evaluations show high success on cup arrangement and substantial sensitivity to ablated sensing, action, and proprioceptive interfaces.

Dynamic tossing tests whether human rapid motions can be transferred into robot trajectories when target bins lie outside the robot’s kinematic reach. The 87.5% object-level success rate demonstrates that UMI can represent velocity-dependent behavior and synchronize release with the arm trajectory. The no-latency baseline shows that this capability is conditional on accurate timing compensation.

Cloth folding requires coordinated sleeve folding, hem lifting, rotation, and final folding of a deformable sweater. A centralized policy controlling both arms achieves 70% success. The inter-gripper ablation falls to 30%, establishing that relative two-arm proprioception is not a minor auxiliary feature in this task. The result also clarifies the scope of visual imitation: the policy does not simply learn independent arm behaviors; it uses explicit geometric coupling to coordinate simultaneous contact events.

Dish washing is a seven-step sequential task involving a faucet, plate, sponge, water, ketchup, and recovery behavior when additional sauce is introduced. A CLIP-pretrained ViT-B/16 vision encoder is fine-tuned with Diffusion Policy, producing 14/20 successful episodes, or 70%. By contrast, a ResNet-34 trained from scratch achieves 0/10. The paper therefore makes a strong claim about representation capacity and initialization: for visually complex, long-horizon manipulation with limited task-specific data, pretrained ViT features are necessary in this experiment. The comparison is not a controlled test of architecture alone because it also contrasts pretraining regimes and training configurations, so the result should not be interpreted as establishing universal superiority of ViT over ResNet.

Figure 8

Figure 8: UMI transfers dynamic tossing, bimanual cloth folding, and long-horizon dish washing in addition to cup arrangement.

In-the-wild generalization and collection efficiency

The most extensive generalization experiment uses 1,400 cup-arrangement demonstrations collected by three demonstrators in 30 locations within 12 person-hours. The data include homes, offices, restaurants, and outdoor environments, and 15 cup styles during training. Evaluation takes place in an outdoor cafe and on a black water fountain with a continuously flowing water film, using both training and unseen cup designs.

The policy obtains 28/40 successes on training cups, 15/20 on held-out cups, and 43/60 overall, or 71.7%. On the unseen cups specifically, it achieves 75%, slightly higher than its 70% performance on training cups. The paper describes this as zero-shot generalization because no demonstrations are collected in the evaluation environments or with the held-out cup styles. A model trained only on narrow-domain laboratory data, despite using the same pretrained ViT backbone, achieves 0% in these environments and does not move toward the cup.

Figure 9

Figure 9: Diverse in-the-wild demonstrations support transfer to novel environments and objects, whereas narrow-domain data fail in the same evaluation settings.

The comparison supports the conclusion that pretrained visual representations alone do not substitute for environmental and object diversity. Nevertheless, the evaluation is limited in scale: two unseen environments, two held-out cup styles, and 60 total trials do not characterize generalization over a broad task distribution. Moreover, the evaluation states are manually aligned and success is judged by an operator, which makes the metric operationally meaningful but introduces subjectivity.

UMI also improves data collection throughput relative to space-mouse teleoperation. For cup arrangement, UMI operates at 48% of bare-hand speed and is reported to be more than three times faster than teleoperation. For tossing, it operates at 64% of bare-hand speed, while the space-mouse interface produces no successful demonstration within 15 minutes. The comparison captures reset time, object randomization, and robot faults, making it more representative of end-to-end collection cost than a pure trajectory-recording rate. However, UMI remains slower than direct human demonstration because of its mass, bulk, and lower degrees of freedom.

Figure 10

Figure 10: UMI improves practical demonstration throughput over space-mouse teleoperation while remaining slower than bare-hand demonstration.

Limitations and open questions

The framework depends on kinematic filtering because the deployment robot is not necessarily known during data collection. Filtering removes demonstrations that are infeasible for a target embodiment, but it does not solve the more general problem of adapting valid human trajectories to different kinematic structures, joint limits, or dynamic capabilities. The 90% FR2 result, with failures caused by joint-limit violations, directly illustrates this limitation.

The visual–inertial SLAM pipeline requires sufficient environmental texture and can be vulnerable to feature scarcity, repeated patterns, or severe motion blur. Marker-enhanced mapping improves initialization but does not eliminate the underlying dependence on visual features. The paper leaves open whether a hybrid sensing configuration could preserve portability while supporting texture-deficient spaces.

The mechanical interface is also less dexterous than the human hand. Continuous parallel-jaw control and compliant fingers enable a broader behavior set than binary grippers, but the demonstrated tasks do not establish transfer of highly dexterous in-hand manipulation, force-sensitive assembly, or multi-contact hand behaviors. Finally, the principal success metrics are manually judged, the number of evaluation episodes is modest, and the operator can terminate episodes for safety, faults, or timeout. These conditions are appropriate for physical robot experiments but complicate exact comparison and reproducibility.

Conclusion

UMI presents a coherent systems approach to robot teaching without requiring robots during demonstration collection. Its contribution is the integration of portable hardware, fisheye and mirror-based observation, visual–inertial SLAM, continuous gripper control, relative trajectory representations, explicit inter-gripper geometry, latency matching, and diffusion-based multimodal policy learning. The empirical results show 87.5% object-level success for dynamic tossing, 70% success for bimanual cloth folding and dish washing, 90% cross-robot success on cup arrangement, and 71.7% success in out-of-distribution environments. The ablations indicate that these results depend materially on temporal synchronization, relative action representations, corrected mirror observations, broad visual context, and bimanual proprioception. The framework’s unresolved issues are embodiment feasibility, texture-dependent tracking, limited dexterity, and evaluation subjectivity, but the paper establishes a technically viable route for collecting transferable manipulation data outside conventional robot laboratories.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 2 likes about this paper.