July 28th report: Today, Beyond, a company specializing in universal embodied basic models, officially launched...The world's first implicit haptic world motion model, Being-H0.8.
It is understood that Being-H0.8This is the first time that tactile modalities have been introduced into large-scale model pre-training.It unifies vision, touch, movement, and future state changes into the same latent space.
This enables robots to further understand "how the world has changed as a result," thereby forming a more complete understanding and predictive ability regarding physical interaction processes.
Wisdom without boundaries isThe team that pioneered the use of large-scale human video data to train a general embodied basic model framework.
It is reported that Zhizai Wujie has established the UniHand 3.0 human interaction database, which covers natural object interaction, long-term tasks, rich hand gestures, and diverse scenarios.The company has amassed over 500,000 hours of first-person video data and has also partnered with dozens of data partners.
Wisdom in the Boundless Realm is said to have been completed.The entire stack of infrastructure construction, from data pipeline, model pre-training, post-training, evaluation to edge deployment.
01.
The model deploys a dual-robotic arm platform, capable of squeezing toothpaste and picking grapes.
Being-H0.8 was deployed on two dual-arm robotic platforms.
The first platform uses two ROKAE AR5 robotic arms and supports the replacement of different dexterous hands for cross-embodiment evaluation.
This platform includes the LinkerHand L25, which has 16 active degrees of freedom and 21 total degrees of freedom, and the DexHand 021, which is tendon-driven and has 12 active degrees of freedom and 22 total degrees of freedom.
Each fingertip of the L25 integrates a 12×6 piezoresistive pressure array, while the DexHand 021 integrates visual tactile sensors.
The second platform consists of two RealMan RM65 robotic arms equipped with Inspire grippers and Daimon DM-TAC W2M tactile sensors.
In the video provided by Zhizai Wujie, the robotic arm can complete tasks through tactile perception.Squeezing toothpaste, finding hidden items in a bag, writing calligraphyOperations such as...
▲A robotic arm squeezing toothpaste (Source: Wisdom Without Borders)
▲A robotic arm searches for a water cup (Source: Wisdom Without Boundaries)
▲A robotic arm writing calligraphy (Source: Wisdom Without Boundaries)
▲A robotic arm picking grapes (Source: Wisdom Without Borders)
02.
Three major breakthroughs, introducing the slow and fast motion expert
It is understood that the core breakthrough of Being-H0.8 compared to its predecessor, Being-H0.7, is...For the first time, tactile modalities are systematically introduced into the implicit world model architecture.It is also the first embodied foundation model to be pre-trained by combining large-scale human video and tactile information.
The addition of tactile feedback enables the model to perceive contact states, force changes, and object constraints that are difficult to observe directly with vision, thereby enabling it to complete contact-intensive tasks that are difficult to solve stably by relying solely on vision.
To make full use of high-frequency tactile feedback, Being-H0.8 further proposes the Slow-Fast Action Expert.
Building upon the standard full action block generation mechanism, this model introduces a lightweight, fast update flow: the system first calculates and caches the "world-action context" at the start of each action horizon.
Subsequently, at a few anchor points in the time domain, the latest observed ontological state and haptic feedback are injected to dynamically generate or correct the currently executing action segment.
To address the bottleneck of limited tactile data volume, Zhizai Wujie has independently developed...TactoHand, a dense tactile pseudo-labeling system for large-scale unlabeled human videos.
This system can infer the contact probability and continuous proximity of dense spatial points during hand-object interaction from ordinary human videos without requiring data collectors to wear tactile gloves or additional sensors, thus supplementing video with tactile supervision.
In Being-H0.8, Zhi proposed the second generation of [the technology/intelligence].TopoHand (Unified Action Space)It provides a unified kinematic interface for human hands, dexterous hands, and parallel grippers.
TopoHand employs a fixed-topology "screw-hand" representation, consisting of 20 semantic keypoints and 20 canonical joint variables, and establishes a unified coordinate system centered on the wrist:
The x-axis points in the direction of the projection of the middle fingertip onto the palm plane, the z-axis is in the same direction as the palm normal, and the y-axis is used to construct the right-hand coordinate system.
Under this representation, the joint naming, mechanical structure, and mesh topology specific to different ontologies are all isolated from the policy interface.
Therefore, this model no longer directly relies on the original joint definitions of a specific robot or human hand model, but...Learning operational rules within a unified semantic topology and kinematic space.
Compared to the first-generation UAS, TopoHand achieves compatibility at the control variable level, enabling operational priors in large-scale human videos to be transferred more efficiently to various robot bodies, and significantly improving the efficiency and scalability of cross-body pre-training.
03.
Establish a database of human interactions, totaling 500,000 hours.
To train Being-H0.8, Zhi collected over [number missing] data within the boundless [data missing].500,000 hoursThe original first-person human video footage.
According to Zhizai Wujie, this is the largest human video dataset to date, and every sample in the dataset can be traced back to its source.
The database transforms raw first-person video data into structured embodied training data, which includes interactive video clips, temporally coherent hand motion trajectories, and 3D spatial information.
According to Zhizai Wujie, its data curation pipeline removes redundant episodes and segments, while retaining the "manipulation-centric" structure required for embodied learning, making the UniHand 3.0 dataset more closely resemble general multimodal videos in terms of the distribution of diverse features.
This pipeline is specifically designed forLarge-scale in-the-wild dataThe design significantly reduces serious tracking errors, hand shape distortion, and unusable tracks.Automated quality control stageIt can detect more than [number] cases before pre-training.92%The erroneous task trajectory fragment.
Wisdom in the Boundless Realm also achievedStandardization of heterogeneous robot dataTo manage this heterogeneity, the company breaks down each data source into reusable platform-level embodiment specifications and dataset-level metadata.
In this way, the standardization process is simplified into two controllable steps: first, different robot platforms are mapped to a unified shared embodiment interface, and then calibration, synchronization, and validation are carried out for specific datasets.
For each supported platform,Standard URDF file (Canonical URDF)The kinematics, bimanual geometry, joint conventions, and wrist or end-effector frames of the robotic arm are defined.
By reconstructing calibration parameters from the robot's geometry, recorded metadata, and visual observation results, Zhizai Wujie is able to map the original state and action data to the unified URDF specification.
Furthermore, each aligned dataset undergoes checks for data integrity, cross-stream synchronization, kinematic consistency, and language-video agreement.
Samples that fail the validation are either filtered out directly or sent to a dedicated reprocessing pipeline; while high-quality task trajectory segments are enhanced using geometry-consistent transformations to maintain consistency between images, states, actions, and language.
04.
Conclusion: Moving from "being able to see" to "being able to accurately grasp"
The inflection point for industrialization may be approaching.
The release of Being-H0.8 shows that"Haptic perception" and "cross-embodied generalization" are becoming key breakthroughs for the generalization of embodied intelligence basic models..
Zhizai Wujie solves the challenge of scaling tactile data through TopoHand, breaking down the barriers to unified motion representation across different robot hardware. Combined with UniHand 3.0, a large-scale database, it provides the industry with a complete paradigm from underlying algorithms and data cleaning to hardware implementation.
As the implicit world model continues to deepen its understanding of the laws of physical interaction, "vision + touch" is further propelling robots out of the laboratory and specific working conditions, demonstrating higher generalization and robustness in real-world scenarios such as home services and complex industrial precision operations.
The industrialization inflection point of embodied intelligence, moving from "visible" to "accurately graspable and precisely controlled," may come faster than expected.
