智在无界发布具身世界动作模型Being-H0.8,基于50万小时第一人称人类视频数据实现跨本体机器人控制
英文摘要
Chinese embodied AI startup BeingBeyond, founded by PKU professor Lu Zongqing, released Being-H0.8, a world action model trained on 500,000 hours of first-person human video—currently the largest such dataset in the world. The model integrates latent touch information during pre-training, supports control of over 30 robot embodiments, and achieves 30 Hz inference on edge devices. Unlike many competitors, the company focuses purely on a general-purpose embodied foundation model without building its own robots, aiming to license the model to robot makers and system integrators. The release marks a step toward a model that combines visual understanding, tactile interaction, action generation, and cross-body generalization.
中文摘要
由北大教授卢宗青创立的具身智能公司智在无界发布了Being-H0.8世界动作模型,使用50万小时第一人称人类视频进行预训练,是目前全球最大的该类数据集。该模型在预训练阶段加入了隐式触觉信息,可跨30多种机器人本体进行控制,并在端侧实现30赫兹的实时推理。与众多软硬一体公司不同,智在无界坚持只做通用具身基础模型,通过售卖模型授权和后训练服务商业化。此次发布标志着模型能力从视觉理解向触觉交互和跨本体泛化演进。
关键要点
Being-H0.8 is trained on 500,000 hours of egocentric human video, the largest known dataset for embodied AI pretraining, enabling broad scene coverage.
Being-H0.8基于50万小时第一人称人类视频预训练,是目前全球已知最大的具身智能预训练数据集,覆盖广泛场景。
The model incorporates latent tactile information by annotating contact points in video, moving beyond purely visual understanding.
模型通过在视频中标注接触点引入隐式触觉信息,使预训练不仅包含视觉理解,还融合了初步的触觉交互。
It supports cross-embodiment control—deployable on over 30 different robot platforms—and runs inference at 30 Hz on edge GPUs.
支持跨本体部署,可控制超过30种机器人,端侧推理速度达到30赫兹,满足动态任务需求。
BeingBeyond adopts a latent world action model paradigm, requiring only ~1% of the compute of video-generation-based world models, enabling cost-efficient scaling.
采用隐式世界动作模型范式,算力消耗仅为基于视频生成的世界模型的大约1%,降低了规模化的训练成本。
The company remains a pure model provider, rejecting hardware development to focus on a general-purpose embodied foundation model, with commercialization via model licensing and post-training services.
公司坚持纯模型路线,不做硬件,专注于通用具身基础模型,商业化方式包括模型授权和面向高价值场景的后训练服务。