Yingdong Hu

Lead of Frontier Technology Exploration · Spirit AI

Yingdong Hu

I lead Frontier Technology Exploration at Spirit AI. I received my Ph.D. from the Institute for Interdisciplinary Information Sciences (IIIS), Tsinghua University, advised by Prof. Yang Gao. Previously, I obtained my bachelor's degree from Beijing University of Posts and Telecommunications (BUPT).

My research focuses on Embodied AI, a frontier domain at the intersection of machine learning, robotics and computer vision. My goal is to develop general-purpose robots endowed with human-like agility and dexterity, and capable of generalizing learned behaviors to novel tasks and unseen real-world environments.

Yingdong Hu
01

Research

PDFFull research statementOct 2025

My work builds toward what I call the Keystone Arch Architecture: two pillars, locked together by a keystone.

Keystone

Generalist Robot Foundation Model

A large-scale robot policy that unifies the modules of the stack in a single architecture and feeds on the data engine — generalizing broadly across tasks and embodiments.

Left pillar

Learning-First Robotics Stack

Renovating the classical perception → planning → control pipeline by embedding scalable learning-based methods at every layer.

Right pillar

Robot Data Engine

There is no “Internet for Robots.” I study which data actually matters for physical intelligence, and how to scale it.

02

News

2026.06

We are organizing the Data-Centric Robotics Workshop at RSS 2026.

2026.06

We release OpenHLM, an empirical recipe for whole-body humanoid loco-manipulation.

2025.05

VLAs need to reason — but they need to know when to reason (or not)! We release OneTwoVLA.

2024.10

We release Data Scaling Laws in Imitation Learning for Robotic Manipulation.

2023.08

Semantic-Geometric Representation (SGR) is accepted at CoRL 2023.

2023.04

Our work on pre-trained vision models in motor control is accepted at ICML 2023.

2022.07

SFC is accepted at ECCV 2022 as an oral presentation!

03

Publications

Representative papers (* denotes equal contribution). For the full list, see my Google Scholar.

Data Scaling Laws in Imitation Learning for Robotic Manipulation

Yingdong Hu*, Fanqi Lin*, Pingyue Sheng, Chuan Wen, Jiacheng You, Yang Gao

International Conference on Learning Representations (ICLR), 2025

Oral Presentation ★ Best Paper · CoRL 2024 Workshop

We demonstrate that a policy's generalization ability to new objects, new environments, or both scales approximately as a power law with the number of training objects, training environments, or training environment–object pairs, respectively.

GR-3 Technical Report

ByteDance Seed

Technical Report, 2025

GR-3 is a large-scale vision-language-action (VLA) model. It showcases exceptional capabilities in generalizing to novel objects, environments, and instructions involving abstract concepts.

PVM

For Pre-Trained Vision Models in Motor Control, Not All Policy Learning Methods are Created Equal

Yingdong Hu, Renhao Wang, Li Erran Li, Yang Gao

International Conference on Machine Learning (ICML), 2023

We conduct the first thorough evaluation of pre-trained vision model performance across different downstream policy learning methods and environments, discovering that the effectiveness of pre-training is highly dependent on the choice of the downstream policy learning algorithm.

SFC

Semantic-Aware Fine-Grained Correspondence

Yingdong Hu, Renhao Wang, Kaifeng Zhang, Yang Gao

European Conference on Computer Vision (ECCV), 2022

Oral Presentation

We show that fine-grained features learned with pixel-level self-supervised learning (SSL) objectives are complementary to semantic features from image-level SSL methods. Fusing these features can significantly improve performance on visual correspondence tasks.