Masatoshi Tateno

舘野 将寿

Masatoshi Tateno

I am a third-year Ph.D. student at the Graduate School of Information Science and Technology, The University of Tokyo, advised by Prof. Yoichi Sato. I am currently interning at AIRoA.

Before that, I graduated from the same department in 2022 and got Master’s Degree of Information Science and Technology. I also worked as Research Assistant at AIST (2023-2024).

Email: masatate[at]iis.u-tokyo.ac.jp

I am actively looking for research positions starting April 2027, after completing my Ph.D. — feel free to reach out!

News

older news (2025 and earlier)

Research Interest

Research

Project Image

Do Egocentric Video-Language Models Capture Both Hand- and Object-Centric Cues?

Masatoshi Tateno, Alexandros Stergiou, Risa Shinoda, Yoichi Sato, Dima Damen

arXiv 2026

Oral Presentation at MIRU 2026

We propose a learning paradigm combining hand-object masked training with an HOI-dynamics-aware decoder, and introduce the DEHOI testbed that disentangles hand- and object-related observations through inpainting to evaluate cue-specific reasoning. Our approach exploits both cues more effectively than existing models, improving on DEHOI, action recognition, object state recognition, and robot manipulation action recognition.

Project Image

EgoYUBIs: Portable Robot Data Collection with Bimanual Handheld Grippers via Egocentric Perception

Takehiko Ohkawa, Masatoshi Tateno*, Kengo Ikeuchi*, Kei Ota, Yoichi Sato (*equal contribution)

Meta Research Summit for Egocentric Intelligence 2026 - Oral & Poster, Best Poster Award 🏆

EgoYUBIs combines the YUBI gripper with egocentric perception to enable portable, wearable data collection for bimanual dexterous manipulation. Our Aria-based setup integrates marker-cube tracking, wrist-camera SLAM, and interaction perception, making in-the-wild robot data collection possible without any stationary capture rig.

Project Image

YUBI: Yielding Universal Bidigital Interface for Bimanual Dexterous Manipulation at Scale

Takehiko Ohkawa, Jumpei Arima, Yuki Noguchi, Masatoshi Tateno, ..., Kei Ota

Beyond Teleoperation Workshop @ ICRA 2026

We introduce YUBI, a finger-aligned gripper whose finger-driven actuation enables ergonomic and scalable data collection for bimanual dexterous manipulation, and curate a UMI-based dataset of unprecedented scale: 8,434 hours across 1.20M episodes and 119 tasks. A single policy trained on this data transfers across multiple bimanual robots simply by mounting the gripper, and we release the hardware, software, and dataset as one integrated open stack.

Project Image

HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction Dynamics

Masatoshi Tateno, Gido Kato, Hirokatsu Kataoka, Yoichi Sato, Takuma Yagi

CVPR 2026 - Highlight ✨

Invited Oral Presentation at MIRU 2026

Oral Presentation at MIRU 2025

We introduce HanDyVQA, a fine-grained video question-answering benchmark covering both the manipulation and effect aspects of hand-object interaction, with 11.1K multiple-choice QA pairs across six question types and 10.3K segmentation masks. Even the best-performing model, Gemini-2.5-Pro, reaches only 73% average accuracy against 97% human performance, revealing remaining challenges in spatial relationship, motion, and part-level geometric understanding.

Project Image

Learning Multiple Object States from Actions via Large Language Models

Masatoshi Tateno, Takuma Yagi, Ryosuke Furuta, Yoichi Sato

WACV 2025

Oral Presentation at MIRU 2024

We formulate object state recognition as a multi-label task handling the multiple states an object holds simultaneously (e.g., an egg can be both raw and whisked), and leverage large language models to derive these states from narrations that rarely mention them, accumulating past context into the pseudo-labels. On our newly collected Multiple Object States Transition (MOST) dataset, our model trained with LLM-generated pseudo-labels significantly outperforms strong vision-language models.

Project Image

Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

Kristen Grauman, Andrew Westbry, Lorenzo Torresani, Kris Kitani, Jitendra Malik, ..., Masatoshi Tateno, ..., Michael Wray

CVPR 2024

We present Ego-Exo4D, a large-scale multimodal dataset of simultaneously-captured egocentric and exocentric video of skilled human activities, with 1,286 hours from 740 participants accompanied by audio, eye gaze, 3D point clouds, camera poses, IMU, and expert commentary from coaches and teachers. We also provide benchmark tasks and annotations for fine-grained activity understanding, proficiency estimation, cross-view translation, and 3D hand/body pose, all open sourced for the community.