Zhiding Yu

Zhiding Yu

Principal Research Scientist & Research Lead @ NVIDIA

About

I am interested in building general autonomy and intelligence across both virtual and physical domains. My recent focus lies in Vision Transformers, LLMs, multimodal LLMs, and vision-language-action (VLA) models, with applications spanning open-world understanding, reasoning, AV/robot perception-planning, and agentic systems. I have led or contributed to numerous flagship research efforts and products at NVIDIA, including SegFormer (Most Influential NeurIPS Papers), VoxFormer, FB-BEV/FB-OCC (CVPR23 3D Occ Pred Challenge winner), Hydra-MDP (CVPR24 E2E Driving Challenge winner), the Eagle VLM project, Nemotron, Llama-Nemotron-VL, Nemo Retriever, GR00T N1 and GR00T N1.5 (NVIDIA’s foundation models for humanoid robots). I also participated in designing NVIDIA’s next-generation end-to-end autonomous driving system. My works are characterized by state-of-the-art performance, scalable architectures, and data-centric strategies towards real-world generalization.

Country

United States

City

Santa Clara

Industry

Research

Skill

-

Experience

NVIDIA

Principal Research Scientist & Research Lead

NVIDIA

LinkedIn
2024-4 - Present · 2 yrs 5 mos

San Francisco Bay Area

I conduct research in multimodal learning and intelligent data strategies. I lead the Eagle VLM project which develops a family of frontier vision-language models with public training/data recipes and state-of-the-art performance matching or outperforming existing top-tier VLMs. Our work has laid the core VLM foundation and data strategy behind several flagship NVIDIA products/projects, including Llama-Nemotron-VL, Nemo Retriever Multimodal Embedding, GR00T N1, and GR00T N1.5.

NVIDIA

Staff Research Scientist

NVIDIA

LinkedIn
2023-6 - 2024-3 · 10 mos

San Francisco Bay Area

Participated in a multi-org effort to design NVIDIA’s next-generation AV system. Led a team to design and develop a Transformer-based 3D perception system for joint 3D object detection, tracking and online mapping at long distances. Developed Transformer-based neural planner and DriveVLM for E2E driving, with community-recognized works such as BEV-Planner, OmniDrive, and Hydra-MDP.

NVIDIA

Senior Research Scientist

NVIDIA

LinkedIn
2020-6 - 2023-5 · 3 yrs

San Francisco Bay Area

Led the earliest effort to develop Vision Transformers at NVIDIA which partially shaped the landscape of NVIDIA’s internal AI product. Applications of my work include scene understanding, robust general purpose backbone, autonomous driving perception and scalable auto-labeling pipelines. Proposed multiple Transformer-based bird’s-eye view (BEV) perception frameworks with SOTA results in 3D object detection, tracking, and 3D occupancy prediction. Some works from this period with community impact include SegFormer (Most Influential NeurIPS Papers, 3K Stars), VoxFormer (CVPR23 Highlight, 1.1K Stars), FB-BEV/FB-OCC (ICCV23, 735 Stars), and FocalFormer3D (ICCV23, Ranked 1st on nuScenes LiDAR 3D Detection and Tracking Leaderboard (Mar. 2023)). Successful tech transfers to numerous NVIDIA products, including AV and NVIDIA TAO Toolkit.

NVIDIA

Research Scientist

NVIDIA

LinkedIn
2018-1 - 2020-5 · 2 yrs 5 mos

San Francisco Bay Area

Worked extensively on label-efficient learning and transfer learning. Proposed weakly supervised, semi-supervised and self-supervised learning frameworks, with SOTA performance in visual recognition applications. Proposed various unsupervised domain adaptation and synthetic-to-real generalization methods for improved model robustness and generalization “in the wild”.

Mitsubishi Electric Research Laboratories

Research Intern

Mitsubishi Electric Research Laboratories

LinkedIn
2016-7 - 2016-11 · 5 mos

Cambridge, Massachusett

Deep category-aware semantic edge detection Proposed an end-to-end deep semantic edge learning framework with state-of-the-art performance One paper accepted to CVPR 2017

Microsoft

Research Intern

Microsoft

LinkedIn
2015-5 - 2015-8 · 4 mos

Redmond

Deep CNN based static facial expression recognition Partly integrated to the Microsoft Emotion Recognition API under Project Oxford First Runner Up at the EmotiW-SFEW Challenge 2015 Full-length paper accepted at ACM-ICMI 2015 as Oral

Adobe

Research Intern

Adobe

LinkedIn
2013-6 - 2013-8 · 3 mos

San Jose, CA, USA

Intern Project: PixelTone

Education

Carnegie Mellon University

Carnegie Mellon University

LinkedIn

Electrical & Computer Engineering

2012 - 2016 · 4 yrs
The Hong Kong University of Science and Technology

The Hong Kong University of Science and Technology

LinkedIn

Electronic & Computer Engineering

2009 - 2012 · 3 yrs
South China University of Technology

South China University of Technology

LinkedIn

Information Engineering (Talented Student Program)

2005 - 2008 · 3 yrs

Zhiding Yu's Contact Information

Email

******@***.com

Phone

(**) *** ****

Find the Right Leads
Find Verified Contact Data

Try with: Jensen Huang @ nvidia.com Click to autofill
LeadContact awards, five-star ratings, and GDPR compliance badges

What LeadContact does well

Find verified emails, phone numbers, and decision-makers with 98% accuracy.

Find Leads

Find Leads

Find the right people by company, role, industry, location, and more.

925M+ professional profiles

Find Leads
Find Emails

Find Emails

Access verified email addresses for your target contacts.

657M+ emails

Find Emails
Find Phone Numbers

Find Phone Numbers

Get cross-validated phone data from multiple top sources.

239M+ phone numbers

Find Phone Numbers

More Accurate. Lower Cost.

Find contact data in 1 tool with 98% accuracy

LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.

LeadContact Logo
Competitor Tools

All these = $289 per month

Great conversations start with the right contact.

It’s time to find yours.