Yu Rong
Research Engineer @ Meta
About
I am currently a Research Engineer at Meta Reality Labs Research, working on Multimodal AI Agents involvingmultimodal LLMs, audio/TTS models, video generation models, and human avatar creation. My researchinterests include Multimodal LLMs, Generative Models, Video Generation, and 3D Computer Vision.
United States
Kirkland
Computer Software
PyTorch, Artificial Intelligence (AI), Computer Vision, Research, Machine Learning, English, Mandarin, Python (Programming Language), C++, JAVA, Git, Python, Algorithms, Deep Learning, Computer Graphics
Experience

Research Engineer
Redmond, WA
Working on multimodal AI agents, embodied AI, video generation, human avatar creation, and 3D computer vision. - Building toward real-time multimodal embodied AI agents that understand text, speech, facial expression, and visual signals, and generate speech plus photorealistic digital human behavior. - Developed components across the multimodal GenAI stack, including unified RVQ tokenizers for human motion and semantic behavior modeling, LLM-based speech / TTS model extensions with human-latent modalities, audio-controlled avatar animation models with flow matching, and WAN-based video generation models controllable from audio and text. - Work involving data preprocessing and curation, model training, evaluation pipeline development, and training / inference infrastructure for multimodal avatar agent research. - Built human-tracking foundation models for large-scale image and video understanding, using ViT backbones pretrained on billion-scale image data and post-trained with synthetic data. Added prediction heads for human motion and dense landmark estimation, and owned model training and pipeline curation for million-scale video processing. - Developed high-fidelity offline face-tracking algorithms and production-scale processing pipelines for indoor multi-view and in-the-wild datasets, including algorithm design, implementation, quantitative analysis, and large-scale GPU processing. - Built a live AR avatar-driving system that takes desk-mounted camera inputs and predicts headposes and facial-expression latents in real time. Work involving raw data collection, training data generation, deep tracking model training, and live deployment.
Research Intern
Menlo Park, California, United States
We use SMPL-X to represent whole body 3D motion including face, hands and body. Separate modules is adopt to predict independent body part first. Part motion predictions are then composed by the integration module to obtain unified whole body motion results. We design three different integration modules which trade off between accuracy and speed. Our system, both fast mode and slow-yet-precise mode, largely surpass the previous methods for in-the-wild scenarios. Demo, code and models are available at the project page.
Yu Rong's Contact Information
Phone
Find the Right Leads
Find Verified Contact Data
What LeadContact does well
Find verified emails, phone numbers, and decision-makers with 98% accuracy.
Find Leads
Find the right people by company, role, industry, location, and more.
925M+ professional profiles

Find Emails
Access verified email addresses for your target contacts.
657M+ emails

Find Phone Numbers
Get cross-validated phone data from multiple top sources.
239M+ phone numbers

More Accurate. Lower Cost.
Find contact data in 1 tool with 98% accuracy
LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.
Great conversations start with the right contact.
It’s time to find yours.



