Zhe Wang
AI Research Scientist @ Meta
About
I am a Senior Machine Learning Engineer at Adobe Firefly, where I focus on building text-to-video foundation models and their downstream applications. My work lies at the intersection of generative AI and visual understanding, aiming to develop intelligent systems that can comprehend and generate dynamic visual content from natural language. I bring over 10 years of experience in computer vision, deep learning, and multi-modal learning, with a PhD in Computer Vision. My expertise spans generative modeling, self-supervised learning, and human-centric scene understanding. At Adobe, I’m developing scalable foundation models that power real-world creative tools and enable AI-driven content generation. My research has been published in top-tier conferences and journals such as CVPR, ICCV, and IEEE TPAMI, and has received over 8,000 citations. I also curated one of the largest 3D human pose estimation datasets with scene-level ground truth. Prior to Adobe, I worked on large-scale perception systems and sensor fusion technologies. I’ve also contributed to applied machine learning projects at AWS AI and Adobe Research, including fine-grained video understanding and mobile-level human pose estimation. I’m passionate about building intelligent visual systems that enhance human creativity and expression through generative AI. Google Scholar: https://scholar.google.com/citations?user=FL-t3JEAAAAJ&hl=en
United States
San Francisco
Computer Software
Large Language Models (LLM), PEFT, RLHF, Pal, LORA, Machine Learning, Computer Vision, Python, Image Processing, Computer Science, Algorithms, C++, Matlab, Java, JavaScript, Photoshop, C, OpenGL, Android, HTML 5
Experience

Senior Applied Research Scientist
San Francisco Bay Area
I mainly work on multi-modality and multi-platform AV perception. It includes offboard perception, and on board perception. For architectures I work on camera major sensor fusion. For data consumption I work on self-supervised learning, semi-supervised learning to leverage data more efficiently and more effectively.

Computer Vision Researcher
1. Combined model-based method and non-parametric approaches to leverage the best of both worlds to get better 3d human pose and shape performance. 2. Analyzed the bias in 5 popular 3d human pose datasets, and utilized the bias and achieve SOTA on the 5 datasets. 3. Demonstrated the effectiveness of integrating scene geometry into three popular 3d human pose estimators on 3 popular and SOTA methods. https://arxiv.org/abs/1905.07718 4. Curated the world's largest 3d human pose estimation dataset (1M images) with scene ground truth (gold-standard maya mesh) using Vicon Blade mocap studio, KinectV1, HD cameras. http://wangzheallen.github.io/GPA 5. Explored the good practice in training visual question answering model: https://github.com/wangzheallen/STL-VQA

Joint Research Assistant
CUHK and SIAT
Shenzhen
Contests: --Ranked 1st in ActivityNet Action Recognition Challenge (CVPR16) --Ranked 3rd in Chalearn cultural event detection track (ICCV15) --Ranked 1st in Chalearn action and gesture recognition track (CVPR15) --Ranked 1st in Chalearn cultural event detection track (CVPR15) --Ranked 5th in THUMOS 2015: Video Classification Task (CVPR15) --Ranked 5th in video terror detection competition Paper: One paper accepted by ECCV 2016 --Temporal Segment Networks: Towards Good Practices for Deep Action Recognition One paper accepted by CVPR 2016 --Real-time Action Recognition with Enhanced Motion Vector CNNs One paper accepted by ICASSP 2016 (as oral) --Codebook Enhancement of VLAD Representation for Visual Recognition One paper accepted by ICCV workshop 2015 --Better Exploiting OS-CNNs for Better Event Recognition in Images Two papers accepted by CVPR workshop 2015 --Exploring Fisher Vector and Deep Networks for Action Spotting --Object-Scene Convolutional Neural Networks for Event Recognition in Images Five Technical Report: --CUHK&SIAT Submission for THUMOS15 Action Recognition Challenge --Towards Good Practices for Very Deep Two-Stream ConvNets --CUHK & ETHZ & SIAT Submission to ActivityNet Challenge 2016 --Weakly Supervised PatchNets: Describing and Aggregating Local Patches for Scene Recognition (TIP 2018) --Transferring Object-Scene Convolutional Neural Networks for Event Recognition in Still Images (IJCV 2018) --Temporal Segment Networks for Action Recognition in Videos (TPAMI 2018) Awards: --“CAS Dean's Outstanding Scholarship”, 2015/2016 --“Research Assistantship from CUHK-MMLAB”, 2015/2016
Zhe Wang's Contact Information
Phone
Find the Right Leads
Find Verified Contact Data
What LeadContact does well
Find verified emails, phone numbers, and decision-makers with 98% accuracy.
Find Leads
Find the right people by company, role, industry, location, and more.
925M+ professional profiles

Find Emails
Access verified email addresses for your target contacts.
657M+ emails

Find Phone Numbers
Get cross-validated phone data from multiple top sources.
239M+ phone numbers

More Accurate. Lower Cost.
Find contact data in 1 tool with 98% accuracy
LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.
Great conversations start with the right contact.
It’s time to find yours.







