Yan Li
Head of Multimodal AI and AIGC @ JD.COM
About
As an AI researcher and technical leader, I specialize in computer vision and generative AI. My career began with a BSc in Computer Science from Nankai University (2010), followed by a Ph.D. from the Institute of Computing Technology, Chinese Academy of Sciences (2017), where my research focused on computer vision, deep learning, and facial recognition. My doctoral work was recognized with the Director’s Special Award — the institute’s highest honor for outstanding academic contributions.After my Ph.D., I joined Tencent WeChat as a Researcher, leading multimodal content understanding for core features including Moments and Favorites. In 2019, as a founding member of WeChat Channels (Video Accounts), I designed its full-scene content intelligence framework, covering short-video analysis, live-streaming comprehension, and user behavior modeling.In 2021, I moved to Kuaishou to lead content understanding for its global expansion, serving regions such as Southeast Asia, Latin America, and the Middle East. Since 2023, I have been leading the image generation project for Kuaishou’s Kling (可灵) AI, driving end-to-end development of its text-to-image foundation model from research to product integration.In 2025, I joined JD.com as the Head of Multimodal AI and AIGC. I lead initiatives across product, operations, algorithm, and engineering to redefine e-commerce experiences. This includes developing multimodal understanding and controllable generation models, launching the JD AIGC platform for merchants, and advancing virtual try-on with personalized outfit recommendations. We are now pioneering the next paradigm of dynamic, AI-generated product content tailored to individual user preferences.My experience includes building and mentoring technical teams to deliver scalable AI solutions. I am committed to advancing AGI with ethical deployment, and welcome collaborations to shape the future of intelligent systems.
China
Beijing
Internet
Technical Leadership, Generative AI, Computer Vision, Artificial Intelligence
Experience

Head of Multimodal AI and AIGC
北京市, 中国
I currently lead the Multimodal AI and AIGC initiatives at JD.com, overseeing product strategy, operations, design, algorithm development, and engineering R&D. On the algorithm side, my team focuses on advancing e-commerce multimodal understanding models, controllable visual generation models, and unified models tailored for online retail scenarios. Through our product efforts, we launched JD Oxygen Vision, an AIGC-powered platform designed to help merchants create, test, and optimize product visuals and advertising materials. The platform now serves a wide range of sellers on JD.com. We’ve also introduced a virtual try-on capability that enables customers to visualize clothing fit and receive personalized outfit recommendations. Looking forward, we’re working toward the next-generation content paradigm in e-commerce: fully personalized, AI-generated visual materials for each user. By delivering unique product visuals tailored to individual preferences, we aim to enhance the shopping experience while driving better conversion and business growth for our merchants.

Senior Staff Algorithm Engineer / Director
Beijing, China
As Tech Lead at Kuaishou’s Kling AI Team, I led the development of data infrastructure for collecting, cleaning, and managing multi-billion-scale high-quality image/video datasets. My team engineered the foundational text-to-image generative model (Kolors), innovating core components including multilingual text encoders, variational autoencoders (VAE), and diffusion model architectures. Leveraging this proprietary technology, we developed ID-preserving controllable image generation and e-commerce AIGC solutions—such as virtual try-on capabilities. To date, our Kolors model has generated over 340 million premium images for diverse applications, serving millions of users while ensuring consistent brand identity in generated assets.

Senior Staff Algorithm Engineer / Director
Beijing, China
At Kwai Global, I led multimodal content understanding across all scenarios — including short videos, live streams, e-commerce products, and commercial ad creatives — for regions spanning Latin America (Portuguese/Spanish), Southeast Asia (Indonesian/Urdu), and the Middle East (Arabic/Turkish). This required developing multilingual models to handle linguistic and cultural nuances in 10+ languages. Additionally, I applied AIGC capabilities to repurpose existing videos through automated editing and localized dubbing, enriching Kwai’s overseas content ecosystem. I also generated commercial ad creatives (e.g., synthetic product demos, regionally adapted visuals) to support user acquisition campaigns. My work established a scalable framework for multimodal content understanding and generation across diverse global markets.

Senior Researcher / Lead of Multimodal Content Understanding
Beijing, China
At WeChat Video Channels, I led the development of core content understanding capabilities, including video classification from broad categories to fine-grained tags, video duplicate detection algorithms, and facial recognition systems. These multimodal models formed the foundational features for the recommendation system, providing critical signals for new content during cold-start phases.

Senior Researcher
Beijing, China
As a core AI researcher at Tencent WeChat, I led the development of early-stage facial recognition algorithms deployed across WeChat’s ecosystem, enhancing security and user authentication for hundreds of millions of users. My work included optimizing algorithms for accuracy in challenging scenarios like cross-age verification, contributing to industry-leading performance on benchmarks. For WeChat Moments, I spearheaded multimodal content understanding systems, designing a comprehensive knowledge framework with 40,000+ semantic tags to analyze image-text relationships and improve content recommendations. Additionally, I drove AI innovation for WeChat 7.0’s flagship feature—AI-powered video soundtracking—developing algorithms that dynamically matched music to video content, which became one of the industry’s most acclaimed video enhancement tools during its release. These projects leveraged computer vision, multimodal learning, and generative AI to elevate user experiences at scale.
Yan Li's Contact Information
Phone
Find the Right Leads
Find Verified Contact Data
What LeadContact does well
Find verified emails, phone numbers, and decision-makers with 98% accuracy.
Find Leads
Find the right people by company, role, industry, location, and more.
925M+ professional profiles

Find Emails
Access verified email addresses for your target contacts.
657M+ emails

Find Phone Numbers
Get cross-validated phone data from multiple top sources.
239M+ phone numbers

More Accurate. Lower Cost.
Find contact data in 1 tool with 98% accuracy
LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.
Great conversations start with the right contact.
It’s time to find yours.



