Shawn Song
CEO & Co-Founder @ Sonilo
About
Multimodal Recommendation, Multimodal Music Generation, Large Language Model (LLM)
United States
San Francisco Bay Area
Internet
machine learning, image processing, speech processing, C++, Keras, Pytorch, Python, Matlab, Java, SQL, Linux, TensorFlow, Natural Language Processing, Natural Language Understanding, Speech Recognition, C, Microsoft Office
Experience

CEO & Co-Founder
San Francisco Bay Area
Sonilo is an AI video-to-music platform that automatically generates perfectly synchronized soundtracks for your videos. By analyzing visuals, pacing, and emotion, Sonilo creates music that fits your content frame-by-frame—no editing or manual timing needed. Built for creators, filmmakers, and creative tools, it delivers fast, high-quality results and integrates seamlessly into modern video workflows.

CTO, Founding Engineer
San Francisco Bay Area
Leading a 50+ person team across research and engineering, focused on advancing state-of-the-art AI music generation. Owned the full-stack technology strategy for high-fidelity and controllable music generation, including model design, data systems, evaluation, and scalable training and inference infrastructure. Scaled Mureka.ai to $20M+ ARR, positioning it among the leading AI music generation platforms globally. Representative Work: Musicot: Analyzable Chain-of-Musical-Thought Prompting for High-Fidelity Music Generation

Head of Gen AI
San Francisco Bay Area
Leading a 50 members, including music generation (Mureka.ai), world model, multimodal understanding across algorithms, data, evaluation, AI infrastructure, and engineering. Focused on building the world’s best AI-powered song generator, driving innovation from cutting-edge research to large-scale production. Representative Work: Matrix-game 2.0: An open-source real-time and streaming interactive world model (https://arxiv.org/abs/2508.13009) Skywork UniPic 3.0: Unified Multi-Image Composition via Sequence Modeling (https://arxiv.org/abs/2601.15664) Skywork-R1V4: Toward Agentic Multimodal Intelligence through Interleaved Thinking with Images and DeepResearch (https://arxiv.org/pdf/2512.02395)

Tech Lead, Staff Research Engineer
Tech lead and staff research engineer at ByteDance Seed Audio/Music; Leading the TikTok Multimodal Team to conduct cutting-edge tech in large-scale multimodal learning, Large Language Model (LLM) and end-to-end music generation.

Researcher
Sunnyvale, CA
• Designed and enhanced multi-engine Speech Recognition (ASR), Natural Language Processing (NLP), and Natural Language Understanding (NLU) modules for Bosch's in-car voice assistant. • Developed intent detection and slot filling systems, improving user interaction for millions of users. • Collaborated with cross-functional teams to ensure seamless integration of voice technology into automotive applications.
Shawn Song's Contact Information
Phone
Find the Right Leads
Find Verified Contact Data
What LeadContact does well
Find verified emails, phone numbers, and decision-makers with 98% accuracy.
Find Leads
Find the right people by company, role, industry, location, and more.
925M+ professional profiles

Find Emails
Access verified email addresses for your target contacts.
657M+ emails

Find Phone Numbers
Get cross-validated phone data from multiple top sources.
239M+ phone numbers

More Accurate. Lower Cost.
Find contact data in 1 tool with 98% accuracy
LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.
Great conversations start with the right contact.
It’s time to find yours.



