Vasu Sharma
Head of Artificial Intelligence @ Pocket FM
About
I am presently working as a Senior Staff Research Scientist at Tesla Optimus, working on building Causal, Dynamic, Real time World Models to enable closed loop RL training for humanoid robotics. I am also working on building Multimodal foundation models, efficient video generation models, Robotics Foundation models and high fidelity autolabelling pipelines working at massive scale to enable reliable training for Robotic action models and to develop true generalization capabilities in a diverse set of real world usecases. Previously I worked as an Applied Research Scientist Lead at Facebook AI research, working on building Multimodal foundational generative AI models. I am also interested in the domain of self supervised learning. I have published 100+ papers across top AI conferences like NeurIPS, CVPR, ACL, EMNLP, TMLR, ICLR, NAACL, COLM, EACL, WACV, Interspeech among others garnering over 15k+ citations. I routinely work with multi-billion scale datasets to train these massive multimodal models. In the past, I have also worked as Quantitative Researcher at Citadel where my work involved leveraging the power of Machine Learning and Statistical methods in an attempt to fathom the enigmatic world that is the financial markets. I have also worked at Amazon Alexa AI on large scale multimodal models and Embodied AI applications to bring smart robot intelligence to Alexa devices. I actively advise several early stage startups and often guest lecture at Stanford, CMU, MIT, Oxford, Oreilly among others. I graduated from Indian Institute of Technology, Kanpur completing my Bachelors in Computer Science and Engineering and then completed my graduate school in Machine Learning and Artificial Intelligence at the Language Technologies institute at Carnegie Mellon University.I am deeply passionate about research in my field. My research interests include: Multimodal foundation models and it's uses in the field of Computer Vision, Speech and Music Processing and Natural Language Processing. I have had the good fortune of working with some amazing people at some fantastic places and have learnt a lot from them. I hope to continue learning, travelling to new places and meeting new people. My mantra in life is- "Live life with passion - Love what you do, do what you love". Besides being a Technology enthusiast, I am also very passionate about sports. I was a part of the IIT Kanpur Aquatics team and love to swim, play Water Polo, Soccer and Cricket.
United States
San Francisco Bay Area
Information Technology & Services
Generative AI, Large Language Models (LLM), Natural Language Processing (NLP), Multimodal machine learning, Deep Learning, Computer Vision, Speech Processing, Machine Learning, Data Mining, Software Development, Programming, Swimming, Artificial Intelligence, Algorithms, Computer Science, Data Structures, Image Processing, C++, C, Python
Experience

Head of Artificial Intelligence
San Francisco Bay Area
Building the worlds largest AI native media unicorn. Leading AI at Pocket Entertainment, a profitable, late-stage, Lightspeed-backed unicorn startup with ~500M$ ARR, 300M+ users across 20 countries and 100B+ minutes streamed in 2025. Built and scaled a 60-person AI org across GenAI Research, Applications, and Personalization to power a 75,000-title media catalog spanning audio, comics, and novels, where AI was foundational to content generation and user experience. - Directed the development and deployment of custom LLMs and diffusion models for long-form, multimodal storytelling, optimized for coherence, cultural adaptation, and multilingual output capable of creating hours of high fidelity audio/video content. - Creating the worlds best creative writing model allowing true democratization of long form content creation, enabling writers to generate hundreds of hours of creative narrative content while maintaining novelty, coherence and writing style. - Architected agentic content generation pipelines ingesting billions of tokens monthly to enable scalable, creator-augmented narrative production. - Working on creating the most performant multilingual narrative TTS model which will power our audio content library. - On the personalization side, led modeling efforts for real-time ranking, churn prediction, and LTV optimization, driving measurable uplifts in engagement, retention, and monetization. Our AI efforts were foundational to the companys growth from 200M$ ARR to 450M$ ARR in under a year while becoming profitable!

Startup Advisor
• I advise multiple startup founders and VCs on technical and AI strategies for scaling from ideation to growth stage • Advising companies in the domains of Multimodal Foundation model training/post training, AI evals, World models, RL environments, Data annotations, Synthetic data generation, Video generation models, Robotic foundation models, Audio foundation model training, Industrial Computer vision applications and horizontal RL stack creation. • Designed AI infrastructure and built AI native products to drive customer acquisition and growth. • Facilitated connections between founders and VCs and assisted in hiring top talent to support company expansion.

Senior Staff Machine Learning Scientist
Palo Alto, CA
Optimus Robotics team Working on Elon's vision to bring humanoid robots to life - Building causal, real time and dynamic World models for Closed loop RL training of Robotics Foundation models - Worked on designing the flagship Multimodal Foundational model (TRON) to enable multimodal reasoning capabilities for Optimus and Tesla FSD - Created a high fidelity auto labelling pipeline to enable automated annotations at scale - Improved throughput of DiT based video generation backbones by ~ 100x through use of efficient attention computations with barely any drop in visual fidelity powering both Autopilot and Optimus usecases. - Significantly improved visual quality on Wan3.x video generation models by enhancing the MoE setup by scaffolding high motion expert, specializing in generating visually dynamic sequences leading to higher visual fidelity in these long tail but crucial scenarios.

Applied Research Scientist Lead
Menlo Park, CA
Working with Facebook AI Research (FAIR) as a part of Meta Superintelligence labs (MSL), on research on Large scale multimodal foundational models on trillion scale datasets. Particularly interested in Generative AI research and production use cases and exploring the realm of Self supervised learning. Published list of papers: - DINOv2: Learning Robust Visual Features without Supervision (Published at TMLR ) https://dinov2.metademolab.com/ - Chameleon: Mixed-Modal Early-Fusion Foundation Models (https://about.fb.com/news/2024/06/releasing-new-ai-research-models-to-accelerate-innovation-at-scale/) - Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning (CM3Leon) (https://ai.meta.com/blog/generative-ai-text-images-cm3leon/) - Demystifying CLIP Data (MetaCLIP) (Published at ICLR 2024): (https://github.com/facebookresearch/MetaCLIP) - Mavil: Masked audio-video learners (Published at NeurIPS 2023): (https://ar5iv.labs.arxiv.org/html/2212.08071) - A Picture is Worth More Than 77 Text Tokens: Evaluating CLIP-Style Models on Dense Captions (Published at CVPR 2024): (https://openaccess.thecvf.com/content/CVPR2024/papers/Urbanek_A_Picture_is_Worth_More_Than_77_Text_Tokens_Evaluating_CVPR_2024_paper.pdf) - Seamless Interaction (https://ai.meta.com/research/publications/seamless-interaction-dyadic-audiovisual-motion-modeling-and-large-scale-dataset/) - FLAP: Fast Language-Audio Pre-training (Published at ASRU 2023) (https://arxiv.org/abs/2311.01615) - An Introduction to Vision-Language Modeling (https://arxiv.org/abs/2405.17247) - Text Quality-Based Pruning for Efficient Training of Language Models (https://arxiv.org/pdf/2405.01582) and DMLR 2025 - Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM (Published at COLM 2024) (https://arxiv.org/abs/2403.07816)

AI Research Director
Using my personal time to lead the development of a cutting-edge AI program to empower students with industry-relevant skills, leveraging top AI research experience from leading labs. • Our students got admitted to CMU, Stanford, MIT, UCBerkeley among others and several received job offers from top companies like Anthropic and Tiktok. • Collaborated with teams to create unparalleled opportunities in AI education, nurturing future innovators in the field allowing students as young as high school to publish to ICML, Neurips, EMNLP, EACL among others. Published papers: - FrontierScience Bench: Evaluating AI Research Capabilities in LLMs (ICML: REALM 2025) - Rosetta-PL: Propositional Logic as a Benchmark for Large Language Model Reasoning: https://arxiv.org/pdf/2505.00001 - FaceSafe: An Inpainting Pipeline for Privacy-Compliant Scalable Image Datasets (ICML 2025 : DIG-BUGS) - COREVQA: A Crowd Observation and Visual Entailment Visual Question Answering Benchmark (ICML 2025 : DIG-BUGS) - Pause-Tuning for Long-Context Comprehension: A Lightweight Approach to LLM Attention Recalibration (ICML 2025: LCFM 2025) - NovelHopQA: Diagnosing Multi-Hop Reasoning Failures in Long Narrative Contexts (ICML 2025: LCFM 2025) - Rewrite-to-Rank: Optimizing Ad Visibility via Retrieval-Aware Text Rewriting (ICML 2025: MoFA 2025) - TRUTH DECAY: Quantifying Multi-Turn Sycophancy in Language Models:https://arxiv.org/pdf/2503.11656 - Deconstructing bias: A multifaceted framework for diagnosing cultural and compositional inequities in text-to-image generative models: https://arxiv.org/pdf/2505.01430 - Advancing Uto-Aztecan Language Technologies: A Case Study on the Endangered Comanche Language: https://aclanthology.org/2025.americasnlp-1.4.pdf

Applied Scientist
Sunnyvale, CA
Working with Alexa AI on a variety of problems like: - A dialog enabled visual-language navigation bot leveraging the multimodal data sources to faithfully navigate a virtual environment based on user instruction. Created a new benchmark for visual language navigation as a part of the Alexa Prize Simbot challenge and designed benchmark models for the same - Designing efficient multimodal transformers to speed up their training and deployment by improving the computational complexity of the self attention mechanism - Video processing applications like video action recognition, video question answering, video summarization, moment retrieval etc working directly with compressed video streams - Created a benchmark for cooperative heterogenous multi agent reinforcement learning platform including open sourcing the collected dataset and it's benchmark models - Working on creating a massively multimodal transformer pipeline capable of handling a wide range of input modalities with modality agnostic transformer blocks which work well for a several tasks leveraging a multitude of modalities Published list of papers: - Alexa arena: A user-centric interactive platform for embodied ai (Published at Neurips 2023) (https://www.amazon.science/publications/alexa-arena-a-user-centric-interactive-platform-for-embodied-ai) - Alexa, play with robot: Introducing the first Alexa Prize SimBot Challenge on embodied AI (https://www.amazon.science/alexa-prize/proceedings/alexa-play-with-robot-introducing-the-first-alexa-prize-simbot-challenge-on-embodied-ai) - CHMARL: A Multimodal Benchmark for Cooperative, Heterogeneous Multi-Agent Reinforcement Learning (Published at RSS 2022) (https://www.amazon.science/publications/chmarl-a-multimodal-benchmark-for-cooperative-heterogeneous-multi-agent-reinforcement-learning) - ε-ViLM: Efficient Video-Language Model via Masked Video Modeling with Semantic Vector-Quantized Tokenizer (Published at WACV 2024) (https://arxiv.org/abs/2311.17267)

Graduate Research Assistant
Multicomp Lab, Language Technologies Institute, School of Computer Science
I worked with Prof. Loius Phillipe Morency on a wide range of project related to Multi Modal Machine Learning and on building robust, explainable Deep Learning models - We designed adversarial attack mechanisms on Visual Question Answering models to identify their vulnerabilities and then fix the same by proposing a variety of robust training mechanisms. - We worked on developing a neural network model which uses a Deep Convolutional Neural Network based pipeline alongside a geometrically conditioned point distribution model for Facial Landmark Detection. - We alsondeveloped the first fully ecologically validated models of visual perception. We will combine intracranial EEG (iEEG)recordings captured during long stretches of natural visual behavior with cutting-edge computer vision, machine learning, and statistical analyses to understand the neural basis of natural, real-world visual perception. - We also explored facial expression recognition in extreme face scenarios like profile face views, occluded faces, non centric and rotated faces alongside recognition for gender, age and racially diverse faces.

Graduate Research Assistant
Articulab, Language Technologies Institute, School of Computer Science
I worked on the SARA and the Yahoo! InMind projects at the ArticuLab which focus on building a socially aware robotic assistant. My primary focus was on trying to combine the user’s visual, vocal and verbal cues to better gauge the ‘rapport’ between the user and the conversational agent and using it to enable the agent to become socially more aware to the user’s emotional needs.
Vasu Sharma's Contact Information
Phone
Find the Right Leads
Find Verified Contact Data
What LeadContact does well
Find verified emails, phone numbers, and decision-makers with 98% accuracy.
Find Leads
Find the right people by company, role, industry, location, and more.
925M+ professional profiles

Find Emails
Access verified email addresses for your target contacts.
657M+ emails

Find Phone Numbers
Get cross-validated phone data from multiple top sources.
239M+ phone numbers

More Accurate. Lower Cost.
Find contact data in 1 tool with 98% accuracy
LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.
Great conversations start with the right contact.
It’s time to find yours.




