Vasu Sharma

Vasu Sharma

Head of Artificial Intelligence @ Pocket FM

About

I am presently working as a Senior Staff Research Scientist at Tesla Optimus, working on building Causal, Dynamic, Real time World Models to enable closed loop RL training for humanoid robotics. I am also working on building Multimodal foundation models, efficient video generation models, Robotics Foundation models and high fidelity autolabelling pipelines working at massive scale to enable reliable training for Robotic action models and to develop true generalization capabilities in a diverse set of real world usecases. Previously I worked as an Applied Research Scientist Lead at Facebook AI research, working on building Multimodal foundational generative AI models. I am also interested in the domain of self supervised learning. I have published 100+ papers across top AI conferences like NeurIPS, CVPR, ACL, EMNLP, TMLR, ICLR, NAACL, COLM, EACL, WACV, Interspeech among others garnering over 15k+ citations. I routinely work with multi-billion scale datasets to train these massive multimodal models. In the past, I have also worked as Quantitative Researcher at Citadel where my work involved leveraging the power of Machine Learning and Statistical methods in an attempt to fathom the enigmatic world that is the financial markets. I have also worked at Amazon Alexa AI on large scale multimodal models and Embodied AI applications to bring smart robot intelligence to Alexa devices. I actively advise several early stage startups and often guest lecture at Stanford, CMU, MIT, Oxford, Oreilly among others. I graduated from Indian Institute of Technology, Kanpur completing my Bachelors in Computer Science and Engineering and then completed my graduate school in Machine Learning and Artificial Intelligence at the Language Technologies institute at Carnegie Mellon University.I am deeply passionate about research in my field. My research interests include: Multimodal foundation models and it's uses in the field of Computer Vision, Speech and Music Processing and Natural Language Processing. I have had the good fortune of working with some amazing people at some fantastic places and have learnt a lot from them. I hope to continue learning, travelling to new places and meeting new people. My mantra in life is- "Live life with passion - Love what you do, do what you love". Besides being a Technology enthusiast, I am also very passionate about sports. I was a part of the IIT Kanpur Aquatics team and love to swim, play Water Polo, Soccer and Cricket.

Country

United States

City

San Francisco Bay Area

Industry

Information Technology & Services

Skill

Generative AI, Large Language Models (LLM), Natural Language Processing (NLP), Multimodal machine learning, Deep Learning, Computer Vision, Speech Processing, Machine Learning, Data Mining, Software Development, Programming, Swimming, Artificial Intelligence, Algorithms, Computer Science, Data Structures, Image Processing, C++, C, Python

Experience

Pocket FM

Head of Artificial Intelligence

Pocket FM

LinkedIn
2026 - Present

San Francisco Bay Area

Building the worlds largest AI native media unicorn. Leading AI at Pocket Entertainment, a profitable, late-stage, Lightspeed-backed unicorn startup with ~500M$ ARR, 300M+ users across 20 countries and 100B+ minutes streamed in 2025. Built and scaled a 60-person AI org across GenAI Research, Applications, and Personalization to power a 75,000-title media catalog spanning audio, comics, and novels, where AI was foundational to content generation and user experience. - Directed the development and deployment of custom LLMs and diffusion models for long-form, multimodal storytelling, optimized for coherence, cultural adaptation, and multilingual output capable of creating hours of high fidelity audio/video content. - Creating the worlds best creative writing model allowing true democratization of long form content creation, enabling writers to generate hundreds of hours of creative narrative content while maintaining novelty, coherence and writing style. - Architected agentic content generation pipelines ingesting billions of tokens monthly to enable scalable, creator-augmented narrative production. - Working on creating the most performant multilingual narrative TTS model which will power our audio content library. - On the personalization side, led modeling efforts for real-time ranking, churn prediction, and LTV optimization, driving measurable uplifts in engagement, retention, and monetization. Our AI efforts were foundational to the companys growth from 200M$ ARR to 450M$ ARR in under a year while becoming profitable!

Startup Advisor

2022 - Present · 4 yrs

• I advise multiple startup founders and VCs on technical and AI strategies for scaling from ideation to growth stage • Advising companies in the domains of Multimodal Foundation model training/post training, AI evals, World models, RL environments, Data annotations, Synthetic data generation, Video generation models, Robotic foundation models, Audio foundation model training, Industrial Computer vision applications and horizontal RL stack creation. • Designed AI infrastructure and built AI native products to drive customer acquisition and growth. • Facilitated connections between founders and VCs and assisted in hiring top talent to support company expansion.

Tesla

Senior Staff Machine Learning Scientist

Tesla

LinkedIn
2025 - 2026 · 1 yr

Palo Alto, CA

Optimus Robotics team Working on Elon's vision to bring humanoid robots to life - Building causal, real time and dynamic World models for Closed loop RL training of Robotics Foundation models - Worked on designing the flagship Multimodal Foundational model (TRON) to enable multimodal reasoning capabilities for Optimus and Tesla FSD - Created a high fidelity auto labelling pipeline to enable automated annotations at scale - Improved throughput of DiT based video generation backbones by ~ 100x through use of efficient attention computations with barely any drop in visual fidelity powering both Autopilot and Optimus usecases. - Significantly improved visual quality on Wan3.x video generation models by enhancing the MoE setup by scaffolding high motion expert, specializing in generating visually dynamic sequences leading to higher visual fidelity in these long tail but crucial scenarios.

Meta

Applied Research Scientist Lead

Meta

LinkedIn
2022 - 2025 · 3 yrs

Menlo Park, CA

Working with Facebook AI Research (FAIR) as a part of Meta Superintelligence labs (MSL), on research on Large scale multimodal foundational models on trillion scale datasets. Particularly interested in Generative AI research and production use cases and exploring the realm of Self supervised learning. Published list of papers: - DINOv2: Learning Robust Visual Features without Supervision (Published at TMLR ) https://dinov2.metademolab.com/ - Chameleon: Mixed-Modal Early-Fusion Foundation Models (https://about.fb.com/news/2024/06/releasing-new-ai-research-models-to-accelerate-innovation-at-scale/) - Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning (CM3Leon) (https://ai.meta.com/blog/generative-ai-text-images-cm3leon/) - Demystifying CLIP Data (MetaCLIP) (Published at ICLR 2024): (https://github.com/facebookresearch/MetaCLIP) - Mavil: Masked audio-video learners (Published at NeurIPS 2023): (https://ar5iv.labs.arxiv.org/html/2212.08071) - A Picture is Worth More Than 77 Text Tokens: Evaluating CLIP-Style Models on Dense Captions (Published at CVPR 2024): (https://openaccess.thecvf.com/content/CVPR2024/papers/Urbanek_A_Picture_is_Worth_More_Than_77_Text_Tokens_Evaluating_CVPR_2024_paper.pdf) - Seamless Interaction (https://ai.meta.com/research/publications/seamless-interaction-dyadic-audiovisual-motion-modeling-and-large-scale-dataset/) - FLAP: Fast Language-Audio Pre-training (Published at ASRU 2023) (https://arxiv.org/abs/2311.01615) - An Introduction to Vision-Language Modeling (https://arxiv.org/abs/2405.17247) - Text Quality-Based Pruning for Efficient Training of Language Models (https://arxiv.org/pdf/2405.01582) and DMLR 2025 - Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM (Published at COLM 2024) (https://arxiv.org/abs/2403.07816)

Algoverse

AI Research Director

Algoverse

LinkedIn
2024 - 2024

Using my personal time to lead the development of a cutting-edge AI program to empower students with industry-relevant skills, leveraging top AI research experience from leading labs. • Our students got admitted to CMU, Stanford, MIT, UCBerkeley among others and several received job offers from top companies like Anthropic and Tiktok. • Collaborated with teams to create unparalleled opportunities in AI education, nurturing future innovators in the field allowing students as young as high school to publish to ICML, Neurips, EMNLP, EACL among others. Published papers: - FrontierScience Bench: Evaluating AI Research Capabilities in LLMs (ICML: REALM 2025) - Rosetta-PL: Propositional Logic as a Benchmark for Large Language Model Reasoning: https://arxiv.org/pdf/2505.00001 - FaceSafe: An Inpainting Pipeline for Privacy-Compliant Scalable Image Datasets (ICML 2025 : DIG-BUGS) - COREVQA: A Crowd Observation and Visual Entailment Visual Question Answering Benchmark (ICML 2025 : DIG-BUGS) - Pause-Tuning for Long-Context Comprehension: A Lightweight Approach to LLM Attention Recalibration (ICML 2025: LCFM 2025) - NovelHopQA: Diagnosing Multi-Hop Reasoning Failures in Long Narrative Contexts (ICML 2025: LCFM 2025) - Rewrite-to-Rank: Optimizing Ad Visibility via Retrieval-Aware Text Rewriting (ICML 2025: MoFA 2025) - TRUTH DECAY: Quantifying Multi-Turn Sycophancy in Language Models:https://arxiv.org/pdf/2503.11656 - Deconstructing bias: A multifaceted framework for diagnosing cultural and compositional inequities in text-to-image generative models: https://arxiv.org/pdf/2505.01430 - Advancing Uto-Aztecan Language Technologies: A Case Study on the Endangered Comanche Language: https://aclanthology.org/2025.americasnlp-1.4.pdf

Citadel

Quantitative Researcher (Non compete)

Citadel

LinkedIn
2021 - 2024 · 3 yrs

Chicago, IL

Citadel

Quantitative Researcher

Citadel

LinkedIn
2019 - 2021 · 2 yrs

Chicago

Demistifying the black box of financial markets through AI

Amazon Lab126

Applied Scientist

Amazon Lab126

LinkedIn
2021-8 - 2022-8 · 1 yr 1 mo

Sunnyvale, CA

Working with Alexa AI on a variety of problems like: - A dialog enabled visual-language navigation bot leveraging the multimodal data sources to faithfully navigate a virtual environment based on user instruction. Created a new benchmark for visual language navigation as a part of the Alexa Prize Simbot challenge and designed benchmark models for the same - Designing efficient multimodal transformers to speed up their training and deployment by improving the computational complexity of the self attention mechanism - Video processing applications like video action recognition, video question answering, video summarization, moment retrieval etc working directly with compressed video streams - Created a benchmark for cooperative heterogenous multi agent reinforcement learning platform including open sourcing the collected dataset and it's benchmark models - Working on creating a massively multimodal transformer pipeline capable of handling a wide range of input modalities with modality agnostic transformer blocks which work well for a several tasks leveraging a multitude of modalities Published list of papers: - Alexa arena: A user-centric interactive platform for embodied ai (Published at Neurips 2023) (https://www.amazon.science/publications/alexa-arena-a-user-centric-interactive-platform-for-embodied-ai) - Alexa, play with robot: Introducing the first Alexa Prize SimBot Challenge on embodied AI (https://www.amazon.science/alexa-prize/proceedings/alexa-play-with-robot-introducing-the-first-alexa-prize-simbot-challenge-on-embodied-ai) - CHMARL: A Multimodal Benchmark for Cooperative, Heterogeneous Multi-Agent Reinforcement Learning (Published at RSS 2022) (https://www.amazon.science/publications/chmarl-a-multimodal-benchmark-for-cooperative-heterogeneous-multi-agent-reinforcement-learning) - ε-ViLM: Efficient Video-Language Model via Masked Video Modeling with Semantic Vector-Quantized Tokenizer (Published at WACV 2024) (https://arxiv.org/abs/2311.17267)

Carnegie Mellon University

Graduate Research Assistant

Carnegie Mellon University

LinkedIn
2018-8 - 2019-8 · 1 yr 1 mo

Multicomp Lab, Language Technologies Institute, School of Computer Science

I worked with Prof. Loius Phillipe Morency on a wide range of project related to Multi Modal Machine Learning and on building robust, explainable Deep Learning models - We designed adversarial attack mechanisms on Visual Question Answering models to identify their vulnerabilities and then fix the same by proposing a variety of robust training mechanisms. - We worked on developing a neural network model which uses a Deep Convolutional Neural Network based pipeline alongside a geometrically conditioned point distribution model for Facial Landmark Detection. - We alsondeveloped the first fully ecologically validated models of visual perception. We will combine intracranial EEG (iEEG)recordings captured during long stretches of natural visual behavior with cutting-edge computer vision, machine learning, and statistical analyses to understand the neural basis of natural, real-world visual perception. - We also explored facial expression recognition in extreme face scenarios like profile face views, occluded faces, non centric and rotated faces alongside recognition for gender, age and racially diverse faces.

Citadel

Research Intern

Citadel

LinkedIn
2018-5 - 2018-8 · 4 mos

Greater Chicago Area

Working on using Machine Learning/ Deep Learning techniques to better model Financial Time series data and ensure scalability of the algorithms to arbitrary number of input features

Carnegie Mellon University

Graduate Research Assistant

Carnegie Mellon University

LinkedIn
2017-8 - 2018-5 · 10 mos

Articulab, Language Technologies Institute, School of Computer Science

I worked on the SARA and the Yahoo! InMind projects at the ArticuLab which focus on building a socially aware robotic assistant. My primary focus was on trying to combine the user’s visual, vocal and verbal cues to better gauge the ‘rapport’ between the user and the conversational agent and using it to enable the agent to become socially more aware to the user’s emotional needs.

Education

Carnegie Mellon University

Carnegie Mellon University

LinkedIn

PhD dropout Masters in Language Technologies (MLT program) Language Technologies Institute, School of Computer Science

Indian Institute of Technology, Kanpur

Indian Institute of Technology, Kanpur

LinkedIn

Computer Science and Engineering

Vasu Sharma's Contact Information

Email

******@***.com

Phone

(**) *** ****

Find the Right Leads
Find Verified Contact Data

Try with: Jensen Huang @ nvidia.com Click to autofill
LeadContact awards, five-star ratings, and GDPR compliance badges

What LeadContact does well

Find verified emails, phone numbers, and decision-makers with 98% accuracy.

Find Leads

Find Leads

Find the right people by company, role, industry, location, and more.

925M+ professional profiles

Find Leads
Find Emails

Find Emails

Access verified email addresses for your target contacts.

657M+ emails

Find Emails
Find Phone Numbers

Find Phone Numbers

Get cross-validated phone data from multiple top sources.

239M+ phone numbers

Find Phone Numbers

More Accurate. Lower Cost.

Find contact data in 1 tool with 98% accuracy

LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.

LeadContact Logo
Competitor Tools

All these = $289 per month

Great conversations start with the right contact.

It’s time to find yours.