Chen Chen
Machine Learning Engineer @ TikTok
About
As a Machine Learning Engineer at TikTok/Bytedance, my work connects evaluation-driven workflows with data curation, model training, and fine-tuning, resulting in improved video quality and performance. Leveraging expertise in reinforcement learning, LLMs, and VLM-based evaluators, I contribute to enhancing content generation capabilities for thousands of daily users. Previously, I held roles as a Data Scientist at Tiktok and Seagen, where I developed innovative solutions such as multimodal classifiers, large language model workflows, and natural language-driven tools. My passion lies in transforming data-driven insights into impactful, scalable solutions to address complex challenges in machine learning and AI.
United States
San Jose
Internet
Positive Reinforcement Learning, ChatGPT, Named Entity Recognition (NER), Data Monetization, MLOps, Text-to-Image Generation, Post-editing, Work Sampling, Model-view-viewmodel (MVVM), Apache Spark ML, Voice Search Optimization, Generative AI, Knowledge Extraction, Data Science, Applied Linguistics, Machine Learning, Multi-level Communication, Core ML, Query Languages, Transformer Models
Experience

Machine Learning Engineer
‣ Drove evaluation-driven data flywheels for production text-to-video generation models, connecting model evaluation, bad-case mining, data curation, human review, post-training, and re-evaluation across 20+ model iterations. ‣ Trained and helped productionize one-stage and two-stage text-to-video post-training models, including direct 480p/720p generation and cascaded 480p-to-720p super-resolution, powering 3K+ daily generations after launch. ‣ Fine-tuned a Qwen-7B-based LLM prompt rewriting model for deployed text-to-video generation, reducing bad-case rate by 13% compared with GPT-4.1 and improving online A/B export rate by 4.28%. ‣ Developed and productionized VLM-based video quality evaluators for text-to-video, image-to-video, and reference-to-video workflows, improving recall by +0.42 on human distortion and +0.36 on physics violation over closed-source VLM baselines while enabling screening of 50K generated videos/day. ‣ Designed agentic video generation and editing workflows for AI short-drama production, including automated video evaluation, repair/regeneration decisions, VLM/ASR-based post-editing, and RAG-based voice retrieval services for natural-language voice matching.

Data Scientist
美国
‣ Built a travel content understanding taxonomy using named entity recognition (NER) and KMeans clustering, boosting label coverage from 51% to 78%; deployed a multimodal travel ads classifier with weighted tier-1 precision above 0.80. ‣ Led training-data preparation and quality strategy across 15 labeling projects and 71 queues, supporting 2M+ high-precision samples while saving 870 labeling tasks, 11,312 review hours, and $485K through sampling optimization.

Data Scientist
西雅图, WA
‣ Developed a GPT-powered chatbot using Streamlit that enables users to search clinical trial database using natural language to assist with clinical trial and regulatory submission planning. ‣ Utilized prompt engineering with GPT-4 to extract clinical trial information from 700+ FDA oncology label documents. Build data pipelines to represent extracted data as knowledge graphs in Neo4j. Achieved significant cost savings for clinical trial data collection utilizing LLM workflows. ‣ Crafted a Python-based data visualization tool that automatically query data in SQL from Azure Databricks with PySpark and generates visuals based on user requests in natural language. The tool won first prize in the Research Data Hub Hackathon at Seagen.

Data Science Ambassador
Tucson, Arizona, United States
My role includes: 1. Helping researchers work on data management plans. 2. Addressing questions related to data science through consults, training, and/or referrals to the right resources and experts on-campus or online.

Graduate Research Associate
Tucson, Arizona, United States
My role includes: 1. Maintaining language extraction rules using the Odin event extraction framework to extract task-relevant events from natural language utterances. 2. Using Python to display the language hierarchy extracted from human team communications. 3. Training language annotators with other team members and scheduling weekly meetings to identify and resolve annotation issues.
Education
Chen Chen's Contact Information
Phone
Find the Right Leads
Find Verified Contact Data
What LeadContact does well
Find verified emails, phone numbers, and decision-makers with 98% accuracy.
Find Leads
Find the right people by company, role, industry, location, and more.
925M+ professional profiles

Find Emails
Access verified email addresses for your target contacts.
657M+ emails

Find Phone Numbers
Get cross-validated phone data from multiple top sources.
239M+ phone numbers

More Accurate. Lower Cost.
Find contact data in 1 tool with 98% accuracy
LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.
Great conversations start with the right contact.
It’s time to find yours.







