Mark C.
Research Engineer @ TikTok
About
Former CMUer, THUer. Currently working on Large Language Models and multimodal foundation models. Continuosly to be amazed by ChatGPT and trying to figure out how it was made Interested in various topics including investment and quant
United States
San Jose
Information Technology & Services
Search, Generative AI, Natural Language Generation, multimodal, ChatGPT, deepspeed, megatron, Natural Language Processing (NLP), Large Language Models (LLM), Distributed Machine Learning, GPT, Python (Programming Language), C++, Python, Computer Vision, Machine Learning, Computer Science, Algorithms, C, Matlab
Experience

Research Engineer
Since 2024.09 I'm in charge of the Tiktok AI Search, which serves hundreds of millions of American users everyday. We aim to use LLM techniques to improve the experience of users. I worked on everything related to AI Search. ---- I work on LLM foundation models and code generation models. I have been working on pretraining, finetuning and evaluation of the models. Our product --- Doubao, has more than 30 million MAU right now. Our code generation model is on par with the SOTA opensource code models.

Research Engineer
CA, United States / Beijing, China
Fusion: Design and implement multi-sensor fusion models, algorithms, and pipelines. Using techniques like DeepLearning, XGboost, Graph Neural Network and etc via ML frameworks such as MXnet and Pytorch. Deploy algorithms and models onto our trucks using c++ with GPU support. Tracking Design and implement multi-frame multi-object tracking models,algorithms and pipelines. Using techniques like DeepLearning, XGboost, GCN, RNN and kalman filters Machine Learning Platform/Pipeline: Developed the first auto sampling, auto training, and auto benchmarking platform for perception modules—processing Terabytes of road test datasets daily and updating our models when deemed necessary by the algorithm.

Software Development Engineer
Palo Alto, California, United States
1. Designed and built a git-style data management system that provides data versioning and can be used both in the command line and raw python project. 2. Used recursive hashing to compare folder difference and enabled the system to track terabytes size datasets. 3. Made a central data sharing system that can push and pull(download) datasets using a single command. Greatly reduce the time needed for sharing datasets (instantly for pull). 4.Reduced 20~30% of disk usage on EFS by providing symlinks/refinks to users and save $12~20K per month for the team potentially. 5. Help to build one of Amazon Search's first Bert based relevance model
Education
Mark C.'s Contact Information
Phone
Find the Right Leads
Find Verified Contact Data
What LeadContact does well
Find verified emails, phone numbers, and decision-makers with 98% accuracy.
Find Leads
Find the right people by company, role, industry, location, and more.
925M+ professional profiles

Find Emails
Access verified email addresses for your target contacts.
657M+ emails

Find Phone Numbers
Get cross-validated phone data from multiple top sources.
239M+ phone numbers

More Accurate. Lower Cost.
Find contact data in 1 tool with 98% accuracy
LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.
Great conversations start with the right contact.
It’s time to find yours.




