Yashu L.

Yashu L.

Senior Machine Learning Engineer @ Stripe

About

Data Scientist with passion for constantly learning new things in advanced machine learning. I have experience in applying machine learning /predictive modeling algorithms in fraud and credit industry. Have experience working with large datasets , analyzing and building models using spark. Languages : C++, SQL, Python Data mining tools: R ,Rapid miner Big data tools: Hadoop, Hive, Spark,Scala, AWS Visualization software: Tableau, R Shiny

Country

United States

City

Seattle

Industry

Internet

Skill

RapidMiner, R, SQL, Python, JMP, Matlab, C++, Business Analysis, Market Research, C, Microsoft Excel, Chemical Engineering, Process Simulation, Data Mining, Machine Learning, Stata, Amazon Web Services (AWS)

Experience

Stripe

Senior Machine Learning Engineer

Stripe

LinkedIn
2025-6 - Present · 1 yr 4 mos

Seattle, Washington, United States

Work on advance ML applications to tackle fraud at Stripe. 1. Designed a system to annotate the payments with anomaly attributes such as Identity mismatch and gibberish/enumerated email using GPT5 and effective prompt engineering which would be used in the internal rules to block fraudulent transactions. 2. Built an LLM agent to create fraud or non_fraud explanations about the payment. These explanations were used to finetune an LLM using LoRA to predict with reasoning if a payment is fraudulent. 3. Built an LLM as a judge to evaluate the clustering annotation of the suspicious payments. This led to gaining confidence in the clustering annotation leading to production launch of the product. 4. Designed ML Ops and monitoring system for the card testing detection model by creating Prometheus alerts, model performance dashboard and setting a strict SLO for model refresh and rollouts. 5. Feature-engineered and productionalized a sequence classification model using the Llama sequence classification architecture to predict whether a sequence of transactions was undergoing a card testing attack, outperforming the existing solution by over 500 bps in recall resulting in 10% reduction in card testing attacks over the existing detection system. The model is set up for asynchronous scoring using queues. 6. Productionalized the payments foundational model which is a masked LLM with contrastive loss fine tuning. The embeddings of the model are being used for multiple downstream tasks to detect fraud.

Stripe

Senior Data Scientist

Stripe

LinkedIn
2021-11 - 2025-6 · 3 yrs 8 mos

Greater Seattle Area

1. Developed and trained a doubly robust causal model to predict incremental churn following a risk intervention and implemented the training and evaluation code in flyte to automate the process. The model is integrated into production within a risk framework that determines intervention strategies. This model results in $3M incremental revenue for the company. 2. Created the equations used in the risk framework that determines intervention strategies and developed a method to evaluate the effectiveness of the framework itself. The framework considers both the riskiness and the value of a merchant to inform intervention strategies, rather than relying solely on risk 3. Engineered second-hop graph features and trained an identity graph fraud model(XGBoost), achieving an 8% increase in recall at a fixed precision 4. Experimented with various deep neural network architectures, including ResNet, FT-Transformer, and MLP, leading to the productionalization of the FT-Transformer architecture for predicting fraud in SEPA transactions. Both ResNet and Transformer models achieved over 100 bps improvement in recall at a fixed precision over existing solutions

Google

Data scientist

Google

LinkedIn
2019-6 - 2021-11 · 2 yrs 6 mos

Sunnyvale, California

Provided data driven solutions to accelerate revenue of Google Cloud Platform through sales and marketing: 1. Implemented and trained an ensemble of XGBoost classification model and RNN to prioritize GCP free trial leads to sales reps, resulting in 25% higher AUC. This resulted in ~30% increase in the conversion of free trials to paid users. 2. Designed and analyzed A/B experiments to understand the impact of the new GCP free trial sales prioritization model. 3. Created a tensorflow pipeline and trained multiple uplift models(GBDT, DNN) for email marketing 4. Trained a startup targeting model for GCP using boosted positive - unlabeled learning, for sales outreach resulting in increased sales in the product. 5. Developed an Apache Beam pipeline and python package to extract the named entities, summarize the news articles and map to GCP’s existing customers to provide insights to sales. 6. Created GCP free trial customer segments using unsupervised techniques to device targeted marketing initiatives. 7. Developed pipelines in SQL and python packages for training and scoring the ML models

ID Analytics

Data Scientist

ID Analytics

LinkedIn
2018-7 - 2019-6 · 1 yr

Greater San Diego Area

Worked on creating features and training tree based models to predict synthetic identity fraud

ID Analytics

Associate Data Scientist

ID Analytics

LinkedIn
2016-6 - 2018-6 · 2 yrs 1 mo

San Diego, California

Provide data driven, identity theft solutions to banks, telecom and retail clients. Majority of the work includes: • Building predictive models to detect fraudsters and to provide credit scores for the clients • Performing data studies to understand the anomalies in the score distributions and performance • Work on research and development projects to enhance the IDA products and work towards building new products Projects: 1. Developed boosted decision tree models to predict third party fraud in the retail industry which resulted in 14% higher fraud detection rate 2. Engineered email variables using Pyspark and trained DNN and GBDT, resulting in 2-3% performance gain 3. Trained CNNs with numerical data to predict credit risk applications, resulting in 5% higher KS 4. Implemented One Class SVMs, Auto encoders and other unsupervised algorithms to detect synthetic frauds as part of research and development 5. Built data pipelines in PySpark to clean and generate data for model building.

Leading Consulting Firm

Data Science Intern

Leading Consulting Firm

2016-1 - 2016-5 · 5 mos

Greater Minneapolis-St. Paul Area

Objective is to develop a reference architecture using AWS, and a working prototype to demonstrate how to harness the power of IoT. During the course of intern I will be working on: ● Integration of simulated data streams using real world parameters (of an equipment) to AWS Kinesis and storing the data in Dynamodb ● Build a static predictive model using simulated static data and implement on simulated streaming data to detect failures using aws lambda

Orbees Business Solutions Pvt. Ltd.

Associate Consultant

Orbees Business Solutions Pvt. Ltd.

LinkedIn
2013-7 - 2015-4 · 1 yr 10 mos

Hyderabad, Telangana

Provided diverse set of strategic and data driven marketing and sales strategy solutions to medical devices and pharmaceutical companies, enabling them to venture into emerging markets. • Designed strategic initiatives for penetration of a medical device company by clustering (segmenting), analyzing, and visualizing hospital and national level therapy data • Performed various hypothesis tests on physicians survey data to validate hypothesis around market entry and developed market development and penetration strategies for key emerging markets with the help of prioritized actionable activities and market models for a leading medical devices company • Forecasted the revenue for 10 years, laid out detailed investment analysis for an emerging market to estimate the investment and the ROI for market penetration • Performed conjoint analysis and designed adjacency mapping to prioritize the technology pipeline in neurovascular and neurosurgery space, which led to acquisition of new technologies Have a comprehensive knowledge of various disease spaces such as neurovascular, orthopedics and peripheral artery diseases. Additionally, have experience in several developed and emerging markets including US, EU5 (mainly Germany and UK), Brazil, India, Taiwan and Mexico

Indian Institute of Petroleum

Research Intern

Indian Institute of Petroleum

LinkedIn
2012-5 - 2012-6 · 2 mos

Dehradun

Developed a reformer model based on kinetic rate equation Performed simulation to yield data of catalytic reformer using MATLAB

Indian Institute of Science (IISc)

Research Intern

Indian Institute of Science (IISc)

LinkedIn
2011-5 - 2011-6 · 2 mos

Bengaluru Area, India

Education

UMN Carlson School of Management

UMN Carlson School of Management

LinkedIn

Master of science in Business Analytics

2015 - 2016 · 1 yr

The Master of Science in Business Analytics program imparts a specialized combination of analytical know-how and business acumen. The curriculum shapes professionals who leverage big data to identify business opportunities, and apply their findings to take action. Engage with leading faculty whose expertise in business fundamentals and information systems positions you to derive invaluable insights from big data.

Indian Institute of Technology (Banaras Hindu University), Varanasi

Indian Institute of Technology (Banaras Hindu University), Varanasi

LinkedIn

Chemical Engineering

2009 - 2013 · 4 yrs

Department of Industrial Chemistry was established in 1921 at Banaras Hindu University. Subsequently, it was renamed as the Department of Chemical Engineering and Technology in 1956. The Department has established several benchmarks of achievements in teaching and research. It modernizes its programmes to impart education in upcoming areas of chemical engineering. Currently it is ranked among the nation's top departments in chemical engineering.

National Public School

National Public School

LinkedIn

Yashu L.'s Contact Information

Email

******@***.com

Phone

(**) *** ****

Find the Right Leads
Find Verified Contact Data

Try with: Jensen Huang @ nvidia.com Click to autofill
LeadContact awards, five-star ratings, and GDPR compliance badges

What LeadContact does well

Find verified emails, phone numbers, and decision-makers with 98% accuracy.

Find Leads

Find Leads

Find the right people by company, role, industry, location, and more.

925M+ professional profiles

Find Leads
Find Emails

Find Emails

Access verified email addresses for your target contacts.

657M+ emails

Find Emails
Find Phone Numbers

Find Phone Numbers

Get cross-validated phone data from multiple top sources.

239M+ phone numbers

Find Phone Numbers

More Accurate. Lower Cost.

Find contact data in 1 tool with 98% accuracy

LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.

LeadContact Logo
Competitor Tools

All these = $289 per month

Great conversations start with the right contact.

It’s time to find yours.