Yashraj Jadhav
Student Assistant – International Services @ University of Southern California
About
Data Scientist and ML Engineer with 2+ years of industry experience building end-to-end machine learning systems that drive measurable business outcomes. At Capgemini, I built classification models on 500K+ records, designed causal inference pipelines supporting $4.2M pharmaceutical investment decisions, and applied clustering and association rule mining on 2.5TB of data to reduce misdiagnosis rates by 23%. Currently building production-grade ML systems — an ICU census forecasting system deployed on GCP Cloud Run, a churn prediction model with causal uplift measurement, and LLM-based agent systems using LangGraph, Gemini API, and RAG. Core stack: Python · SQL · PyTorch · scikit-learn · XGBoost · LangChain · FastAPI · Docker · GCP Actively seeking Data Scientist and ML Engineer roles at companies where I can own end-to-end model development and drive measurable impact.
United States
Los Angeles
Information Technology & Services
Gemini API, Large Language Models (LLMs), Microsoft Power BI, Pandas (Software), NumPy, Retrieval-Augmented Generation (RAG), Google Cloud Platform (GCP), LangGraph, LangChain, Machine Learning, Python, Data Science, SQL, Statistical Modeling, Scikit-Learn, XGBoost, Random Forest, PyTorch, Transformers, Classification
Experience

Data Scientist
Pune, Maharashtra, India
• Built a patient readmission classifier comparing Random Forest vs XGBoost on 500K+ records in scikit-learn; handled ~15% positive-class imbalance using stratified k-fold cross-validation and selected the final model by optimizing Precision-Recall AUC over default accuracy, cutting readmission rates by 18% across three hospital departments • Designed a causal inference pipeline for pharmaceutical trial evaluation: defined treatment/control cohorts, ran stratified subgroup analysis across demographic and clinical segments, and computed effect sizes with confidence intervals to support $4.2M go/no-go investment decisions • Applied K-Means clustering (elbow method + silhouette scoring) and Apriori association rule mining on 2.5TB of diagnostic records using Python and SQL; surfaced high-confidence co-occurrence patterns that reduced misdiagnosis rates by 23% • Engineered SQL-based feature extraction and ETL pipelines processing 50K+ daily records with schema validation, null-rate monitoring, and incremental loads; improved pipeline throughput by 35% and downstream feature quality by 28% • Delivered model predictions (risk scores, cluster assignments, subgroup effects) via Power BI decision dashboards to clinical executives; tied every visual to an upstream model output, reducing operational costs by 15%

Data Science Intern
Pune, Maharashtra, India
• Profiled 300K+ patient records using Python and SQL to assess feature distributions, cardinality, null rates, and class balance; findings directly shaped feature selection and encoding decisions for the readmission and clustering models built in the subsequent full-time role • Conducted hypothesis testing and statistical analysis on patient safety event logs across 3 clinical units, identifying significant outcome patterns that improved clinical protocol adherence by 22%
Education
Yashraj Jadhav's Contact Information
Phone
Find the Right Leads
Find Verified Contact Data
What LeadContact does well
Find verified emails, phone numbers, and decision-makers with 98% accuracy.
Find Leads
Find the right people by company, role, industry, location, and more.
925M+ professional profiles

Find Emails
Access verified email addresses for your target contacts.
657M+ emails

Find Phone Numbers
Get cross-validated phone data from multiple top sources.
239M+ phone numbers

More Accurate. Lower Cost.
Find contact data in 1 tool with 98% accuracy
LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.
Great conversations start with the right contact.
It’s time to find yours.







