Chirag Madhukar

Chirag Madhukar

Data Engineer @ Goldman Sachs

About

Motivated and dedicated to providing efficient, prudential decisions to help solve complex problems I am a results-driven data engineering and analytics leader with 5+ years of experience converting raw data into actionable insights, data modeling, and complex ETL pipeline development. Expertise extends to implementing end-to-end analytics solutions for cross-functional teams utilizing cloud platforms and big data tools. Proficient in data visualization, adept at delivering insightful narratives tailored for the SLT I graduated from the University of California, Irvine - The Paul Merage School of Business with an MS in Business Analytics. I previously worked at Capital One as a Senior Data Analyst, and I am currently employed as a Data Engineer at Goldman Sachs Alongside a strong programming background and a flair for business, management, and analytics, I possess good leadership, analytical, and decisive problem-solving skills. My career interests lie in the field of Data Science, Data Engineering, and Analytics I am deeply passionate about working with data and thrive in an innovative, data-driven environment that allows me to utilize my skills and elevate my career as a Data Engineer/Data Scientist while helping the organization reach its ultimate goal Feel free to reach me at chirag.madh09@gmail.com Skills Programming Languages: Python, SQL, R, Scala, HiveQL, Java, C, C++, JavaScript, HTML, CSS Libraries: Pandas, NumPy, SciPy, Scikit-Learn, TensorFlow, PyTorch, Matplotlib, Seaborn, Plotly Data Engineering & ETL Data Warehousing & Pipelines: Snowflake, Databricks, Airflow, Redshift, BigQuery, dbt Big Data & Streaming: Hadoop, Spark (PySpark, Scala), Kafka, Presto, Hive, Cassandra ETL & Data Processing: SQL transformations, Python ETL scripts, Alteryx Analytics & BI Data Visualization: Tableau, Power BI, Jupyter Notebook Business Tools: Salesforce, MS Excel (Power Query, VBA, PivotTables) Statistics & ML: Hypothesis Testing, A/B Testing, Predictive Modeling Machine Learning & AI Supervised & Unsupervised Learning: Classification, Clustering, Recommendation Systems Models: Linear & Logistic Regression, Random Forest, XGBoost, K-Means, PCA Deep Learning & NLP: Neural Networks, Transformers, Sentiment Analysis Cloud ML Platforms: AWS SageMaker Cloud & Infra Cloud Platforms: AWS (S3, Lambda, Redshift, Glue), GCP (BigQuery, Dataflow) DevOps & Deployment: Docker, Kubernetes, CI/CD Pipelines Project Management & Agile Agile Tools: JIRA, Scrum, Kanban Data Strategy: Business Intelligence, Data Governance, Data Modeling (Kimball, Star Schema)

Country

United States

City

New York City Metropolitan Area

Industry

Financial Services

Skill

Python (Programming Language), SQL, Snowflake, Databricks, Tableau, Apache Airflow, PySpark, Continuous Integration and Continuous Delivery (CI/CD), Software Development Life Cycle (SDLC), Gitlab, GitHub, YAML, Amazon Web Services (AWS), Amazon EC2, JavaScript, Apache Spark, Amazon S3, NetBeans, Java, Hadoop

Experience

Goldman Sachs

Data Engineer

Goldman Sachs

LinkedIn
2024-5 - Present · 2 yrs 5 mos

New York, New York, United States

• Platform Solutions/Consumer Banking: Data Engineering & Card Analytics (Consumer Card BI Engineering Team) • Architected a centralized Airflow-based observability framework (meta-DAG) monitoring 50+ BI pipelines across Apple Card and Savings, enforcing SLA compliance, automating PagerDuty alerts; reduced on-call MTTR and cut silent reporting failures by 90% • Engineered scalable data pipelines in Databricks, PySpark, and Snowflake to integrate raw and curated layers; delivered customer behavior, risk, and KPI insights for Apple Card and Savings, eliminating 40+ hours/​month of manual reporting • Orchestrated ETL workflows in PySpark and SQL via Airflow DAGs to deliver tranche-level data across 18+ GM Card segments; enabled dashboards that supported a $285M+ loss-account sale and seamless BAU portfolio transfer to Barclays •Implemented reconciliation controls and data validation frameworks in SQL and PySpark (unit tests, control reports) to ensure integrity between Apple Savings and Apple Card; improved consistency by 40% and reduced reporting discrepancies by 25% • Developed and productionized end-of-month adoption score pipelines using XGBoost and PySpark to segment 5M+ customers into deciles; operationalized scoring outputs to S3 and Snowflake to power targeted retention strategies • Optimized Airflow-triggered workloads by auditing and tagging DAG queries and tuning PySpark and Snowflake jobs in Databricks, boosting observability and query efficiency, resulting in a 30% reduction in monthly compute costs • Contributed to the Top of Wallet (TOW) scoring framework using LightGBM, generating primary payment likelihood scores and SHAP reason codes; operationalized outputs to power personalized marketing campaigns • Migrated processes from username-password to key-pair authentication for all Spark and Snowflake connectors, boosting security and reducing vulnerabilities by 80%

Capital One

Senior Data Analyst

Capital One

LinkedIn
2022-10 - 2023-12 · 1 yr 3 mos

McLean, Virginia, United States

• Sr. Associate - Data Analytics & Engineering on the Business Cards and Payments Team (Small Business Cards and Commercial Cards) • Built a governed analytics layer on Snowflake and Databricks using PySpark and SQL, integrating 40+ internal sources for Small Business & Commercial Cards; standardized schemas, SCD handling, and data validations (row-level QA, reconciliation), improving data quality and time-to-insight for downstream analytics • Productionized KPI pipelines in Airflow using dbt-style patterns for acquisition, funnel, and portfolio metrics with SLAs, alerting, and backfills; reduced manual MBR prep by ~10 hours/​month and improved reporting consistency • Delivered interactive Tableau Dashboards integrated with Salesforce, enabling sales teams to analyze client portfolios, resulting in a 12% YoY boost in customer spending and earning SVP recognition for data-driven insights • Implemented campaign eligibility marts (Python/​PySpark) to exclude non-safelisted customers; reduced compliance risk by $3M+/​quarter and improved marketing efficiency by 15% • Architected a reusable ETL automation framework in Python and PySpark, improving scalability (30% faster runtimes) and reducing manual coding effort by 40%, accelerating adoption across teams

Blue Shield of California

Data Science Intern

Blue Shield of California

LinkedIn
2022-1 - 2022-8 · 8 mos

Irvine, California, United States

•Built Python web scrapers (API requests, BeautifulSoup) to ingest and normalize NPI registry data across 100K+ providers, integrating it with claims data to reduce manual prep by ~50% and support executive dashboards • Executed data cleaning, manipulation, and feature engineering of over 9M rows of healthcare claims data using Pandas to provide a framework for descriptive analytics and advanced reporting capabilities • Performed hypothesis testing (t-tests) in AWS SageMaker to validate cost-saving potential of Ambulatory Surgical Centers, identifying $400+ savings per outpatient procedure • Identified prevalent comorbidities using ICD-10 diagnostic codes, analyzed their patterns, and implemented long-term cost-saving strategies for affected patients, resulting in a 10% reduction in annual patient spending • Built interactive Tableau Dashboards using global filters, LOD expressions, and advanced charts by mapping web-scraped NPI registry data to identify regions with elevated comorbidity rates, enabling rapid insights for the C-Suite on patient data

Start-Up

Data Engineer

Start-Up

2019-8 - 2021-7 · 2 yrs

Bangalore Urban, Karnataka, India

• Implemented data quality gates in Airflow with Python checks and Great Expectations–style validations (row-count, null, referential integrity); blocked ~200+ bad loads annually and reduced downstream defects by 35%. • Introduced CDC ingestion pipelines from operational stores into S3 and Snowflake using Kafka and Spark Structured Streaming; enabled near-real-time dashboards and reduced data freshness lag from ~6 hours to under 10 minutes. • Enhanced Airflow DAGs with observability (retry/backoff, SLA alerts, lineage tagging), raising on-time delivery to 99%+ and cutting incident resolution time by 40% with better logs and metrics. • Established versioned data contracts between producers and consumers via schema registry, reducing breaking changes by 70% and shortening analyst onboarding from weeks to days

Tata Consultancy Services

Data Engineer Intern

Tata Consultancy Services

LinkedIn
2018-6 - 2018-7 · 2 mos

Bangalore Urban, Karnataka, India

• Data warehousing/Big Data: Created accelerators for various components in the platform • Assisted in the design and development of ETL (Extract, Transform, Load) processes to efficiently transfer data from various sources to a centralized data repository • Collaborated with the data engineering team to design and implement ETL processes, with a specific emphasis on integrating MongoDB and Oracle databases into the data processing workflows. • Participated in the creation and maintenance of data pipelines, ensuring smooth and reliable data flow for business analytics and reporting • Conducted data profiling and cleansing activities to identify and rectify inconsistencies, missing values, and outliers in datasets

Education

University of California, Irvine - The Paul Merage School of Business

University of California, Irvine - The Paul Merage School of Business

LinkedIn

Business Analytics (Data Science)

Relevant Coursework: Predictive Analytics, Machine Learning for Analytics, Statistics for Data Science, Data and Programming for Analytics, Natural Language Processing, Business Data Management, Management Science for Analytics, The Art and Science of Applied Forecast Modeling

PES University

PES University

LinkedIn

Computer Science - Specialization in Data Science

Relevant Coursework: Introduction to Data Science, Design Analysis & Algorithms, Data Structures, Machine Learning, Database Management Systems, Data Analytics, Computer Networks, Python Programming

National Hill View Public School

National Hill View Public School

LinkedIn

Physics, Chemistry, Mathematics and Computer Science

Chirag Madhukar's Contact Information

Email

******@***.com

Phone

(**) *** ****

Find the Right Leads
Find Verified Contact Data

Try with: Jensen Huang @ nvidia.com Click to autofill
LeadContact awards, five-star ratings, and GDPR compliance badges

What LeadContact does well

Find verified emails, phone numbers, and decision-makers with 98% accuracy.

Find Leads

Find Leads

Find the right people by company, role, industry, location, and more.

925M+ professional profiles

Find Leads
Find Emails

Find Emails

Access verified email addresses for your target contacts.

657M+ emails

Find Emails
Find Phone Numbers

Find Phone Numbers

Get cross-validated phone data from multiple top sources.

239M+ phone numbers

Find Phone Numbers

More Accurate. Lower Cost.

Find contact data in 1 tool with 98% accuracy

LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.

LeadContact Logo
Competitor Tools

All these = $289 per month

Great conversations start with the right contact.

It’s time to find yours.