Chirag Madhukar
Data Engineer @ Goldman Sachs
About
Motivated and dedicated to providing efficient, prudential decisions to help solve complex problems I am a results-driven data engineering and analytics leader with 5+ years of experience converting raw data into actionable insights, data modeling, and complex ETL pipeline development. Expertise extends to implementing end-to-end analytics solutions for cross-functional teams utilizing cloud platforms and big data tools. Proficient in data visualization, adept at delivering insightful narratives tailored for the SLT I graduated from the University of California, Irvine - The Paul Merage School of Business with an MS in Business Analytics. I previously worked at Capital One as a Senior Data Analyst, and I am currently employed as a Data Engineer at Goldman Sachs Alongside a strong programming background and a flair for business, management, and analytics, I possess good leadership, analytical, and decisive problem-solving skills. My career interests lie in the field of Data Science, Data Engineering, and Analytics I am deeply passionate about working with data and thrive in an innovative, data-driven environment that allows me to utilize my skills and elevate my career as a Data Engineer/Data Scientist while helping the organization reach its ultimate goal Feel free to reach me at chirag.madh09@gmail.com Skills Programming Languages: Python, SQL, R, Scala, HiveQL, Java, C, C++, JavaScript, HTML, CSS Libraries: Pandas, NumPy, SciPy, Scikit-Learn, TensorFlow, PyTorch, Matplotlib, Seaborn, Plotly Data Engineering & ETL Data Warehousing & Pipelines: Snowflake, Databricks, Airflow, Redshift, BigQuery, dbt Big Data & Streaming: Hadoop, Spark (PySpark, Scala), Kafka, Presto, Hive, Cassandra ETL & Data Processing: SQL transformations, Python ETL scripts, Alteryx Analytics & BI Data Visualization: Tableau, Power BI, Jupyter Notebook Business Tools: Salesforce, MS Excel (Power Query, VBA, PivotTables) Statistics & ML: Hypothesis Testing, A/B Testing, Predictive Modeling Machine Learning & AI Supervised & Unsupervised Learning: Classification, Clustering, Recommendation Systems Models: Linear & Logistic Regression, Random Forest, XGBoost, K-Means, PCA Deep Learning & NLP: Neural Networks, Transformers, Sentiment Analysis Cloud ML Platforms: AWS SageMaker Cloud & Infra Cloud Platforms: AWS (S3, Lambda, Redshift, Glue), GCP (BigQuery, Dataflow) DevOps & Deployment: Docker, Kubernetes, CI/CD Pipelines Project Management & Agile Agile Tools: JIRA, Scrum, Kanban Data Strategy: Business Intelligence, Data Governance, Data Modeling (Kimball, Star Schema)
United States
New York City Metropolitan Area
Financial Services
Python (Programming Language), SQL, Snowflake, Databricks, Tableau, Apache Airflow, PySpark, Continuous Integration and Continuous Delivery (CI/CD), Software Development Life Cycle (SDLC), Gitlab, GitHub, YAML, Amazon Web Services (AWS), Amazon EC2, JavaScript, Apache Spark, Amazon S3, NetBeans, Java, Hadoop
Experience

Data Engineer
New York, New York, United States
• Platform Solutions/Consumer Banking: Data Engineering & Card Analytics (Consumer Card BI Engineering Team) • Architected a centralized Airflow-based observability framework (meta-DAG) monitoring 50+ BI pipelines across Apple Card and Savings, enforcing SLA compliance, automating PagerDuty alerts; reduced on-call MTTR and cut silent reporting failures by 90% • Engineered scalable data pipelines in Databricks, PySpark, and Snowflake to integrate raw and curated layers; delivered customer behavior, risk, and KPI insights for Apple Card and Savings, eliminating 40+ hours/month of manual reporting • Orchestrated ETL workflows in PySpark and SQL via Airflow DAGs to deliver tranche-level data across 18+ GM Card segments; enabled dashboards that supported a $285M+ loss-account sale and seamless BAU portfolio transfer to Barclays •Implemented reconciliation controls and data validation frameworks in SQL and PySpark (unit tests, control reports) to ensure integrity between Apple Savings and Apple Card; improved consistency by 40% and reduced reporting discrepancies by 25% • Developed and productionized end-of-month adoption score pipelines using XGBoost and PySpark to segment 5M+ customers into deciles; operationalized scoring outputs to S3 and Snowflake to power targeted retention strategies • Optimized Airflow-triggered workloads by auditing and tagging DAG queries and tuning PySpark and Snowflake jobs in Databricks, boosting observability and query efficiency, resulting in a 30% reduction in monthly compute costs • Contributed to the Top of Wallet (TOW) scoring framework using LightGBM, generating primary payment likelihood scores and SHAP reason codes; operationalized outputs to power personalized marketing campaigns • Migrated processes from username-password to key-pair authentication for all Spark and Snowflake connectors, boosting security and reducing vulnerabilities by 80%

Senior Data Analyst
McLean, Virginia, United States
• Sr. Associate - Data Analytics & Engineering on the Business Cards and Payments Team (Small Business Cards and Commercial Cards) • Built a governed analytics layer on Snowflake and Databricks using PySpark and SQL, integrating 40+ internal sources for Small Business & Commercial Cards; standardized schemas, SCD handling, and data validations (row-level QA, reconciliation), improving data quality and time-to-insight for downstream analytics • Productionized KPI pipelines in Airflow using dbt-style patterns for acquisition, funnel, and portfolio metrics with SLAs, alerting, and backfills; reduced manual MBR prep by ~10 hours/month and improved reporting consistency • Delivered interactive Tableau Dashboards integrated with Salesforce, enabling sales teams to analyze client portfolios, resulting in a 12% YoY boost in customer spending and earning SVP recognition for data-driven insights • Implemented campaign eligibility marts (Python/PySpark) to exclude non-safelisted customers; reduced compliance risk by $3M+/quarter and improved marketing efficiency by 15% • Architected a reusable ETL automation framework in Python and PySpark, improving scalability (30% faster runtimes) and reducing manual coding effort by 40%, accelerating adoption across teams

Data Science Intern
Irvine, California, United States
•Built Python web scrapers (API requests, BeautifulSoup) to ingest and normalize NPI registry data across 100K+ providers, integrating it with claims data to reduce manual prep by ~50% and support executive dashboards • Executed data cleaning, manipulation, and feature engineering of over 9M rows of healthcare claims data using Pandas to provide a framework for descriptive analytics and advanced reporting capabilities • Performed hypothesis testing (t-tests) in AWS SageMaker to validate cost-saving potential of Ambulatory Surgical Centers, identifying $400+ savings per outpatient procedure • Identified prevalent comorbidities using ICD-10 diagnostic codes, analyzed their patterns, and implemented long-term cost-saving strategies for affected patients, resulting in a 10% reduction in annual patient spending • Built interactive Tableau Dashboards using global filters, LOD expressions, and advanced charts by mapping web-scraped NPI registry data to identify regions with elevated comorbidity rates, enabling rapid insights for the C-Suite on patient data

Data Engineer
Start-Up
Bangalore Urban, Karnataka, India
• Implemented data quality gates in Airflow with Python checks and Great Expectations–style validations (row-count, null, referential integrity); blocked ~200+ bad loads annually and reduced downstream defects by 35%. • Introduced CDC ingestion pipelines from operational stores into S3 and Snowflake using Kafka and Spark Structured Streaming; enabled near-real-time dashboards and reduced data freshness lag from ~6 hours to under 10 minutes. • Enhanced Airflow DAGs with observability (retry/backoff, SLA alerts, lineage tagging), raising on-time delivery to 99%+ and cutting incident resolution time by 40% with better logs and metrics. • Established versioned data contracts between producers and consumers via schema registry, reducing breaking changes by 70% and shortening analyst onboarding from weeks to days

Data Engineer Intern
Bangalore Urban, Karnataka, India
• Data warehousing/Big Data: Created accelerators for various components in the platform • Assisted in the design and development of ETL (Extract, Transform, Load) processes to efficiently transfer data from various sources to a centralized data repository • Collaborated with the data engineering team to design and implement ETL processes, with a specific emphasis on integrating MongoDB and Oracle databases into the data processing workflows. • Participated in the creation and maintenance of data pipelines, ensuring smooth and reliable data flow for business analytics and reporting • Conducted data profiling and cleansing activities to identify and rectify inconsistencies, missing values, and outliers in datasets
Education

Business Analytics (Data Science)
Relevant Coursework: Predictive Analytics, Machine Learning for Analytics, Statistics for Data Science, Data and Programming for Analytics, Natural Language Processing, Business Data Management, Management Science for Analytics, The Art and Science of Applied Forecast Modeling
Chirag Madhukar's Contact Information
Phone
Find the Right Leads
Find Verified Contact Data
What LeadContact does well
Find verified emails, phone numbers, and decision-makers with 98% accuracy.
Find Leads
Find the right people by company, role, industry, location, and more.
925M+ professional profiles

Find Emails
Access verified email addresses for your target contacts.
657M+ emails

Find Phone Numbers
Get cross-validated phone data from multiple top sources.
239M+ phone numbers

More Accurate. Lower Cost.
Find contact data in 1 tool with 98% accuracy
LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.
Great conversations start with the right contact.
It’s time to find yours.



