Siddhant Madan
AI Data Engineer @ Omdena
About
MS in Business Analytics @ UC | 5+ yrs in Data Engineering, BI & Analytics | Open to Data Science, DE, & Analytics Roles I build scalable data solutions that turn raw data into actionable insights. My experience spans healthcare, CRM, and retail domains, where I’ve led end-to-end projects—from data pipelines to machine learning models. Tech Stack: 🔹 Python, SQL, R, Java 🔹 Spark, PySpark, Delta Lake, Databricks, Kafka, Airflow 🔹 MySQL, Teradata, AWS Redshift 🔹 Power BI, Tableau, Logi Analytics 🔹 Azure Data Factory, Synapse Analytics 🔹 GitHub, Bitbucket, Salesforce 🔹 ML: Supervised Learning, NLP, Time Series, Feature Engineering
United States
San Francisco Bay Area
Information Technology & Services
Amazon EC2, Streamlit, GitHub, OpenClaw, FastAPI, Postman API, Edtech, Banking, ICD-10, Analytical Skills, Electronic Health Records (EHR), U.S. Health Insurance Portability and Accountability Act (HIPAA), supply chain data, SAP Business Warehouse (SAP BW), POS Data Analysis, Azure Databricks, Delta Lake, Microsoft SQL Server, YAML, pandas
Experience

Analytics Engineer
Cincinnati, OH
• Engineered a Python ETL pipeline integrating Salesforce and QuickBooks APIs to unify data, enabling Power BI tracking of a $50M+ grant portfolio. • Designed DAX-driven Power BI dashboards with drill-through analytics, providing stakeholders real-time visibility into funding trends and proposal submissions. • Optimized AI Document Processing by developing a Python automation script to extract embedded URLs from Word tables and headers, ensuring 100% context retrieval for playbook validation. • Implemented AI auditing workflows using Ironclad (Jurist AI) to automate contract compliance checks, replacing manual reviews and identifying standard-deviation gaps.

Graduate Assistant (Data Science)
Cincinnati, OH
1. Interactive Audit Application Client: Cintas Spearheaded a Python/Voila dashboard processing large SKU datasets with 15+ dynamic rules, reducing anomaly detection time by 50% and minimizing supply chain compliance risk. 2. Legal Document Intelligence System Client: HBH Holdings Engineered a resource-efficient RAG system (Haystack/Mistral-7B) achieving 75% accuracy on 250 legal documents, enabling high-performance local deployment and eliminating cloud costs.

Senior Data Engineer
Gurugram
Client : Amplify Health (MLOps & Feature Engineering) • Implemented a Universal Data Model (UDM) and tailored Feature Store on Databricks for multi-region claims data, standardizing global analytics and reducing team onboarding time by 40%. • Developed post-processing scripts and feature store functionalities, including data normalization and chronic condition classification, to enhance data relevance for patient risk stratification. • Engineered 50+ reusable ML features using Object-Oriented Python (OOP) and PySpark, delivering specific features for Financial (Regression) and Clinical (Classification) models. • Optimized PySpark transformation pipelines by analyzing execution plans (Spark UI) and implementing broadcast joins/caching, achieving a 30% reduction in system latency. • Formulated a comprehensive data logging framework to ensure traceability and reproducibility for model training, evaluation, and drift debugging. • Resolved complex data quality and model performance issues by collaborating with Principal Data Scientists to build a structured root-cause analysis framework. Client : Jefferson Health (Cloud Migration & Architecture) • Engineered Silver-layer pipelines (Synapse/PySpark) for a strategic migration of 10TB+ Epic Caboodle data to Azure, designing bridge tables and collaborating with clinical coders to validate logic, improving consistency by 25%. • Built a HIPAA-aligned reconciliation framework to enforce schema accuracy, ensuring compliant EHR processing and mitigating regulatory risk across the migration pipeline. • Partnered with the engineering team to refactor legacy SQL logging processes into optimized PySpark DataFrames, resulting in a 50% reduction in notebook execution time. • Supported team delivery by providing technical guidance to junior members on Spark best practices, earning Employee of the Month honors for excellence in team-based execution.

Business Intelligence Analyst
Bengaluru
• Analyzed complex CRM datasets to drive customer segmentation, profitability, and personalization strategies, delivering actionable insights through custom web reports and interactive dashboards. • Engineered automated ETL pipelines using Advanced SQL and Python to process millions of records across a multi-tenant architecture, reducing manual reporting overhead by 40%. • Facilitated business requirement discussions and conducted technical feasibility checks to translate stakeholder needs into optimized data models and intuitive user interfaces. • Produced 50+ end-to-end BI solutions for high-scale EdTech and BFSI clients, enabling churn prediction models that improved customer retention by 30%. • Implemented Row-Level Security (RLS) and hybrid data retrieval strategies using Elasticsearch and Redshift to ensure data sovereignty and improve query performance for low-latency reporting. • Collaborated directly with strategic clients to define key metrics for Sales and Marketing funnels, visualizing the end-to-end "Lead Journey" to identify conversion bottlenecks.

Data Engineer / Analyst
Kolkata
Client: The Co-operative Group Project: Retail Data Platform • Orchestrated the strategic migration of a 5TB+ SAP Retail warehouse from Teradata to Azure, leveraging Azure DevOps for CI/CD and Azure Synapse Analytics to ensure data integrity. • Automated the generation of complex SQL scripts using Shell scripting (BASH), which eliminated 2 hours of manual coding effort per day and reduced human error. • Enhanced and optimized existing Linux-based ETL workflows, significantly improving the reliability and throughput of daily retail data ingestion. • Engineered custom Python parsers to digest complex XML, CSV, and SAP flat files, enabling the structured loading of nested data for a high-priority retail business transformation project. • Executed advanced data modeling by developing and maintaining SQL scripts for Type-1 and Type-2 Slowly Changing Dimensions (SCD), ensuring historical accuracy across critical inventory tables. • Optimized warehouse storage efficiency by implementing advanced data compression techniques, reducing the infrastructure footprint while maintaining query performance. • Resolved critical store and customer data discrepancies through deep-dive root cause analysis, utilizing Power BI to visualize data quality trends and report findings. • Analyzed granular basket-level and stock data using SQL to identify supply chain patterns, directly informing optimized inventory replenishment strategies.
Education

Business Analytics
Relevant Coursework : Data Visualization | Data Wrangling | Simulation Modeling and Methods | Probability Models | Statistical Methods | Applied Linear Regression | Big Data | Governance of AI/ML Systems | Marketing Foundations
Siddhant Madan's Contact Information
Phone
Find the Right Leads
Find Verified Contact Data
What LeadContact does well
Find verified emails, phone numbers, and decision-makers with 98% accuracy.
Find Leads
Find the right people by company, role, industry, location, and more.
925M+ professional profiles

Find Emails
Access verified email addresses for your target contacts.
657M+ emails

Find Phone Numbers
Get cross-validated phone data from multiple top sources.
239M+ phone numbers

More Accurate. Lower Cost.
Find contact data in 1 tool with 98% accuracy
LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.
Great conversations start with the right contact.
It’s time to find yours.




