Rakesh P

Rakesh P

Senior Data Engineer – Snowflake · dbt · Azure · HIPAA @ Molina Healthcare

About

As a Senior Data Engineer at Molina Healthcare, I contribute to the design and optimization of scalable data pipelines and Medallion Architecture (bronze/silver/gold) on Azure and Snowflake. My work includes processing over a million healthcare records daily for clinical analytics while ensuring compliance with HIPAA standards. I specialize in implementing Infrastructure as Code (IaC) using Terraform to enhance operational efficiency and reduce provisioning time significantly. Currently pursuing a Master’s in Computer Science at the University of Central Missouri, I bring a strong foundation in Python, SQL, and XML, complemented by certifications in Azure and AWS Data Engineering. My focus is to drive innovation and data-driven decision-making by delivering robust and efficient data solutions.

Country

-

City

United States

Industry

Hospital & Health Care

Skill

Python (Programming Language), XML, SQL Server Management Studio, Tableau, FastAPI, REST APIs, Gitlab, GitHub, github actions, Data Validation, Metadata Management, Data observability, Data Lineage, data contracts, Snowflake schema, Star Schema, Azure Synapse, Google BigQuery, Amazon Elastic MapReduce (EMR), Amazon Redshift

Experience

Molina Healthcare

Senior Data Engineer – Snowflake · dbt · Azure · HIPAA

Molina Healthcare

LinkedIn
2024-3 - Present · 2 yrs 7 mos

Sacramento, CA

• Architected Medallion Architecture (bronze/silver/gold) on Azure and Snowflake, processing 1M+ healthcare records daily for clinical analytics using Delta Lake and Apache Iceberg. • Designed and maintained scalable ETL/ELT pipelines using Snowflake, dbt Cloud, and Azure Data Factory, enforcing Data Contracts and schema evolution policies to prevent breaking changes across downstream dependencies. • Implemented Infrastructure as Code (IaC) using Terraform to automate provisioning of Snowflake objects, cloud resources, and RBAC roles reducing provisioning time by 60%. • Optimized Snowflake performance through advanced clustering keys, Z-ordering and automatic clustering, achieving 45% reduction in query latency for complex analytical queries. • Refactored inefficient SQL queries and redesigned data models, achieving 35% reduction in overall pipeline runtime and improved SLAs for downstream consumers. • Established comprehensive data quality framework using dbt and Great Expectations with automated validation, anomaly detection, and data profiling reducing critical data discrepancies by 40%. • Implemented HIPAA-compliant security controls including Row-Level Security (RLS), dynamic data masking, ColumnLevel Security, and encryption at rest and in transit for 2M+ sensitive patient records. • Built real-time data observability dashboards tracking pipeline health, data freshness, volume trends and quality metrics for proactive incident response using Power BI. • Integrated LLM-based pipelines for automated extraction and summarization of unstructured clinical notes, enhancing data availability for advanced analytics and research. Environment: Azure, Snowflake, dbt Cloud, Azure Data Factory, Terraform, Python, PySpark, Apache Airflow, Great Expectations, HIPAA, Delta Lake, Apache Iceberg, Parquet, Git, GitHub Actions, Power BI

T-Mobile

Data Engineer – AWS · GCP · Kafka · PySpark (T-Mobile)

T-Mobile

LinkedIn
2023-9 - 2024-2 · 6 mos

Overland Park, KS

• Designed and delivered scalable data pipelines processing 5TB+ daily across AWS and GCP, supporting critical analytics and operational reporting for 4 engineering teams across multi-cloud environments. • Built distributed data transformations using Python, PySpark, and SQL on Snowflake and BigQuery, handling complexjoins, schema evolution and large-scale aggregations for business intelligence. • Automated cloud infrastructure provisioning using Terraform, managing S3 buckets, IAM roles, Glue jobs, VPC configurations, and secrets management reducing manual deployment effort by 50%. • Orchestrated batch, micro-batch, and Spark Streaming pipelines using Apache Airflow and Apache Kafka with EventDriven Architecture, designing DAGs with Change Data Capture for late-arriving data handling. • Optimized SQL query performance through refactoring, join restructuring, and workload tuning achieving 35% reduction in pipeline runtime and measurable cloud cost savings. • Established Data Lineage tracking and governance standards across multi-cloud pipelines using Unity Catalog, improving end-to-end traceability from source ingestion to final reporting. • Implemented data quality validation frameworks using Great Expectations to detect anomalies, duplicates, and inconsistencies in high-volume Spark Streaming and batch data. • Collaborated with analytics teams, data scientists and business stakeholders to understand requirements and deliver reliable, production-scale data systems. • Configured Snowflake and BigQuery warehouses for cost-performance balance, implementing auto-scaling, clustering, partitioning, and Spark Streaming optimizations. Environment: AWS (S3, Glue, Kinesis, IAM, VPC), GCP (BigQuery, Cloud Storage), Snowflake, Unity Catalog, Terraform, Apache Airflow, Python, PySpark, SQL, Apache Kafka, Spark Streaming, Git, GitHub Actions, Tableau

Capgemini

Data Engineer

Capgemini

LinkedIn
2021-12 - 2023-7 · 1 yr 8 mos

Kolkata

• Designed and delivered enterprise-scale financial data pipelines and data warehouses using Snowflake, dbt, and AWS for a global investment banking client, supporting risk analytics, regulatory reporting (Basel III, CCAR), and trade reconciliation across equities and fixed income portfolios. • Built and maintained ETL/ELT pipelines ingesting 5TB+ of daily trade, position, and market data from upstream systems (Bloomberg, Reuters, internal OMS) into a centralized Snowflake data warehouse for P&L and risk attribution reporting. • Applied Change Data Capture patterns using Databricks and EMR to enable real-time ingestion of transaction ledger updates, reducing end-of-day reconciliation latency by 35%. • Implemented Infrastructure as Code (IaC) with Terraform across development, staging, and production environments, standardizing cloud resource provisioning and reducing deployment failures by 45%. • Refactored complex legacy SQL queries and stored procedures for financial reporting models optimizing join logic, eliminating redundant calculations, and achieving 40% reduction in report generation time for daily risk dashboards. • Enforced data governance controls including Column-Level Security for PII and sensitive financial data, Data Lineage documentation from source trades to aggregated risk positions, and RBAC-based access policies aligned with SOX compliance requirements. • Designed Star Schema and Snowflake Schema dimensional models for financial reporting covering trade lifecycle, counterparty exposure, and portfolio performance improving query performance by 25%. Environment: AWS (S3, Redshift, Glue, IAM), Snowflake, dbt, Terraform, Python, SQL, Apache Airflow, Databricks, Great Expectations, Git, Agile, Jira, Confluence, Tableau

DXC Technology

Software Engineer – Data & Integration

DXC Technology

LinkedIn
2020-8 - 2021-11 · 1 yr 4 mos

Bengaluru

• Developed and maintained data integration solutions for a large financial services client, automating ingestion of trade confirmations, account statements and transaction records from multiple upstream banking systems into a centralized SQL data warehouse. • Built reusable Python modules and REST API integrations to extract structured and semi-structured financial data (CSV, JSON, XML) from third-party vendors and core banking platforms, implementing error handling, pagination and rate limiting for reliable daily loads. • Wrote and optimized complex SQL queries, stored procedures, indexes and views to support financial reporting and reconciliation workflows achieving 25% improvement in query execution times and reducing report generation lag. • Developed data validation and cleansing routines to enforce business rules on incoming financial records flagging duplicate transactions, null account identifiers and type mismatches before loading into downstream reporting tables. • Automated ETL/ELT job scheduling and monitoring using Bash scripting and Cron Jobs on Linux servers, implementing structured logging and alerting to enable proactive failure detection and reduce manual intervention. • Supported software development lifecycle (SDLC) activities including unit testing, code reviews, and deployment scripting, contributing to on-time delivery of 3 data integration modules in an Agile team environment. • Documented data flow diagrams, transformation logic and API integration specifications, building foundational data engineering knowledge that enabled a transition into full-scale financial data pipeline. Environment: Python, Pandas, SQL, MySQL, PostgreSQL, REST APIs, CSV, JSON, XML, Git, Jira, Agile, Linux, Bash, Cron Jobs

Education

University of Central Missouri

University of Central Missouri

LinkedIn

Computer Science

2023-8 - 2025-5 · 1 yr 10 mos
Acharya Nagarjuna University

Acharya Nagarjuna University

LinkedIn

B.Sc. Of Computer Science

2017-6 - 2021-6 · 4 yrs 1 mo

Rakesh P's Contact Information

Email

******@***.com

Phone

(**) *** ****

Find the Right Leads
Find Verified Contact Data

Try with: Jensen Huang @ nvidia.com Click to autofill
LeadContact awards, five-star ratings, and GDPR compliance badges

What LeadContact does well

Find verified emails, phone numbers, and decision-makers with 98% accuracy.

Find Leads

Find Leads

Find the right people by company, role, industry, location, and more.

925M+ professional profiles

Find Leads
Find Emails

Find Emails

Access verified email addresses for your target contacts.

657M+ emails

Find Emails
Find Phone Numbers

Find Phone Numbers

Get cross-validated phone data from multiple top sources.

239M+ phone numbers

Find Phone Numbers

More Accurate. Lower Cost.

Find contact data in 1 tool with 98% accuracy

LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.

LeadContact Logo
Competitor Tools

All these = $289 per month

Great conversations start with the right contact.

It’s time to find yours.