Hetal Vaghela
Data Engineer @ CVS Health
About
Innovative Data Engineer with 4+ years of experience delivering scalable, cloud-native data solutions across healthcare and financial services. I specialize in transforming complex data ecosystems into reliable, high-performance platforms that drive measurable business outcomes. I bring deep expertise in ETL/ELT pipeline design, real-time streaming architectures, and modern data platforms across AWS, GCP, and Azure. My work focuses on building secure, compliant, and analytics-ready systems that improve decision-making, operational efficiency, and data quality. What I do best: Cloud & Data Platform Engineering: Led migrations from legacy systems to AWS, GCP, and Azure, achieving up to 40% efficiency gains and significant cost optimization. Real-Time & Big Data Processing: Designed and optimized pipelines using Apache Kafka, Spark, and Hadoop, reducing batch processing time by 30% and enabling near real-time insights. Healthcare Data Engineering: Currently at CVS Health, building HIPAA-compliant data pipelines supporting prescription analytics, inventory optimization, and operational reporting. Financial Analytics & Warehousing: Previously engineered scalable data warehouses and data marts, improving credit risk analysis accuracy by 35% and reducing query latency by 50%. Recent impact highlights: Accelerated delivery timelines by 20% through Agile execution and cross-functional collaboration Reduced deployment failures by 25% using CI/CD pipelines and DevOps best practices Improved data quality by 30% through automated validation, schema enforcement, and governance frameworks Built predictive ML models (74% accuracy) to support flood preparedness and risk forecasting initiatives Technical strengths: Python, SQL, Java | AWS, GCP, Azure | Spark, Hadoop, Snowflake, BigQuery | Kafka | Airflow, AWS Glue, DBT | CI/CD, Docker, Kubernetes | Tableau, Power BI, QuickSight | TensorFlow, Scikit-learn 🎓 M.S. in Information Technology, University of Massachusetts Boston (GPA: 3.78) I’m passionate about using data to solve real-world problems and thrive in collaborative environments where engineering excellence meets business impact.
United States
Boston
Computer Software
SKYLINE, Find a Hiring Manager, Data Loading, PHP, Data Services, Databases, Customer Relationship Management (CRM), Power BI, Microsoft Office, JavaScript, C#, IntelliJ IDEA, NetBeans, Agile & Waterfall Methodologies, Database Consulting, Statistical Data Analysis, Database Analysis, Software Development Life Cycle (SDLC), Agile Methodologies, Microsoft Excel
Experience

Data Engineer
United States
•Designed and built data pipelines and analytics workflows supporting healthcare reporting, inventory tracking, and operational decision-making in cloud-based environments. •Developed ETL pipelines using Python and SQL within the Hadoop ecosystem, automating data ingestion and transformation and improving processing efficiency by 40%. •Optimized Apache Spark batch pipelines processing warehouse and supply chain datasets, reducing batch processing time by 30% while maintaining data accuracy and reliability. •Implemented real-time streaming pipelines using Apache Kafka to process pharmacy inventory and order events, applying schema validation and governance checks to improve data quality by 30%. •Improved analytics performance by enabling serverless querying with Amazon Athena on S3, reducing query response times by 20%. •Contributed to the development of CI/CD pipelines and DevOps workflows for data applications, reducing deployment errors by 25% and improving release stability. •Collaborated with engineers, analysts, and product partners using Agile practices, helping accelerate development cycles by 20%. •Built interactive dashboards in Amazon QuickSight to visualize prescription trends and inventory metrics, improving operational visibility and decision-making efficiency by 25%. Tech Stack: Python, SQL, AWS (Glue, S3, Athena, QuickSight), Apache Spark, Apache Kafka, Hadoop, CI/CD, DevOps

Data Engineer
Ahmedabad, Gujarat, India
•Developed and maintained cloud-based data pipelines using GCP Dataflow, unifying transactional and structured data sources with production-grade testing and validation. •Integrated Hadoop and Apache Spark with relational databases to support financial data warehousing and OLAP workloads, reducing data retrieval latency by 50%. •Contributed to the migration and adoption of Snowflake as a cloud data warehouse, improving scalability and reducing query execution time by 45% for analytics workloads. •Built data marts and dimensional data models using Google BigQuery, improving reporting accuracy by 35% and enabling efficient financial analytics. •Implemented MapReduce and batch processing jobs to transform large-scale datasets, improving data accuracy and supporting enterprise reporting pipelines. •Applied machine learning models using TensorFlow and Scikit-learn to forecast financial market trends, improving prediction accuracy by 20% and supporting risk analysis use cases. •Supported containerized data workflows using Docker and Kubernetes, improving scalability and operational consistency of data processing services. •Developed financial dashboards using Google Data Studio, providing near real-time insights and improving strategic decision-making efficiency by 40%. Tech Stack: GCP, Python, SQL, Hadoop, Apache Spark, Snowflake, BigQuery, TensorFlow, Scikit-learn, Docker, Kubernetes
Education

Information Technology
Half Tuition Waiver & Graduate Assistantship – Awarded based on academic excellence and research contribution throughout the program. • Core Coursework: • Data Management Systems • Business Intelligence • Business Programming • Big Data Analytics • Project Management • Collaborated on data-driven research projects focusing on predictive analytics and cloud-based ETL optimization. • Actively mentored new international students as part of the International Student Organization, assisting with onboarding, visa workshops, and cultural adaptation. • Led initiatives to organize campus-wide events promoting cross-cultural exchange and community building. • Represented the university in multiple graduate research symposiums, showcasing work in big data processing and machine learning integration.

Information Technology
• Core Technical Coursework: • Programming Languages: Java, C, C++, Python • Web Technologies: HTML, CSS, JavaScript, C#, ASP.NET MVC • Database Management Systems (DBMS), SQL • Big Data Fundamentals and Ecosystem • Software Engineering & Object-Oriented Programming • Operating Systems, Data Structures, and Algorithms • Completed multiple hands-on academic projects in full-stack development and database design, applying industry best practices. • Participated in coding competitions and internal hackathons, achieving recognition for logical problem-solving and efficient code structure. • Served as a peer tutor for foundational programming courses, helping juniors strengthen their coding logic and debugging skills. • Organized tech workshops and seminars as part of the university’s IT Club, focusing on emerging technologies like Big Data and Web APIs.
Hetal Vaghela's Contact Information
Phone
Find the Right Leads
Find Verified Contact Data
What LeadContact does well
Find verified emails, phone numbers, and decision-makers with 98% accuracy.
Find Leads
Find the right people by company, role, industry, location, and more.
925M+ professional profiles

Find Emails
Access verified email addresses for your target contacts.
657M+ emails

Find Phone Numbers
Get cross-validated phone data from multiple top sources.
239M+ phone numbers

More Accurate. Lower Cost.
Find contact data in 1 tool with 98% accuracy
LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.
Great conversations start with the right contact.
It’s time to find yours.

