Srinivasa Rao Boyapati
Senior Data Engineer @ Costco Wholesale
About
Data Engineer with 10+ years of experience in building scalable data platforms across cloud (Azure, AWS, GCP), big data, and real-time streaming ecosystems. Specialized in developing robust ETL pipelines, data lakes, and advanced analytics solutions using tools like Spark, Hadoop, Airflow, and Snowflake. Proficient in Azure Data Factory, Azure Synapse, Data Lake Analytics, Databricks, AWS Glue, Redshift, Athena, S3, DynamoDB, and GCP Compute & Storage. Delivered enterprise data migration and integration projects with optimized performance and high availability. Extensive hands-on expertise with Hadoop ecosystem (HDFS, Hive, HBase, Kafka, Flume, Sqoop, Pig), real-time frameworks (Kafka, Spark Streaming, Storm), and data warehousing (Snowflake, Redshift, Azure Synapse, Teradata). Experience includes setting up role-based access control, encryption, compliance protocols, and disaster recovery using Ansible. Built interactive dashboards using Power BI and Tableau, developed REST APIs, and optimized data models across RDBMS and NoSQL databases (MySQL, Oracle, SQL Server, Cassandra, MongoDB). Strong scripting knowledge in Python, Bash, and SQL. Known for automating workflows, improving data reliability, and delivering actionable insights. Adept at Agile/Scrum processes, stakeholder collaboration, and translating business needs into scalable data solutions.
United States
Dallas
Information Technology & Services
Amazon EC2, Amazon S3, Google Cloud Platform (GCP), Oracle Database, Identity and Access Management (IAM), Google BigQuery, Teradata, Tableau, Microsoft Excel, HDFS, Pandas (Software), AWS Glue, Cassandra, Hadoop, Azure Databricks, Azure Cosmos DB, PySpark, Scala, SQL, Microsoft Power BI
Experience

Senior Data Engineer
Issaquah, Washington, United States
Led real-time data engineering initiatives using Spark Streaming to process data from S3 and Kafka, storing results in Delta Lake and HDFS. Built streaming applications in AWS Databricks and Apache Flink, delivering low-latency insights. Designed and maintained scalable Aerospike clusters for high-performance data access and replication. Developed ETL workflows using AWS Glue, Lambda, Step Functions, and PySpark to automate data ingestion, transformation, and quality checks. Created and orchestrated robust data pipelines using Apache Airflow and AWS Data Pipeline. Implemented custom DAGs, operators, and monitoring for reliable execution. Managed infrastructure with Ansible, automating server provisioning and connecting deployments to CI/CD pipelines. Built scalable search systems using Apache Solr, enhancing query speed and indexing efficiency. Executed data migrations from SQL Server and MySQL to Redshift, leveraging EMR and Snowflake for batch/stream processing from S3. Developed Snowpipe and Matillion pipelines for secure transfers. Used CloudWatch and custom alerts to monitor system performance and implemented failover mechanisms for high availability. Set up Kerberos authentication and Dockerized environments for secure, consistent deployments. Wrote Python and Scala scripts to handle data from APIs, S3, and enterprise DBs (Oracle, DB2, Teradata). Optimized data models in Redshift and Snowflake, supporting compliance and analytics needs. Built Lambda functions and API Gateway endpoints for event-driven data processing. Led the design of BI systems and streamlined infrastructure using Jenkins, CodeCommit, and containerization strategies.

Cloud Data Engineer
New York, New York, United States
Led end-to-end Azure Data Engineering projects across banking, finance, utilities, and tech domains. Built and managed scalable ADF pipelines to move data across Azure SQL, Blob Storage, and Data Warehouse systems. Designed Lambda architectures using Azure Data Factory, Data Lake, HDInsight, Azure SQL, and Power BI to support batch and real-time analytics. Built PySpark and Scala-based Spark jobs in Databricks for data processing, table transformations, and Hive integration. Developed automated workflows using Airflow with advanced scheduling, retry logic, and failure handling. Implemented U-SQL pipelines for loading various file formats into Azure SQL Data Warehouse with quality checks. Created Power BI dashboards and SSRS reports, integrating Azure Analysis Services for enhanced performance and row-level security. Used ADF to integrate data from sources like Sybase, Informix, MongoDB, Cosmos DB, and Blob Storage with robust monitoring and alerting systems. Engineered data ingestion workflows using Sqoop, integrated Spark-Kafka pipelines for real-time ETL, and managed Flink jobs for high-throughput streaming use cases. Built and optimized Solr indexing with secure configurations and custom connectors. Enhanced performance and reliability with partitioned Hive tables, Snappy-compressed Parquet files, and efficient HBase storage design. Managed Hadoop and HBase environments, resolving cluster issues proactively and improving throughput. Automated deployments using Ansible, including RBAC, patching, and server provisioning. Integrated CI/CD workflows with Git, Bitbucket, and Jira for version control, issue tracking, and team collaboration. Created reusable Spark libraries, implemented secure access protocols, and followed best practices in enterprise-grade data solutions.

Data Engineer
New Brunswick, New Jersey, United States
Experienced in designing scalable and secure data engineering solutions across cloud platforms (Azure, AWS, GCP) and big data ecosystems. Developed robust data pipelines using Spark, PySpark, Hive, and Python for both batch and real-time ETL workflows. Built Snowflake data models and implemented data governance frameworks to ensure data quality, compliance, and secure access. Created automation scripts in PowerShell and PySpark for efficient data movement between on-premise and cloud systems (HDFS, AWS S3, Blob). Managed Glue-based ETL processes and built optimized external tables in AWS Redshift and Athena. Processed streaming data using Kafka, Spark Streaming, and stored outputs in NoSQL systems such as HBase and Cassandra. Built real-time applications integrating DB2, Spark, and HBase. Hands-on experience with Hadoop components (HDFS, Sqoop, Hive, Pig, Flume, Oozie, NiFi, YARN, MapReduce), and built distributed applications across Hadoop clusters with advanced partitioning, compression, and monitoring techniques. Developed and monitored MLflow pipelines in Databricks for model tracking and versioning. Built OLAP cubes and dimensional models in Azure SQL DW, developed dashboards and reporting tools in Power BI, SharePoint, and SSRS, and implemented RBAC security for access control. Used Python (Pandas, NumPy) for feature engineering, data transformation, and complex file processing (JSON, CSV, XML). Implemented CloudWatch dashboards, set up SSH tunnels for Google DataProc, and managed resource optimization for Snowflake and EMR clusters. Led cross-functional collaboration with analysts and business users, created system documentation, and delivered actionable insights to stakeholders across marketing and operations.

Data Engineer
Chicago, Illinois, United States
Developed and maintained data solutions using Google Cloud Platform (BigQuery, GCS, Cloud Functions, Dataflow, Dataproc, and Cloud Monitoring). Automated service deployment and management with Cloud Shell SDK scripts. Collaborated with users to gather requirements and created detailed ETL design documentation. Built and optimized data pipelines leveraging Dataflow, Apache Beam, and Spark to process large datasets from diverse sources. Created Python Cloud Functions for automated CSV loading into BigQuery and implemented REST API integrations with robust error handling and authentication. Managed and optimized GCP infrastructure for cost efficiency and scalability. Integrated GCP services with BI tools, analytics systems, and external APIs to streamline data workflows. Designed security protocols with IAM and configured Security Command Center for threat monitoring and compliance. Handled petabyte-scale data storage using HDFS; improved retrieval speeds by 40% through data partitioning and optimized storage. Developed Hadoop workflows with Hive and Pig to deliver actionable insights. Created monitoring dashboards and automated alerting with Google Cloud Monitoring. Built scalable automated workflows to enhance data processing efficiency and collaborated closely with data science teams to translate business needs into Hadoop and cloud solutions. Designed ETL processes using Informatica PowerCenter for SQL Server and Oracle migrations, optimizing transformation logic. Established Master Data Management frameworks to ensure data quality and consistency. Developed custom Tableau dashboards with interactive visualizations and automated report scheduling. Enhanced database performance by optimizing stored procedures, functions, and triggers using SQL and PL/SQL best practices, including parallel processing and query hints.

ETL Developer
Hyderabad, Telangana, India
Led optimization of ETL workflows and database processes, reducing processing time and resource usage by identifying bottlenecks and improving data pipelines. Designed and implemented AWS infrastructure using EC2, S3, and DynamoDB, configuring auto-scaling, load balancers, and backup systems for high availability and fault tolerance. Coordinated project timelines and task distribution, prioritizing workflows to maximize team productivity while ensuring code quality and documentation. Collaborated with data providers to resolve quality issues and restructure source data for better integration. Implemented AWS monitoring with CloudWatch metrics, alerts, and dashboards, setting up automated responses to maintain system health and optimize performance. Supported ETL development and solved complex data integration challenges in an agile environment. Developed performance testing and capacity planning for ETL processes and supported production troubleshooting. Used Business Objects Data Services (BODI) to create and document mappings, ensuring data consistency. Built ETL solutions ingesting CSV, JSON, and XML into SQL databases, with reusable components, error handling, and logging. Applied complex data transformations using source qualifier, expression, filter, aggregator, lookup, stored procedures, and other components. Reduced AWS costs by using Reserved and Spot Instances, monitoring usage to control spend without sacrificing performance. Automated Informatica session loads with shell scripts, including error handling and notifications. Enhanced Informatica server performance by tuning configurations based on metrics analysis.
Srinivasa Rao Boyapati's Contact Information
Phone
Find the Right Leads
Find Verified Contact Data
What LeadContact does well
Find verified emails, phone numbers, and decision-makers with 98% accuracy.
Find Leads
Find the right people by company, role, industry, location, and more.
925M+ professional profiles

Find Emails
Access verified email addresses for your target contacts.
657M+ emails

Find Phone Numbers
Get cross-validated phone data from multiple top sources.
239M+ phone numbers

More Accurate. Lower Cost.
Find contact data in 1 tool with 98% accuracy
LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.
Great conversations start with the right contact.
It’s time to find yours.

