Sindhu P.

Sindhu P.

Azure Data Engineer @ Bell

About

I am an experienced Data Engineer with extensive expertise in building and optimizing data pipelines, ETL processes, and managing data across platforms like Databricks, AWS, Azure, and Snowflake. I am an expert in SQL, Python, and PySpark, with a strong background in cloud data management, machine learning implementations, and big data technologies. Proficient in automating data processes, managing cloud storage, and integrating diverse data storage and processing technologies. Committed to adopting the best practices in quality assurance, and continuous integration/deployment. Adept at working within Agile frameworks and continuously exploring emerging technologies to enhance data security and processing efficiency.

Country

-

City

Canada

Industry

Information Technology & Services

Skill

Azure Data Factory, PySpark, Apache Kafka, Microsoft Azure Machine Learning, Azure Databricks, Snow flake, Azure SQL, Extract, Transform, Load (ETL), Big Data, Data Warehousing, Microsoft Azure, Data Science, Google Cloud Platform (GCP), AWS, Data Analytics Internship Program, SQL, Data Visualization, Pandas, NumPy, Matplotlib

Experience

Bell

Azure Data Engineer

Bell

LinkedIn
2025-1 - Present · 1 yr 9 mos

Ontario, Canada

• Developed Databricks Notebooks to apply transformations using Spark SQL, ensuring seamless data flow and integration across systems, databases, and applications. • utilized Apache Hudi to enable incremental data processing and low-latency data ingestion, significantly improving data access times and storage efficiency. • Constructed and maintained processes for data transformation, data structures, metadata, dependency, and workload management, enhancing data integrity and availability • Provided hands-on experience in data refinement and performance tuning using PySpark, Scala, and SQL within Databricks, optimizing data processing and analytics tasks • Worked with distributed computing frameworks like Apache Spark to process and analyze large-scale datasets, demonstrating expertise in big data technologies and methodologies • Implemented Snowflake features including Snowpipe, UDFs, zero copy clones, time travel, micro-partitions, stored procedures, data import/export, and external tables for efficient cloud data management • Automated data movement in and out of Snowflake using SnowPipe, Streams, and Tasks utilities, ensuring efficient data pipeline constructions and data availability • Executed test plans to identify and analyze defects, integrated new code seamlessly, and identified performance bottlenecks, ensuring high-quality software delivery related Azure Data Factory (ADF) Pipelines and managed custom Azure development, showcasing proficiency in cloud data orchestration and troubleshooting • Integrated Azure services like Azure Databricks and Azure Synapse Analytics for enhanced data processing and analytics capabilities • Managed and configured Azure Blob Storage and Azure Data Lake Storage for data storage solutions • Utilized Snowflake's Data Sharing capabilities to securely share real-time, read-only data with external partners

Vosyn

Azure Data Engineer

Vosyn

LinkedIn
2024-1 - 2024-12 · 1 yr

Canada

• Designed, built, and automated ETL processes utilizing AWS Redshift, S3, alongside AWS Lambda for serverless data processing tasks, ensuring efficient data flow and optimal storage solutions • Leveraged Amazon Athena for ad-hoc query executions directly on data stored in S3, enhancing data accessibility and analytical capabilities • Employed AWS Glue for robust ETL processing, integrating additional tools like Apache Airflow for workflow orchestration, automating and scheduling complex data pipelines • Conducted data profiling and synthesized findings using advanced analytics tools like Amazon QuickSight and Google Data Studio, ensuring comprehensive data quality assessments • Utilized advanced features of Snowflake for data warehousing, implementing data sharing and cloning capabilities to enhance data availability and disaster recovery strategies • Applied sophisticated data modeling techniques using tools like Erwin or dbt (data build tool) for managing data transformations and version control in the data warehouse • Built and managed scalable data pipelines using Azure Databricks in conjunction with PySpark, integrating Azure Data Lake Storage for extensive data volume management, ensuring robust data processing capabilities • Leveraged AWS EMR for large-scale data processing in a highly dynamic environment, building ETL pipelines with Airflow to ensure the smooth processing of financial data. • Utilized AWS Athena for querying high-volume data stored in S3, supporting decision-making processes within a fast-paced, investment-focused environment. • Worked with AWS Glue and Redshift, enabling efficient data transformations for investment and capital market data analytics. Implemented Apache Airflow for workflow orchestration, automating complex data transformation processes and ensuring timely delivery of analytical insights.

Infosys.Ltd

Data Engineer

Infosys.Ltd

2020-11 - 2023-12 · 3 yrs 2 mos

India

• Developed Hive and Bash scripts for validating and transforming source data, automating data loading into HDFS and Hive for preprocessing using One Automation, demonstrating proficiency in scripting and automation • Gathered data from Data warehouses in Teradata and Snowflake, showcasing skills in data extraction and integration from diverse data warehousing technologies • Developed Spark/Scala and PYTHON scripts for a regular expression project in the Hadoop/Hive environment, highlighting expertise in big data processing and programming languages • Designed and implemented an ETL framework to load data from multiple sources into Hive and from Hive into Teradata, showing advanced skills in ETL processes and data warehousing solutions • Generated reports using Tableau, emphasizing capabilities in data visualization and reporting for executive decision-making • Utilized SQOOP, ETL, and Hadoop Filesystem APIs for implementing data ingestion pipelines, showcasing knowledge in data ingestion tools and methodologies • Worked on batch data of different granularity ranging from hourly to monthly, illustrating the ability to manage and process data at varying intervals • Had hands-on experience in Hadoop administration and support activities, including installations and configuration of Apache Big Data Tools and Hadoop clusters using Cloudera Manager, indicating a strong foundation in Hadoop ecosystem management • Handled Hadoop cluster installations in various environments such as Unix, Linux, and Windows, showing versatility in operating systems • Assisted in upgrading, configuration, and maintenance of various Hadoop infrastructures like Ambari, PIG, and Hive, demonstrating maintenance skills in big data platforms

Sindhu P.'s Contact Information

Email

******@***.com

Phone

(**) *** ****

Find the Right Leads
Find Verified Contact Data

Try with: Jensen Huang @ nvidia.com Click to autofill
LeadContact awards, five-star ratings, and GDPR compliance badges

What LeadContact does well

Find verified emails, phone numbers, and decision-makers with 98% accuracy.

Find Leads

Find Leads

Find the right people by company, role, industry, location, and more.

925M+ professional profiles

Find Leads
Find Emails

Find Emails

Access verified email addresses for your target contacts.

657M+ emails

Find Emails
Find Phone Numbers

Find Phone Numbers

Get cross-validated phone data from multiple top sources.

239M+ phone numbers

Find Phone Numbers

More Accurate. Lower Cost.

Find contact data in 1 tool with 98% accuracy

LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.

LeadContact Logo
Competitor Tools

All these = $289 per month

Great conversations start with the right contact.

It’s time to find yours.