
Surya Singh Kafle
Data Engineer / Data Analyst @ International Paper
About
I am a Data Engineer with experience building scalable data pipelines, ETL workflows, and cloud-based data architectures that transform raw data into reliable datasets for analytics and decision-making.My work focuses on designing and optimizing data pipelines using Python, PySpark, Spark, and SQL across modern cloud environments such as AWS and Azure. I have experience developing batch and streaming data workflows, implementing data ingestion frameworks, and building data warehouse solutions using platforms like Snowflake and Redshift. I enjoy solving complex data problems, whether that involves processing large-scale datasets with distributed systems, designing efficient data models, or improving pipeline performance and reliability.Previously, I worked as a Graduate Research Assistant at the University of Memphis where I developed scalable ETL pipelines using PySpark and Spark SQL, implemented data ingestion workflows with Azure Data Factory, and built streaming pipelines using Kafka and Spark Streaming to support near real-time analytics. I also worked as a Data Engineer where I built big data pipelines on AWS using Spark, Hive, and Airflow, developed data ingestion frameworks, and designed warehouse-ready datasets for analytics and reporting.I completed my Master’s in Computer Science from the University of Memphis in December 2025. I am currently seeking opportunities as a Data Engineer where I can contribute to building scalable data platforms, distributed data systems, and reliable analytics pipelines.Technical areas include Python, PySpark, Spark, SQL, Kafka, Airflow, Snowflake, AWS, Azure, data warehousing, ETL/ELT pipelines, and big data systems.
United States
Memphis
Computer Software
Data Mining, Data Modeling, Data Warehousing, Data Architects, Amazon Web Services (AWS), Data Architecture, A/B Testing, Data Analysis, Data Science, Data Intelligence, pandas, NumPy, PostgreSQL, Shell Scripting, Git, Scala, Database Administration, Database Design, Computer Engineering, JavaScript
Experience

Data Engineer / Data Analyst
Memphis, TN
• Designed and implemented scalable ETL and ELT pipelines using PySpark, Spark SQL, and Azure Databricks, processing 50M+ records daily across enterprise datasets for analytics and machine learning workloads. • Built automated data ingestion pipelines using Azure Data Factory (ADF) integrating 10+ structured and semi-structured data sources including Azure SQL Database, Blob Storage, and Azure Data Lake, reducing ingestion latency by 40%. • Implemented streaming data pipelines using Apache Kafka and Spark Streaming to support near real-time analytics and operational monitoring. • Designed and maintained Snowflake data warehouse schemas, creating optimized SQL transformations and dimensional data models for enterprise reporting and analytics. • Migrated legacy ETL processes to cloud-based data pipelines using Azure Databricks and Snowflake, improving pipeline reliability by 35% and reducing average job execution time by 50%. • Built feature engineering pipelines using PySpark and SQL to prepare large datasets used in predictive modeling and analytical experiments. • Supported development of machine learning models using TensorFlow and PyTorch by building training datasets, feature pipelines, and automated data preprocessing workflows. • Engineered multiple predictive features from behavioral and operational datasets to improve model performance and analytical insights. • Developed data pipelines to support A/B testing and experimentation frameworks, enabling teams to evaluate product and operational initiatives using data-driven approaches. • Extracted and processed large datasets using optimized SQL queries across distributed systems handling 100M+ records, achieving up to 60% improvement in query execution time. Technologies Used: Azure Data Factory, Azure Databricks, Azure Data Lake, PySpark, Spark, Snowflake, SQL, Kafka, Terraform, Azure DevOps, Power BI, Dash, TensorFlow, PyTorch, Apache Airflow, dbt, Great Expectations, OpenTelemetry

Data Engineer / Data Analyst
New York City Metropolitan Area
• Developed machine learning models and statistical analyses to evaluate financial and operational datasets and support business forecasting. • Built automated ETL pipelines using Python and SQL to collect, transform, and process data from multiple operational systems. • Reduced manual data processing effort by 60% by automating data preparation workflows across 5+ operational systems, improving data consistency and reliability. • Performed exploratory data analysis and feature engineering to identify key drivers of business performance. • Designed and implemented analytical data models and database schemas optimized for reporting and business intelligence. • Created visualizations and dashboards to communicate analytical insights and operational trends to stakeholders. • Wrote optimized SQL queries for large relational datasets of 10M+ rows, improving data retrieval performance by up to 50% and reducing analytical query runtime. • Collaborated with engineering and business teams to deliver data-driven insights and analytical solutions for operational decision-making. • Assisted in building data pipelines to support predictive analytics and reporting workflows. • Scheduled and monitored batch data workflows using Apache Airflow, ensuring reliable pipeline execution and automated failure alerting across data ingestion jobs. • Containerized data pipeline components using Docker and deployed on Kubernetes clusters, enabling scalable and reproducible data processing environments. Technologies Used: Python, SQL, PostgreSQL, MySQL, ETL Pipelines, Machine Learning, Data Analysis, Visualization Tools, Apache Airflow, Docker, Kubernetes, Scikit-learn, Pandas, NumPy
Surya Singh Kafle's Contact Information
Phone
Find the Right Leads
Find Verified Contact Data
What LeadContact does well
Find verified emails, phone numbers, and decision-makers with 98% accuracy.
Find Leads
Find the right people by company, role, industry, location, and more.
925M+ professional profiles

Find Emails
Access verified email addresses for your target contacts.
657M+ emails

Find Phone Numbers
Get cross-validated phone data from multiple top sources.
239M+ phone numbers

More Accurate. Lower Cost.
Find contact data in 1 tool with 98% accuracy
LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.
Great conversations start with the right contact.
It’s time to find yours.


