Nishanth Reddy Palugula
Data Scientist @ IBM
About
Data Scientist with 4+ years building production forecasting systems and cloud data engineering pipelines across AWS, GCP, and Azure. Specialized in demand forecasting, time series modeling, and large-scale ML infrastructure — with proven impact in fraud detection ($1.85M+ prevented), supply chain analytics, and automated cloud-native pipelines processing 5M+ records. What I Do: I engineer end-to-end data science systems — from raw data ingestion through model deployment — with a specialization in forecasting and time series at scale. My work sits at the intersection of data engineering and applied ML, where production systems matter as much as model accuracy. Core Expertise: Demand forecasting & time series modeling (ARIMA, Prophet, LSTM, gradient boosting) Cloud data engineering & ML pipelines — AWS (S3, SageMaker, Redshift), GCP (BigQuery, Pub/Sub), Azure Large-scale data processing with PySpark / Apache Spark on 300K–5M+ record datasets Anomaly detection, fraud detection ML models — $1.85M+ in prevented losses Automated Python pipelines, model monitoring, drift detection, CI/CD for ML workflows LLM / RAG system architecture and agentic workflow design Recent Impact: IBM: Built forecasting models on 5M+ monthly records — reduced analysis turnaround by 30%, improved forecast stability by 25% Mphasis: Forecasting + classification models on 300K+ daily records — improved KPIs by 10–15%, cut manual effort by 40% iScientific: GCP-native data ingestion pipeline (BigQuery, Pub/Sub) processing 500K+ daily IoT sensor readings
United States
Austin
Computer Software
Data Structures, Business Analysis, Google BigQuery, Demand Forecasting, Anomaly Detection, casual inference, Data Pipelines, Time Series Analysis, Cohort Analysis, Google Cloud Platform (GCP), Predictive Maintenance, ETL Tools, data quality management, Amazon Web Services (AWS), Apache Spark, Kubernetes, real time systems, Fraud Detection, Extract, Transform, Load (ETL), Data Science
Experience

Data Scientist
Austin, Texas, United States
Building production forecasting systems and ML pipelines on cloud infrastructure (AWS, GCP) to support demand planning, cost optimization, and AI-driven decision-making for finance and operations stakeholders at scale. -Built and deployed production time series forecasting models on 5M+ monthly usage records using Python (Prophet, scikit-learn, gradient boosting), supporting demand planning, cost optimization, and capacity decisions for finance and operations stakeholders -Engineered automated ML pipelines on AWS (S3, Redshift) using PySpark for large-scale feature engineering, model training, and batch inference — reducing analysis turnaround time by ~30% -Designed model monitoring and drift-detection dashboards that tracked forecast accuracy over time, reducing post-deployment degradation issues by ~25% and enabling proactive retraining cycles -Built RAG-based AI system using LLMs to automate extraction and summarization of insights from large-scale usage datasets, reducing manual analyst review time and enabling natural language querying of forecasting outputs by non-technical stakeholders -Prototyped agentic AI workflows integrating LLM orchestration with existing forecasting pipelines, enabling automated anomaly flagging, root cause narration, and decision support for operations and finance teams -Analyzed customer behavior, cohort dynamics, and usage patterns across rapidly evolving datasets to surface retention, expansion, and efficiency signals for quarterly business reviews -Partnered with finance, product, and operations teams to translate business assumptions into robust analytical models, KPIs, and executive-ready decision frameworks

Data Scientist
Hyderabad, Telangana, India
Built end-to-end forecasting, classification, and analytics pipelines on large-scale structured datasets, supporting capacity planning, utilization forecasting, and operational efficiency initiatives across finance and business operations teams. -Developed demand forecasting and classification models on 300K+ daily records using Python, PySpark, and SQL — supporting capacity planning, utilization forecasting, and operational efficiency initiatives, improving key business KPIs by 10–15% -Built scalable SQL and PySpark-driven analytics pipelines with automated data workflows, reducing manual reporting effort by 40% and enabling real-time operational decision-making across cross-functional teams -Conducted causal and exploratory analyses to identify drivers of growth, risk, and anomalous behavior — applying regression, clustering, and time series decomposition on large structured datasets to surface actionable signals -Designed production model validation frameworks including back-testing, holdout evaluation, and drift monitoring to ensure forecast accuracy and reliability over time in live environments -Built data quality and consistency frameworks that standardized upstream inputs across downstream analytics, ML model serving pipelines, and reporting systems — ensuring end-to-end data integrity at scale

Data Scientist
Hyderabad, Telangana, India
Built GCP-native data engineering pipelines and ML models on IoT sensor data, delivering real-time anomaly detection, time series analysis, and automated data workflows for operational monitoring and early warning systems. -Built a GCP-native real-time data ingestion pipeline using BigQuery, Cloud Storage, and Pub/Sub — processing 500K+ daily IoT sensor readings with automated transformation, load workflows, and end-to-end data quality checks -Developed anomaly detection and classification models on 300K+ IoT sensor records achieving ~0.84 AUC, reducing manual review effort by 60% and enabling early warning signals for operational and engineering teams -Performed time series analysis and behavioral pattern recognition on sensor data to identify drift, degradation, and anomalous usage patterns across IoT device fleets — applying decomposition, trend analysis, and threshold-based alerting -Documented all analytical workflows, data assumptions, and model architecture to ensure full reproducibility and seamless knowledge transfer across engineering and analytics teams
Education

Business Analytics
STEM-designated Master's program in Business Analytics covering statistical analysis, data modeling, predictive analytics, and business decision frameworks. Applied quantitative methods to real-world business problems across finance, operations, and strategy domains.

Data Science and Artifical Intelligence
Bachelor of Technology in Data Science and Artificial Intelligence covering programming fundamentals, statistical analysis, database management, data structures and algorithms, machine learning, neural networks, natural language processing, and big data technologies. Completed applied projects in anomaly detection and IoT data analytics.
Nishanth Reddy Palugula 's Contact Information
Phone
Find the Right Leads
Find Verified Contact Data
What LeadContact does well
Find verified emails, phone numbers, and decision-makers with 98% accuracy.
Find Leads
Find the right people by company, role, industry, location, and more.
925M+ professional profiles

Find Emails
Access verified email addresses for your target contacts.
657M+ emails

Find Phone Numbers
Get cross-validated phone data from multiple top sources.
239M+ phone numbers

More Accurate. Lower Cost.
Find contact data in 1 tool with 98% accuracy
LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.
Great conversations start with the right contact.
It’s time to find yours.

