Anisha Kande

Anisha Kande

Data Engineer @ Allstate

About

Data Engineer with 3+ years of experience designing and scaling distributed data systems across AWS and Azure, specializing in building high-throughput ETL pipelines, real-time streaming architectures, and analytics-ready data platforms. I have hands-on experience processing large-scale datasets (10M+ records/day, ~500GB) using PySpark, SQL, and modern data frameworks to support critical business and data-driven decision-making. In my current role, I’ve developed event-driven data pipelines and real-time ingestion systems leveraging AWS Lambda, DynamoDB, Kafka, and Elasticsearch, significantly reducing data latency and improving query performance. I’ve also designed governed data models using dbt, Snowflake, and Redshift, enabling standardized KPIs and improving reporting accuracy for business stakeholders. I bring strong experience in enabling AI/ML and Generative AI use cases by building scalable data pipelines for feature engineering, model training, and deployment. I have implemented end-to-end Retrieval-Augmented Generation (RAG) pipelines using LangChain, FAISS, and Pinecone to process and index large volumes of unstructured data, enabling semantic search and improving information retrieval efficiency. My work includes optimizing embedding pipelines, evaluating retrieval quality, and enhancing response accuracy for LLM-based applications. Additionally, I have experience building batch and real-time data ingestion frameworks using Apache Kafka, Apache Flink, and Apache NiFi, as well as orchestrating complex workflows using Apache Airflow. I focus on delivering reliable, production-grade systems with high data quality, scalability, and performance. I’m passionate about solving complex data challenges, optimizing large-scale systems, and building AI-ready data platforms that bridge data engineering with machine learning and Generative AI. I thrive in fast-paced, high-impact environments and enjoy collaborating with cross-functional teams to deliver scalable, data-driven solutions.

Country

United States

City

St Louis

Industry

Primary/Secondary Education

Skill

Job Simulations, python, ETL Tools, Data Transformation, Airflow, EDA, Anomaly Detection, pipeline documentation, Scikit-Learn, Data Pipelines, Data Modeling, TensorFlow, Correlation Analysis, Visual Studio, C++, GNU Compiler Collection (GCC), LaTeX, Ec2, Kubernetes, CI/CD

Experience

Allstate

Data Engineer

Allstate

LinkedIn
2025-6 - Present · 1 yr 4 mos

United States

• Built scalable Spark and PySpark pipelines on AWS EMR and Databricks to process 10M+ records/day (~500GB data) across policy records, claims transactions, customer interaction logs, and risk datasets, improving data availability for underwriting & actuarial teams by 40%. • Developed event-driven ingestion workflows using AWS Lambda, DynamoDB, and Elasticsearch to enable near real-time indexing of 500K+ daily events (claims, behavioral data, agent logs), reducing search latency and operational query time by 35%. • Designed governed analytics models in dbt on Snowflake and Amazon Redshift, standardizing 20+ insurance KPIs such as loss ratios, claim severity, and retention metrics, improving Tableau dashboard accuracy and reducing data reconciliation issues by 30%. • Implemented RAG-based data workflows using LangChain, FAISS, and Pinecone to support semantic search across 100K+ unstructured documents (~150GB data) including policy documents and claims notes, cutting manual review effort by 25%. • Optimized Spark-based feature pipelines for clustering and decision-tree models using multi-year fraud datasets (5M+ records) and customer segmentation data, increasing model precision by 20%. • Collaborated with data science, compliance, and business stakeholders to productionize analytics and AI pipelines, deploying 10+ production workflows and accelerating deployment cycles by 30% while strengthening governance across data products.

Hexaware Technologies

Data Engineer

Hexaware Technologies

LinkedIn
2021-7 - 2023-11 · 2 yrs 5 mos

India

• Built end-to-end data pipelines using Python, PySpark, and Spark SQL on Azure Data Lake, processing enterprise datasets including transactional data, customer records, application logs, and operational metrics, improving data throughput by 35%. • Implemented batch and real-time ingestion frameworks with Apache Kafka, Apache NiFi, and Apache Flink to stream event-driven datasets such as user activity logs, system telemetry, and transaction streams, reducing reporting latency by 40%. • Designed lakehouse architecture using Apache Iceberg and Apache Nessie to manage versioned datasets and schema evolution acro ss structured and semi-structured data, improving consistency and reducing rollback-related incidents by 30%. • Orchestrated 100+ scheduled ETL pipelines using Apache Airflow and Informatica PowerCenter, managing multi-source datasets from APIs, databases, and streaming systems, achieving pipeline success rates of 99.5%. • Created analytics-ready datasets in Snowflake and PostgreSQL using aggregated business metrics, forecasting datasets, and performance indicators, reducing manual analytical effort by 25%. • Delivered dashboards and reports in Power BI and Excel using curated datasets from multiple data sources, improving reporting accuracy by 20% and enabling faster decision-making.

Education

Saint Louis University

Saint Louis University

LinkedIn
2024-1 - 2025-12 · 2 yrs

Electronics & Communication Engineering

2019-6 - 2023-6 · 4 yrs 1 mo

Anisha Kande's Contact Information

Email

******@***.com

Phone

(**) *** ****

Find the Right Leads
Find Verified Contact Data

Try with: Jensen Huang @ nvidia.com Click to autofill
LeadContact awards, five-star ratings, and GDPR compliance badges

What LeadContact does well

Find verified emails, phone numbers, and decision-makers with 98% accuracy.

Find Leads

Find Leads

Find the right people by company, role, industry, location, and more.

925M+ professional profiles

Find Leads
Find Emails

Find Emails

Access verified email addresses for your target contacts.

657M+ emails

Find Emails
Find Phone Numbers

Find Phone Numbers

Get cross-validated phone data from multiple top sources.

239M+ phone numbers

Find Phone Numbers

More Accurate. Lower Cost.

Find contact data in 1 tool with 98% accuracy

LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.

LeadContact Logo
Competitor Tools

All these = $289 per month

Great conversations start with the right contact.

It’s time to find yours.