Ramya Jasthi
Palantir Data Engineer/Senior Data Engineer @ Stanford University
About
I’m a Senior Data Engineer having around 10 years of experience designing and building scalable, cloud-native data platforms across healthcare, financial services, insurance, and retail domains. I specialize in distributed data processing and modern data architectures across multicloud ecosystems.My core expertise includes Apache Spark (PySpark, Spark SQL),Databricks, Delta Lake, Snowflake, and Apache Airflow. I’ve built large-scale batch and near real-time data pipelines, implemented Bronze/Silver/Gold lakehouse architectures, and optimized high-volume workloads processing hundreds of gigabytes to terabytes of data daily.I have strong experience in performance tuning, schema enforcement, incremental processing, data reconciliation, and building production-grade pipelines with high reliability and SLA adherence. I’ve also worked extensively on secure cloud integrations, governance, and regulatory reporting use cases.I enjoy solving complex data challenges, collaborating with cross-functional teams, and transforming raw data into scalable, high-performance data solutions that enable analytics and informed decision-making.Currently open to opportunities where I can contribute my expertise in cloud data engineering, distributed systems, and modern data platform design to build impactful enterprise solutions.
-
United States
Information Technology & Services
SAP Products, Apache Sqoop, Realtional Databases, My SQL, Hadoop, MapReduce, Hive, HiveQL, Amazon S3, AWS EMR, PySpark, Amazon Web Services (AWS), Data Loading, Data Transformation, Amazon Elastic MapReduce (EMR), Databricks Products, Python (Programming Language), Apache Spark, SQL
Experience

Palantir Data Engineer/Senior Data Engineer
California, United States
Designed and developed scalable data pipelines to process large structured and semi-structured datasets. Built modular and reusable data workflows to support production-grade data processing. Developed and maintained semantic data models to define relationships across datasets. Implemented data transformation, validation, and enrichment logic using Python, PySpark, and SQL. Designed and delivered interactive dashboards and reporting solutions for operational analytics. Built and supported application workflows integrating data pipelines with business processes. Performed exploratory data analysis and data processing using notebook-based development. Optimized data pipelines and queries for performance, scalability, and efficiency. Integrated data from multiple sources including distributed systems and relational platforms. Ensured data quality, governance, and consistency using validation and reconciliation frameworks. Developed reusable components for data ingestion, transformation, and pipeline standardization. Collaborated with cross-functional teams to translate business requirements into data solutions. Supported development of data-driven applications and analytics workflows. Improved pipeline reliability through monitoring, debugging, and performance tuning. Delivered curated datasets to support analytics, reporting, and decision-making needs.

Databricks Data Engineer/Senior Data Engineer
Overland Park, Kansas, United States
Client: Databricks Designed and developed scalable data pipelines for large-scale data processing and analytics. Built distributed data processing workflows using PySpark and SQL for structured and semi-structured data. Migrated legacy data workflows into modern, scalable data processing frameworks. Developed and optimized data transformation pipelines for performance and efficiency. Integrated data from multiple sources including databases, files, APIs, and streaming systems. Implemented data validation, reconciliation, and consistency checks across systems. Built reusable and configurable data pipeline frameworks for standardized processing. Orchestrated workflows with scheduling, dependency management, and monitoring. Processed high-volume datasets with complex transformations and aggregations. Supported data platform modernization and migration initiatives. Developed scripts for data validation, deduplication, and anomaly detection. Performed source-to-target validation ensuring accuracy and completeness of data. Evaluated modern data platforms and architectures to support scalable data solutions. Designed data models to support analytics, reporting, and business insights. Collaborated with cross-functional teams to deliver data-driven solutions.

Senior Data Engineer
Client: Health care Developed pipelines for processing high-volume data in near real-time environments. Built and maintained streaming and batch data processing workflows using Azure tools. Processed large datasets for operational analytics and reporting use cases. Implemented data ingestion pipelines using Azure-based storage and processing services. Developed transformation logic for cleansing, enrichment, and aggregation of data. Designed systems to handle late-arriving and historical data processing. Implemented schema validation and data consistency checks. Performed data validation and reconciliation using SQL and distributed processing tools. Ensured fault tolerance and reliability of data pipelines. Built event-driven workflows using Azure-based scheduling and orchestration tools. Integrated data from multiple systems including streaming sources and relational databases. Optimized data pipelines for performance and scalability. Supported data integration and reporting requirements across business teams.

Data Engineer
Missouri, United States
Developed scalable batch data pipelines using PySpark on AWS EMR. Processed large datasets (500GB–1TB) for financial analytics and reporting. Designed data lake architecture (landing, staging, curated layers) in Amazon S3. Orchestrated workflows using Apache Airflow with SLA monitoring and retries. Optimized Hive and Spark queries using partitioning and bucketing. Implemented complex transformations using Spark SQL and PySpark. Integrated Snowflake with AWS S3 for reporting pipelines. Performed data validation and reconciliation using AWS Athena. Tuned EMR clusters for performance and resource optimization. Collaborated with finance teams for regulatory reporting solutions.

Data Engineer
Hyderabad
• Developed PySpark-based batch pipelines on AWS EMR to enhance data processing efficiency. • Built ingestion pipelines using Sqoop for seamless data transfer from relational databases to Hadoop. • Designed and implemented Hive-based analytical data models for comprehensive insurance reporting. • Automated workflows with Apache Airflow, ensuring timely scheduling and monitoring of data processes.

Big Data Developer
India
• Developed robust ETL pipelines utilizing Hive, Sqoop, and various Hadoop ecosystem tools to enhance data processing efficiency. • Designed analytical data models focused on retail sales and inventory reporting, supporting informed business decisions. • Built ingestion pipelines from MySQL and SQL Server into Hadoop, ensuring seamless data integration.
Ramya Jasthi's Contact Information
Phone
Find the Right Leads
Find Verified Contact Data
What LeadContact does well
Find verified emails, phone numbers, and decision-makers with 98% accuracy.
Find Leads
Find the right people by company, role, industry, location, and more.
925M+ professional profiles

Find Emails
Access verified email addresses for your target contacts.
657M+ emails

Find Phone Numbers
Get cross-validated phone data from multiple top sources.
239M+ phone numbers

More Accurate. Lower Cost.
Find contact data in 1 tool with 98% accuracy
LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.
Great conversations start with the right contact.
It’s time to find yours.



