Ritam Ghosh

Ritam Ghosh

Graduate Research And Teaching Assistant @ Penn State University

About

I’m a engineer focused on Software Dev and managing Data Platforms, who ships with business impact. At EY, I learned that great code isn’t enough—it has to move a number that matters. I build end-to-end products (React/Node, Python/TypeScript, APIs, CI/CD) grounded in cloud & data (AWS/Azure, Kubernetes, SQL/PySpark, Databricks, Azure Data Factory). Let the results speak louder 🔊 • Cut client’s (Santander Bank) annual Azure spend by 8% ($24k/year) by optimizing their ETL pipelines—parallelized ingestion by splitting large customer tables into ID-range batches, rewrote multi-way SQL joins on 100M+ row datasets with predicate pushdown to reduce runtime, and created shared cluster pools for multiple ETL jobs, improving startup latency and overall resource efficiency. • ~40 engineer-hours/week freed by architecting parameterized Azure Data Factory → reusable PySpark notebooks with ElastiCache-backed validation; manual-error rate ~0. • >$350k new revenue unlocked by co-building a transaction-reporting analytics platform (Databricks + PySpark + AWS RDS) and an ESG rules engine (Redshift + PostgreSQL + React BI); platform adapted by HSBC, Deutsche Bank, Hargreaves Lansdown. A bit of human 🏏 I’m exploring the art of learning—systems, spaced repetition, deliberate practice. Off the keyboard, I like being on the winning side of pickleball matches or just watching cricket. 👨🏽‍🎓 MS CS @ Penn State. Open to Software Engineer / Data Engineer roles. 📨 If the mix of full-stack + cloud + data impact fits your team—or you just want to connect—email me at imritam16@gmail.com or DM me to set up a quick chat.

Country

-

City

United States

Industry

Computer Software

Skill

SQLAlchemy, Tailwind CSS, Flask, LangGraph, Vector Databases, Retrieval-Augmented Generation (RAG), Prometheus.io, DevOps, AWS X-Ray, Amazon EKS, Grafana, Software Observability, PySpark, Extract, Transform, Load (ETL), Azure Data Lake, Artificial Intelligence (AI), BERT (Language Model), Natural Language Processing (NLP), Neural Networks, Transformers

Experience

Penn State University

Graduate Research And Teaching Assistant

Penn State University

LinkedIn
2024-8 - 2025-12 · 1 yr 5 mos

Pennsylvania, United States

Working with Prof. Dariush Khezrimotlagh, Prof. Sara Imanpour and Prof. Ilya Shvartsman on creating Early Prediction Methodology of 30-Day Hospital Readmissions. Tasks performed: • Built an early-warning system for 30-day hospital readmissions and reached 71% prediction accuracy, enabling clinicians to flag high-risk cases before discharge and setup end-to-end LLMOps pipeline for training and evaluating model • Fine-tuned a transformer-based Large Language Model (DistilBERT) in PyTorch for readmission prediction; built a custom NLP pipeline to convert structured ICD-10/CPT medical codes into narrative text for model input. • Implemented and benchmarked multiple machine learning models (Random Forest, CatBoost, MLP, and a DistilBERT transformer) on ~50,000 patient records, designing experiments to compare limited early-admission features vs. full encounter data, proving early-admission data is 92 % as predictive as complete records. Teaching Assistant for CMPSC 462 : Data Structures and Algorithms by Prof. Bimal Ghimire Empowered students to master core computer science principles by leading interactive sessions to deconstruct complex topics in data structures and algorithms. Elevated student performance by meticulously evaluating assignments and delivering targeted feedback to improve their problem-solving and coding abilities.

EY

Software Engineer

EY

LinkedIn
2022-10 - 2023-11 · 1 yr 2 mos

Kolkata, West Bengal, India

• Cut client’s (Santander Bank) annual Azure spend by 8% ($24k/year) by optimizing their ETL pipelines—parallelized ingestion by splitting large customer tables into ID-range batches, rewrote multi-way SQL joins on 100M+ row datasets with predicate pushdown to reduce runtime, and created shared cluster pools for multiple ETL jobs, improving startup latency and overall resource efficiency. • Led the end-to-end development and migration of 3 production data ingestion pipelines from StreamSets to Merck’s Data Ingestion Framework, owning the translation of complex StreamSets pipelines into parameterized YAML configurations that compiled via Jenkins into Apache Airflow DAGs, reliably ingesting a wide variety of structured and semi-structured data formats (e.g., Parquet, Avro) from upstream source systems into the Merck’s Datahub. • Eliminated 40 engineer-hours/week by automating 67 % of daily data-prep tasks—built parameterized ADF pipelines that invoke reusable PySpark notebooks and auto-validate outputs in AWS ElastiCache, driving manual-error rate to zero. • Secured >$350 k new revenue by co-building a transaction-reporting & analytics platform (Databricks, PySpark, AWS RDS, Redshift, React BI), then leading client knowledge-transfer that enabled production rollouts at HSBC, Deutsche Bank, and Hargreaves Lansdown. • Owned the data governance framework for the client-facing ESG analytics platform, implementing automated PySpark validation scripts and quality checks that achieved 99.9% data accuracy and ensured trusted reporting for clients like HSBC and Deutsche Bank.

Education

Penn State University

Penn State University

LinkedIn

Computer Science

2024-1 - 2025-12 · 2 yrs

Relevant Coursework with Distinction (4.0 GPA): Advanced Data Structures and Algorithms Advanced Database Management Systems Advanced Operating Systems and Distributed Systems Data Mining and Knowledge Discovery

Institute of Engineering & Management (IEM)

Institute of Engineering & Management (IEM)

LinkedIn

Computer Science

2018-8 - 2022-7 · 4 yrs

During my computer science undergraduate program, I gained proficiency in programming languages such as Python and Java, as well as an understanding of software engineering principles and data structures. Additionally, I learned how to design and manage databases using DBMS. Through practical projects and coursework, I honed my skills in developing efficient algorithms, creating scalable applications, and conducting data analysis. My experience has equipped me to excel in a variety of technical roles and has given me a solid foundation to pursue a career in software development or data analysis.

Nava Nalanda

Nava Nalanda

LinkedIn
2005-8 - 2016-8 · 11 yrs 1 mo

Spent 11 years of my life as a student of this school , which gave me the opportunity and freedom to learn and grow.

Ritam Ghosh's Contact Information

Email

******@***.com

Phone

(**) *** ****

Find the Right Leads
Find Verified Contact Data

Try with: Jensen Huang @ nvidia.com Click to autofill
LeadContact awards, five-star ratings, and GDPR compliance badges

What LeadContact does well

Find verified emails, phone numbers, and decision-makers with 98% accuracy.

Find Leads

Find Leads

Find the right people by company, role, industry, location, and more.

925M+ professional profiles

Find Leads
Find Emails

Find Emails

Access verified email addresses for your target contacts.

657M+ emails

Find Emails
Find Phone Numbers

Find Phone Numbers

Get cross-validated phone data from multiple top sources.

239M+ phone numbers

Find Phone Numbers

More Accurate. Lower Cost.

Find contact data in 1 tool with 98% accuracy

LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.

LeadContact Logo
Competitor Tools

All these = $289 per month

Great conversations start with the right contact.

It’s time to find yours.