Ayush Khaneja

Ayush Khaneja

Data Analyst @ Mass General Brigham

About

ML engineer shipping production systems that scale. I build LLM pipelines, optimize models for efficiency, and engineer datasets from raw, messy data - turning research into reliable infrastructure. At Mass General Brigham: Deployed systems processing 200k+ clinical records. Fine-tuned smaller LLMs achieving 88-90% F1 on diagnosis detection - outperforming models 8-10x larger. Built custom RAG and chunking that reduced manual effort by 100+ seconds per task. Worked end-to-end from dataset engineering to AWS infrastructure. At Deloitte: Handled financial analysis on million-record datasets. Automated data pipelines saving 48-72 hours of manual work per project. Built Tableau dashboards and VBA tooling for stakeholder delivery. What I actually do: Make expensive problems cheap. I find where organizations waste compute, annotation time, and engineering cycles - then redesign from first principles. Fine-tune models ruthlessly. Engineer datasets to eliminate bottlenecks. Deploy on infrastructure that scales without breaking budgets. The result: systems that perform like enterprise solutions, built lean enough to iterate fast. Published: AMIA 2025, ICPRAM 2020. MS Data Science, Indiana University (3.87 GPA). 4+ years across production ML, financial analytics, and large-scale data systems. Code-obsessed. Ready for the next hard problem.

Country

-

City

United States

Industry

Hospital & Health Care

Skill

Statistical Data Analysis, Oral Communication, Sentiment Analysis, Amazon Web Services (AWS), Apache Spark, Hadoop, Snowflake, PySpark, Google Cloud Platform (GCP), Regression Models, Skilled Multi-tasker, Presentation Skills, Data Engineering, Statistical Analysis, PyTorch, Interpersonal Skills, Attention to Detail, C++, TensorFlow, Keras

Experience

Mass General Brigham

Data Analyst

Mass General Brigham

LinkedIn
2024-6 - Present · 2 yrs 4 mos

Somerville, Massachusetts, United States

▪ Engineer 200k+ unstructured clinical text records from MIMIC-IV data to create artificially labelled datasets, eliminating annotation costs. ▪ Fine-tune and refine smaller LLMs (Llama, Mistral) for private, resource-efficient inferencing with performance like models 8-10x in size. ▪ Implement custom chunking and RAG techniques to further enhance LLM efficiency. ▪ Build and deploy an AWS-based LLM pipeline using Python, Transformers, HuggingFace, and vLLM to structure medical notes, generate precise summaries, and detect mentions of diagnoses, reducing manual effort by over 100 seconds per note

Mass General Brigham

Data Analyst Intern

Mass General Brigham

LinkedIn
2023-6 - 2024-5 · 1 yr

Somerville, Massachusetts, United States

▪ Developed vector-based predictive models using Python, Scpacy, NLTK and TensorFlow to predict diagnosis occurrence with 84-88% F1 score. ▪ Deployed LLMs locally on AWS EC2 instances to extract actionable diagnostic insights from unstructured clinical notes for predictive analysis.

Deloitte

A&A Analytics Specialist Senior Assistant

Deloitte

LinkedIn
2022-4 - 2022-6 · 3 mos

Hyderabad, Telangana, India

• Delivered financial analysis through SAS programming on client financial data to aid the Audit team in their external audit processes • Developed reusable Tableau visualizations and dashboards from scratch to present relevant findings and insights from audit and other financial data • Managed and maintained a Robotic Process Automation pipeline using UiPath, to automate time-consuming client-related tasks, saving 48-72 hours in manual work per week

Deloitte India (Offices of the US)

A&A Analytics Specialist Assistant

Deloitte India (Offices of the US)

2020-8 - 2022-4 · 1 yr 9 mos

Hyderabad, Telangana, India

• Assisted in the development of an automated pipeline to handle client data, allowing the firm to divert valuable man hours (48-72 hours) to other processes

Deloitte

A&A Specialist Intern

Deloitte

LinkedIn
2019-12 - 2020-3 · 4 mos

Hyderabad Area, India

• Conducted journal entry testing procedures in SAS to prevent management override of control, improving reliability of financial statements • Performed data completeness and data quality checks on large quantities of data

CDAC,Noida

Health Analytics Intern

CDAC,Noida

LinkedIn
2019-5 - 2019-7 · 3 mos

Noida, Uttar Pradesh, India

Worked on highly unbalanced data, came up with distance-based approaches to classify data with high bias and implemented a Siamese neural network architecture. Research was presented at ICPRAM 2020.

STMicroelectronics

Raspberry Pi Specialist Intern

STMicroelectronics

LinkedIn
2018-5 - 2018-6 · 2 mos

Noida, Uttar Pradesh, India

Flying drones using Raspberry Pi and computer vision.

Education

Indiana University Bloomington

Indiana University Bloomington

LinkedIn

Data Science

2022-8 - 2024-5 · 1 yr 10 mos
Vellore Institute of Technology

Vellore Institute of Technology

LinkedIn

Computer Science

2016 - 2020 · 4 yrs

Ayush Khaneja's Contact Information

Email

******@***.com

Phone

(**) *** ****

Find the Right Leads
Find Verified Contact Data

Try with: Jensen Huang @ nvidia.com Click to autofill
LeadContact awards, five-star ratings, and GDPR compliance badges

What LeadContact does well

Find verified emails, phone numbers, and decision-makers with 98% accuracy.

Find Leads

Find Leads

Find the right people by company, role, industry, location, and more.

925M+ professional profiles

Find Leads
Find Emails

Find Emails

Access verified email addresses for your target contacts.

657M+ emails

Find Emails
Find Phone Numbers

Find Phone Numbers

Get cross-validated phone data from multiple top sources.

239M+ phone numbers

Find Phone Numbers

More Accurate. Lower Cost.

Find contact data in 1 tool with 98% accuracy

LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.

LeadContact Logo
Competitor Tools

All these = $289 per month

Great conversations start with the right contact.

It’s time to find yours.