Chhaya Tundwal
Software Engineer - Data & AI Integration @ OurFirm.ai
About
Data & AI Engineer with experience building large-scale data systems, backend platforms, and AI-driven pipelines. I’ve worked on high-throughput financial systems at Fidelity, developing Kafka pipelines processing 100K+ daily events and AWS-based workflows saving $10M+ annually. More recently, I’ve been working on large-scale document intelligence systems at OurFirm.ai, building pipelines that transform millions of legal opinions into structured datasets using Python, SQL, and LLM-based extraction workflows. I enjoy solving complex data problems, designing reliable pipelines, and turning messy data into production-ready systems. I thrive in fast-paced environments where I can move quickly, experiment with new technologies, and build systems that operate at scale. Currently exploring opportunities in Software Engineering, Data Engineering, AI Engineering, and Data & Analytics roles.
United States
New York
Information Technology & Services
AWS Glue, Pinecone.io, n8n, Git, MongoDB, Docker, Nginx, GitHub, Leadership, Problem Solving, Google Search Console, Logistic Regression, Azure Language Service, Text-to-Keyword extraction, Optuna- Model Performance Tuning, Classification, Graph Algorithms, Dynamic Programming, Integer Linear Programming, PostgreSQL
Experience

Software Engineer - Data & AI Integration
New York City Metropolitan Area
OurFirm.ai builds large-scale legal intelligence systems by transforming raw court documents into structured, searchable data. • Built a large-scale AI document processing pipeline for 10M+ legal opinions, transforming court documents into structured JSON and production-ready HTML. • Designed a page number and citation recovery system that preserves reporter citations, footnotes, and metadata across heterogeneous sources including HTML, plain text, and PDFs. • Developed Python data processing pipelines to normalize legal documents and recover missing structural elements such as page markers, paragraph boundaries, and citations. • Wrote complex PostgreSQL queries across the CourtListener dataset to analyze opinion clusters, citations, and metadata relationships across millions of records. • Implemented automated validation checks (12+ rules) to detect rendering issues such as missing page markers, broken footnotes, and dropped text before documents are published. • Optimized LLM extraction pipelines using Gemini (Vertex AI) to process millions of documents while minimizing token cost and reprocessing. • Managed supporting AWS infrastructure (RDS, EC2, IAM) for large-scale document processing pipelines and database environments. Scale: 10M+ legal documents processed, multi-million record datasets.

AI Engineer
Lafayette, Indiana, United States
built a fully functional System Design Tutor web application that guides learners through multi-level system design challenges using dynamic, AI-generated level plans. I engineered the platform end-to-end in Next.js/TypeScript, designed a robust state management layer to handle multi-step workflows, and implemented reliable level-evaluation logic that advances only when the AI explicitly signals readiness.

Software Engineer – GenAI Workflows
West Lafayette, Indiana, United States
Automated prompt-driven code documentation pipeline using Python, Gemini Pro, Cursor and SQL, cutting manual documentation effort 70% and migration time 50%. Designed ETL workflows with n8n + LangChain, aggregating 5K+ monthly customer reviews and did sentiment analysis. Created a prompt-engineering workflow using Cursor and Gemini Pro to auto-generate comprehensive documentation for an AngularJS codebase, significantly reducing manual writing effort. Led the migration of an AngularJS application to React.js using prompt engineering to streamline code translation.

Website Developer – Digital Marketing
West Lafayette, Indiana, United States
Automated ETL workflows with Python & SQL to integrate Google Analytics & Salesforce into unified reporting views. Collaborated with faculty and marketing teams to translate business requirements into dashboards (Tableau, Looker). Improved campaign conversion rates by 12% through actionable insights derived from centralized data pipelines. Built Looker Studio dashboards integrating Google Search Console, LinkedIn, and Salesforce data for real-time campaign insights. Improved PMP website performance from 39 → 90 on PageSpeed Insights, increasing engagement and SEO visibility. Implemented A/B testing & analytics pipelines for digital campaigns, boosting CTR and lead conversion.

AI Engineer (LLMs, Vector DBs, LangChain)
West Lafayette, Indiana, United States
Designed Python pipeline to transform LinkedIn data into structured formats, boosting lead qualification accuracy by 80%. Integrated data into NoSQL stores (MongoDB, DynamoDB) and streaming agents for scalable enrichment and retrieval. I worked on simulating AI-powered B2B sales pipelines using large language models (LLMs) integrated with LangChain and Prediction Guard’s secure GenAI platform. My responsibilities included designing and implementing agentic AI modules to automate key sales tasks such as lead research, qualification. Evaluate AI agent performance against human benchmarks across metrics like cost, speed, and accuracy. Additionally, contributed to building a simulation environment that helped the company benchmark AI autonomy levels, optimize internal sales workflows, and create a case study demonstrating responsible AI adoption for enterprise clients.

Software Engineer (Data Pipelines & Cloud Migration)
India
Built and scaled data pipelines using Kafka (Pub/Sub), DynamoDB Streams, and Oracle PL/SQL, transporting and validating 10M+ financial transactions/day with automated governance checks that prevented $12.5M in potential losses. Migrated 100+ legacy workflows from COBOL and on-prem systems to AWS Step Functions, Batch, and EKS, enabling event-driven automation and reducing manual intervention. Designed REST APIs and backend services in Java (Spring Boot) and AWS Lambda to securely access customer and transaction data for analytics, audit, and compliance. Engineered ETL and data validation layers in Python, implementing schema checks, lineage documentation, and Pytest-based quality gates for reliable ingestion and transformation. Optimized analytical workloads by designing Snowflake and BigQuery schemas, reducing query runtimes by 40% and improving data discoverability across teams. Standardized CI/CD pipelines with Terraform, Jenkins, and Docker, cutting deployment times by 40% and improving developer velocity by 60%. Implemented Kafka event publishers streaming 100K+ daily account events, improving downstream observability and near-real-time reporting. Monitored production workloads with Datadog, Splunk, and CloudWatch, maintaining 99.9% uptime and <5-minute MTTR during on-call rotations. Partnered with compliance and business teams to modernize Account Opening and EFT workflows, increasing scalability and delivering $10M+ annual savings.

Software Engineer
Bengaluru, Karnataka, India
Developed a robust automation pipeline in Python and PL/SQL stored procedures to process 1M+ transactions/day, performing CRUD operations on Oracle databases and triggering alerts when business thresholds were exceeded.

Software Engineer - Supply Chain
National Institute of Technology Calicut
Calicut Area, India
Developed a Decision Support System (DSS) for Multimodal Freight Transportation Planning for Food Corporation of India(FCI) Kerala Region, which helped reduce the cost by $2 million and transporting time of grains by months. Deployed the web app, database and Nginx using Docker Compose, GitHub Actions and scripting. Social Impact of the project was that it reduced the food wastage as the time and weight of the freight was distributed according to the capacity and requirement in the region.
Chhaya Tundwal's Contact Information
Phone
Find the Right Leads
Find Verified Contact Data
What LeadContact does well
Find verified emails, phone numbers, and decision-makers with 98% accuracy.
Find Leads
Find the right people by company, role, industry, location, and more.
925M+ professional profiles

Find Emails
Access verified email addresses for your target contacts.
657M+ emails

Find Phone Numbers
Get cross-validated phone data from multiple top sources.
239M+ phone numbers

More Accurate. Lower Cost.
Find contact data in 1 tool with 98% accuracy
LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.
Great conversations start with the right contact.
It’s time to find yours.



