Sri Harsha Singudasu

Sri Harsha Singudasu

Site Reliability Engineer @ NVIDIA

About

I am Sri - a software engineer and a self-propelled engine with over 14 years of experience, driven by a thirst for making services reliable at any cost!

Country

United States

City

San Francisco Bay Area

Industry

Computer Software

Skill

Java, Cassandra, Hadoop, Cloud Computing, MapReduce, blueflood, metrics query engine, Linux, Bash, Sqoop, Hive, Flume, grafana, RDBMS, C, C++, Data Structures, JSON, Apache Pig, JMX

Experience

NVIDIA

Site Reliability Engineer

NVIDIA

LinkedIn
2025-5 - Present · 1 yr 5 mos

Santa Clara, CA

GPU Clusters -

Twilio

Site Reliability Engineer

Twilio

LinkedIn
2022-5 - 2025-5 · 3 yrs 1 mo

San Francisco, California, United States

@Segment, Data Platform Engineering - First SRE in the Segment Data platform Engineering team - Contributed in re-architecting the existing monolith Data Ingestion platform to a scalable and fault tolerant shard architecture - Participated in stress testing the Data platform clusters to brace for crucial customer events and corresponding traffic surges - Involved in planning and migration of Kafka clusters from AWS ECS to AWS EKS and/or AWS MSK - Collaborated with teams across the R&D org to ensure their services have multi-region foot print - Contributed to open source libraries and tools that help with Kafka management. - Focussed on automation heavy solutions by doing various Proof of concepts and presented to whole organization. These solutions reduce the toil dramatically for day to day operations. - Helped in reducing costs to the company by right sizing the Data Platform clusters - Improved and Enhanced in house tools that help operations - Enhanced observability by adding new metrics to the services and fixed dashboards to remove any monitoring gaps - Worked and helped my team and also various cross teams with code review processes - Provided on-call rotation and RCA support for our business-critical infrastructure - Personally involved in triaging and troubleshooting production level SEVs to find root cause and help teams to fix their applications to make the eco system reliable. - Involved in the hiring process for new additions to the team Environment: Site Reliability Engineering, RHEL, Open source contribution, Shell Scripting, Python, Java, GoLang, Docker, Kubernetes, Helm, ArgoCD, Argo Workflows, Automation solutions, Git, Terraform, AWS ECS, AWS EKS, AWS MSK, Kafka, Zookeeper, Prometheus, DataDog, Grafana, Micro services, JSON, Rest-APIs, NSQ

Apple

Service Reliability Engineer

Apple

LinkedIn
2020-5 - 2022-5 · 2 yrs 1 mo

Cupertino, California, United States

AMP (Apple Media Products) - Owned and managed the operational aspects of crucial Apple Media Products applications in Production and various QA environments - Automated tasks to ensure very less/no toil in day to day tasks - Built and enhanced tools to ensure the reliability of services - Worked with various teams in finding right KPIs to achieve good visibility on the applications - Provided Oncall support to the owned application and did RCA for the issues and performed necessary actions so that we won’t face the same issues again - Involved in Database failover activity with zero downtime to applications - Performed decommission of applications from datacenter with no impact to users - Made applications complaint with security team standards - Involved in migration activities from baremetal/cloud to Kubernetes Environment: Site Reliability Engineering, Python, Java, Cloud computing, Cassandra, Linux, Shell Scripting, Ansible, Rest-API, JSON, ElasticSearch, Kibana, Zookeeper, Kafka, Grafana, Git, Node.js, Prometheus, Hadoop, Hive, Docker, Kubernetes and MicroServices

Walmart

Site Reliability Engineer

Walmart

LinkedIn
2017-4 - 2020-5 · 3 yrs 2 mos

Sunnyvale, California, United States

Global eCommerce @WalmartLabs, a division of Wal-Mart Stores, Inc. ("Company") - Design and build tools to improve reliability, latency, availability and scalability of Wal-Mart e-Commerce products. - Build tools and automate to prevent reoccurrence of problem to critical products and services. - Provide on call support to all the services of Wal-Mart Global eCommerce. - Perform root cause analysis for issues occurring in complex applications involving multiple services with temporary and permanent solutions. - Debug and improve the performance of critical services of Wal-Mart Global eCommerce sites which includes deep diving into the operating systems internals and network-related issues. - Collect the health and status of Load balancers and VIPs for the services. - Build and drive automation systems that maintain health. Check the status of the applications and keep them notified by creating action items in the JIRA board. - Develop intelligent alerts and monitors on top of key metrics for early failure detection, which involves data analysis and machine learning. Assist in "always up" services for Wal-Mart eCommerce. - Involved in BCDR for services and infrastructure. Environment: Site Reliability Engineering, OneOps, Automation, Cloud computing, BCDR, Linux, Shell Scripting, Ansible, Consul, RestFul-API, API Framework with Swagger, JSON, ElasticSearch, Kibana, Seyren, Cerebro, Kafka, Grafana, Graphite, Maven, Git, JIRA, Tomcat, Javascript, NodeJS, Angular, Redis, Prometheus, Java, jmx_exporter, Jenkins, OpenTSDB, Hadoop, HDFS, R, Rstudio, Python, Hive, HiveQL, MySql, Cassandra, DataStax, MongoDb, Docker, and MicroServices.

Staples

Global Technology Intern / Site Reliability Engineering

Staples

LinkedIn
2016-5 - 2016-12 · 8 mos

Framingham, Massachusetts, United States

Real time monitoring (RTM) - an open source monitoring stack - Analyzed and performed enhancements in Blueflood (Open source java application by Rackspace). - Designed and implemented Restful-APIs using Java, Servlet, Jolokia (Http Bridge) and Embedded Jetty to capture JMX (metrics) of Blueflood, Cassandra and WebSphere(SOAP way) applications. - Developed Bash scripts to start RTM applications and activate load balanced applications by changing their health status. - Developed adapters in Java to find the Health Status of RTM applications. - Coded Java program to find pending processes in Blueflood application. - Created Java program to clean Metric Query Engine(GoLang tool) data in Cassandra. - Performed load testing on Blueflood and Cassandra before pushing to production environment. - RTM has achieved more than million successful hits to the database by just integrating open source tools also facilitating dashboards with Grafana. - Used BitBucket as version management tool to maintain the enhancements and tools developed. - Logged, raised and cleared various tickets in JIRA for the issues and day to day tasks. - Documented analysis on open source tools in Staples internal pages (Confluence). Environment: Java, Cassandra, DataStax, Linux - RHEL, Shell Scripting, Blueflood, RestFul-API, JSON, Grafana, Jolokia, Maven, Stash, BitBucket, Metrics Query Engine, JMX, JIRA, Embedded Jetty, Tomcat, Servlet, JSP, WebSphere, GoLang, Node.js -------- Helios - Client product suggestion system for Staples internal Customer Care - Generated R-Scripts to perform simple analytics on huge customer data in HDFS - Hadoop framework. - Developed Python programs to trigger R-Scripts which store results in Hive partitioned tables. - Implemented Python programs to transfer data present in Hadoop Framework to MySql using Sqoop. - Received appreciation for on time bug free project delivery. Environment: Hadoop, HDFS, R, Python, Hive, HiveQL, Sqoop, MySql

University of New Haven

Student IT helper/worker - Marvin K. Peterson Library

University of New Haven

LinkedIn
2015-9 - 2016-12 · 1 yr 4 mos

West Haven, Connecticut, United States

[ Sep 2015 - May 2016 ] and [ Aug 2016 - Dec 2016 ] - Helped students who have trouble accessing course work materials in Blackboard. - Assisted students and elders in filling forms. - Worked with IT team during system updates/upgrades. - Checked if printers and scanners are up and running. - Performed activities like software installations, word/excel/pdf/jpeg etc. format conversion and printing in desired format.

Capgemini

Associate Consultant

Capgemini

LinkedIn
2014-12 - 2015-7 · 8 mos

Hyderabad Area, India

Euroclear, Belgium - MM Parties - Performed development activities for Euroclear post trading environment with unit, system and regression testing. - Mentored a team of fresher’s in successful knowledge transmission on Telon an internal tool. - Eliminated performance issues in system by finding modules that have slower processing and excessive memory consumption. Domain: Capital Markets

Tata Consultancy Services

Systems Engineer

Tata Consultancy Services

LinkedIn
2011-8 - 2014-9 · 3 yrs 2 mos

Hyderabad Area, India

State Farm Insurance, Bloomington, Illinois - Policy Master Record - Performed maintenance activities by handling priority (I, II & III) incidents (Bug Fixes) and requests (Problem Tickets). - Logging all the tickets in the HPSM tool for future reference. - Handled requirement gathering, development and enhancements with unit, system and regression testing for huge data. - Assisted in sun setting databases for existing applications which helped in migration. - Analyzed systems for the root cause for the problem tickets raised and performed various code fixes. - Testing and pushing the Javascript code for front end change requests received. - Setting up test environment to perform the enhancement activities. - Developed Java classes for the State Farm auto insurance back end modules. - Received customer appreciations for the solutions provided which improved State Farm auto insurance business. Domain: Auto Insurance - Property & Casualty.

Education

University of New Haven

University of New Haven

LinkedIn

Computer Science

2015 - 2016-1 · 1 yr

Area of focus: - Programming - Databases [Advanced and Distributed] - Independent Study on Big Data - Cloud Computing and - Internship @Staples on Real Time Monitoring - Site Reliability Engineering -------- Teaching Assistant for Cloud Computing [UnPaid] [ Jan 2016 - May 2016 ] Helped students with: - Cloud computing concepts Iaas-Paas- Saas. - Microsoft Azure - Virtual machine set up, adding databases etc. - Hadoop framework set up. - Projects.

Gayatri Vidya Parishad College of Engineering (Autonomous), 530048(CC-13)

Gayatri Vidya Parishad College of Engineering (Autonomous), 530048(CC-13)

LinkedIn

Computer Science

2007 - 2011 · 4 yrs

Sri Harsha Singudasu's Contact Information

Email

******@***.com

Phone

(**) *** ****

Find the Right Leads
Find Verified Contact Data

Try with: Jensen Huang @ nvidia.com Click to autofill
LeadContact awards, five-star ratings, and GDPR compliance badges

What LeadContact does well

Find verified emails, phone numbers, and decision-makers with 98% accuracy.

Find Leads

Find Leads

Find the right people by company, role, industry, location, and more.

925M+ professional profiles

Find Leads
Find Emails

Find Emails

Access verified email addresses for your target contacts.

657M+ emails

Find Emails
Find Phone Numbers

Find Phone Numbers

Get cross-validated phone data from multiple top sources.

239M+ phone numbers

Find Phone Numbers

More Accurate. Lower Cost.

Find contact data in 1 tool with 98% accuracy

LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.

LeadContact Logo
Competitor Tools

All these = $289 per month

Great conversations start with the right contact.

It’s time to find yours.