
Tarun Jangam
Site Reliability Engineer @ Wells Fargo
About
As a Site Reliability Engineer at DBS Bank, I leverage my skills in Linux, Node.js, and Systems Management to ensure the reliability and performance of critical applications. I have successfully established comprehensive monitoring solutions using Grafana, ELK stack, Filebeat, and Metricbeat, which have enhanced system visibility and reliability. I also implemented robust alerting mechanisms for application services, ensuring immediate notifications during downtimes to maintain optimal service availability. Additionally, I conduct daily health checks of critical services, proactively identifying and resolving potential issues to minimize disruptions. I am currently pursuing a Master of Science in Computer Science at the University at Buffalo, where I am expanding my knowledge and skills in the field. I am passionate about applying my technical expertise to solve complex problems and improve system reliability.
United States
Buffalo
Computer Software
Paython, log, Docker, Ansible, Site Reliability Engineering, Git, Infrastructure as code (IaC), Terraform, Microsoft Azure, GitHub, Technical Support, Microsoft SQL Server, Elasticsearch, AppDynamics, Amazon Web Services (AWS), Kubernetes, Docker Products, Cascading Style Sheets (CSS), HTML5, Flask
Experience

Site Reliability Engineer
Buffalo, New York, United States
Working as a Site Reliability Engineer (SRE) supporting mission-critical banking applications with a strong focus on system reliability, availability, performance, and automation. Key Responsibilities & Achievements: -Monitored and maintained highly available production systems, ensuring uptime and performance in a regulated financial environment -Implemented and supported incident management and on-call operations, performing root cause analysis (RCA) and driving permanent fixes -Automated operational tasks using shell scripting / Python, reducing manual effort and operational risk -Collaborated with DevOps, Development, and Infrastructure teams to improve deployment stability and system resilience -Managed and optimized monitoring and alerting using tools such as Prometheus, Grafana, Splunk, CloudWatch, or similar -Supported CI/CD pipelines and infrastructure changes to improve release reliability -Ensured compliance with security, audit, and change management standards -Documented runbooks, SOPs, and reliability metrics (SLIs, SLOs, SLAs) Environment / Tools: Linux, AWS/Azure, Kubernetes, Docker, CI/CD, Terraform, Ansible, Monitoring & Logging Tools, ITSM, Git

Site Reliability Engineer
Hyderabad, Telangana, India
- Monitoring Infrastructure Setup: Spearheaded the establishment of comprehensive monitoring solutions for application teams using Grafana, ELK stack, Filebeat, and Metricbeat, enhancing system visibility and reliability. - Alert Configuration: Implemented robust alerting mechanisms for application services, ensuring immediate notifications during downtimes to maintain optimal service availability. - Health Check Management: Conducted daily health checks of critical services, proactively identifying and resolving potential issues to minimize disruptions. - Grafana High Availability: Successfully set up a high availability (HA) configuration for Grafana, integrated with MariaDB, to ensure continuous monitoring service and data redundancy. - Proof of Concept for Grafana Mimir: Executed a Proof of Concept (POC) for Grafana Mimir, demonstrating its capabilities for long-term storage and high scalability, paving the way for future-proof monitoring solutions.

Site Reliability Engineer
Hyderabad, Telangana, India
Worked as a Site Reliability Engineer (SRE) focusing on system stability, performance, automation, and operational excellence across development and production environments. Key Responsibilities & Contributions: -Ensured high availability and reliability of applications by proactively monitoring systems and responding to incidents -Supported production and non-production environments, performing troubleshooting and root cause analysis (RCA) -Automated routine operational tasks using Shell scripting and Python, improving efficiency and reducing manual intervention -Assisted in building and maintaining CI/CD pipelines to support smooth application deployments Implemented and managed monitoring and alerting systems to detect and resolve issues early -Collaborated with development and infrastructure teams to improve system performance, scalability, and reliability -Managed Linux servers, including deployments, patching, and performance tuning -Created and maintained documentation, SOPs, and runbooks for operational processes Tools & Technologies: -Linux, AWS/Azure (basic), Docker, Kubernetes (basic), Jenkins, Git, Shell Scripting, Python, Monitoring & Logging Tools
Education

Computer Science
Fall"24: Data Intensive Computing Intro to Machine Learning Computer Security Analysis of Algorithm Spring"25: Operating System Computer Architecture Statistical Data Mining II Modern Network Concept's Web Analytics and Optimization Techniques Fall"25: DataBase Management Systems (Capstone Project)
Tarun Jangam's Contact Information
Phone
Find the Right Leads
Find Verified Contact Data
What LeadContact does well
Find verified emails, phone numbers, and decision-makers with 98% accuracy.
Find Leads
Find the right people by company, role, industry, location, and more.
925M+ professional profiles

Find Emails
Access verified email addresses for your target contacts.
657M+ emails

Find Phone Numbers
Get cross-validated phone data from multiple top sources.
239M+ phone numbers

More Accurate. Lower Cost.
Find contact data in 1 tool with 98% accuracy
LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.
Great conversations start with the right contact.
It’s time to find yours.


