Brice I.
Senior DevOps / Site Reliability Engineer @ Paramount+
About
Senior DevOps / Site Reliability Engineer with 8+ years of experience building and operating secure, resilient cloud platforms across AWS, GCP, and Azure. Specializes in Kubernetes (EKS/GKE/AKS), Terraform/Terragrunt infrastructure automation, CI/CD engineering, and GitOps-driven delivery using ArgoCD and progressive rollout strategies. Strong command of observability and reliability practices, including SLO/SLI design, alert tuning, distributed tracing with OpenTelemetry, and incident response discipline focused on prevention and fast recovery. Known for partnering effectively with engineering, security, and product teams to modernize release workflows, enforce least-privilege access, and harden platforms through policy and automation. AWS Certified Solutions Architect – Professional and AWS Certified Machine Learning – Specialty with a consistent track record of delivering stable, compliant systems that support rapid, predictable delivery.
United States
Katy
Information Technology & Services
Infrastructure as code (IaC), Helm Charts, Helm (Software), OCI, Cloud Infrastructure, GITHUB ACTIONS, Docker Products, Site Reliability Engineering, Troubleshooting, Google Kubernetes Engine (GKE), Object-Oriented Programming (OOP), Amazon CloudWatch, AWS Lambda, Google Cloud Platform (GCP), Infrastructure, AWS CloudFormation, Bash, Red Hat Linux, Linux System Administration, Linux Server
Experience

Senior DevOps / Site Reliability Engineer
United States
• Architected and operated production workloads on AWS EKS with multi-AZ resilience, supporting large-scale microservices and meeting 99.99% availability targets. • Designed reusable Terraform modules for VPC, IAM, EKS, RDS, ALB/Ingress, and Route53, cutting provisioning time by 60% and reducing environment drift. • Implemented GitOps delivery using ArgoCD and progressive delivery via Argo Rollouts (blue/green and canary), reducing deployment failures by 40%. • Built and standardized CI/CD pipelines using GitHub Actions and Jenkins, decreasing deployment time by 75% and improving release frequency from weekly to daily. • Established SRE operating model using SLO/SLI definitions and error budgets, reducing Sev-1 incidents by 35% through reliability-focused engineering decisions. • Implemented end-to-end observability using Prometheus, Grafana, Alertmanager, ELK Stack, and CloudWatch, reducing MTTR from 90 minutes to less than 30 minutes. • Tuned alerting and dashboards to improve signal-to-noise, lowering false-positive pages by 30% and improving on-call effectiveness. • Strengthened security controls with least-privilege IAM, Kubernetes RBAC, Vault secrets management, TLS/WAF, and automated scanning with Trivy, improving compliance readiness. • Implemented policy guardrails using OPA patterns for preventative controls, reducing misconfigurations reaching production by 25%. • Led incident response and postmortems (RCA, corrective actions, prevention plans), reducing repeat incidents by 45% through permanent fixes. • Improved traffic management by standardizing Ingress controllers and ALB routing, reducing routing-related production issues by 30%. • Optimized cloud cost via right-sizing, autoscaling, tagging governance, and S3 lifecycle policies, delivering $180K+ annual savings. • Designed and validated HA/DR processes including backup/restore testing and failover drills, meeting defined RTO/RPO targets consistently.

Sr DevOps Engineer
United States
• Automated infrastructure provisioning using Terraform and CloudFormation, accelerating environment builds by 50% and improving repeatability across teams. • Containerized applications using Docker and deployed to EKS/GKE clusters, improving scalability and reducing infrastructure cost by 30%. • Built and maintained CI/CD pipelines in Jenkins, GitLab CI, and Azure DevOps, increasing deployment frequency by 300% while reducing manual release steps. • Standardized deployment packaging using Helm charts and Kustomize overlays, improving environment consistency across dev/test/prod. • Implemented controlled promotion and release workflows (approvals, artifact versioning, rollback strategy), reducing production release risk and downtime. • Integrated monitoring/logging with ELK and CloudWatch, improving time-to-detect by 40% and speeding up troubleshooting. • Reduced pipeline break-fix cycles by hardening build jobs and eliminating recurring failure causes through root-cause remediation and automation. • Improved source control governance with Git branching standards, PR workflows, and review gates, lowering merge conflicts and improving code quality. • Implemented security hardening for Kubernetes using RBAC, secrets management practices, and vulnerability scanning, reducing high/critical findings by 50% before release. • Tuned alerts and dashboards to reduce noise by 30%, aligning monitoring to operational outcomes and service health. • Automated routine operational tasks using Python/Bash, reducing manual work by 40% and improving execution consistency.

System Engineer
Houston, Texas, United States
• Supported enterprise Linux environments (RHEL/Ubuntu/CentOS) for production systems, ensuring stability through disciplined operational support and maintenance. • Performed troubleshooting across OS, services, and connectivity layers, improving platform availability and reducing repeated escalations. • Automated operational checks and routine tasks using Bash and Python, reducing manual workload by 30% and improving consistency. • Supported network and platform services including DNS, load balancing concepts, and service connectivity during incident triage. • Participated in incident response and RCA activities, contributing to a 20% reduction in recurring infrastructure failures through corrective actions. • Assisted with capacity and performance tuning efforts, helping prevent saturation and reliability degradation during peak usage windows. • Supported change and release processes, coordinating maintenance activities to reduce unplanned downtime and improve execution quality. • Helped maintain backup/restore procedures and validation routines, improving recovery reliability and reducing restore failures. • Produced and maintained runbooks, SOPs, and operational documentation in Confluence, improving knowledge transfer and reducing tribal dependency. • Collaborated with cross-functional teams in Agile delivery environments, improving operational alignment with application release timelines.
Brice I.'s Contact Information
Phone
Find the Right Leads
Find Verified Contact Data
What LeadContact does well
Find verified emails, phone numbers, and decision-makers with 98% accuracy.
Find Leads
Find the right people by company, role, industry, location, and more.
925M+ professional profiles

Find Emails
Access verified email addresses for your target contacts.
657M+ emails

Find Phone Numbers
Get cross-validated phone data from multiple top sources.
239M+ phone numbers

More Accurate. Lower Cost.
Find contact data in 1 tool with 98% accuracy
LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.
Great conversations start with the right contact.
It’s time to find yours.




