Principal Site Reliability Engineer

Hybrid
Principal
🇺🇸 United States
Site Reliability Engineer
Technology

We are currently looking for a Principal Site Reliability Engineer to join our growing team. In this role, you will implement and maintain monitoring systems to track the performance and availability of business-critical systems and infrastructure using metrics to identify trends and potential issues. You will also work closely with development teams, operations, and other stakeholders to ensure that new services and features are reliable and scalable.

As a Principal Site Reliability Engineer, your duties and responsibilities will include:

  • Implement and maintain monitoring systems to track the performance and availability of Business-critical systems and infrastructure. Use metrics to identify trends and potential issues.
  • Respond to system outages and performance issues, performing root cause analysis to prevent recurrence
  • Develop scripts and tools to automate repetitive tasks, such as deployment, scaling, and monitoring
  • Work closely with development teams, operations, and other stakeholders to ensure that new services and features are reliable and scalable
  • Work on reducing latency and improving the speed of data transmission across the network
  • Define and measure Service Level Objectives (SLOs) and Service Level Indicators (SLIs) to ensure services meet required performance and availability targets+
  • Conduct postmortems after incidents to identify what went wrong and what can be improved
  • Work with Lead Application owners and internal Change Management to review code changes and support deployments
  • Lead the team of site reliability engineers onshore/offshore, mentor them for support activities required for system reliability
  • Must have ability to communicate and abstract the messaging to multiple target audiences including Sr business & IT leadership, technology, and business teams.

Requirements

WHAT IT TAKES TO CATCH OUR EYE:

  • Master’s degree in computer science, telecommunications, or similar areas, with a minimum of 10 years software engineering experience, including a minimum of 5 years as a site reliability engineer
  • Proven track record of managing mission critical customer facing applications for reliability
  • 5+ years of experience supporting operations and maintenance for cloud-native applications in production that are fault-tolerant, self-healing, scalable and high available
  • Excellent troubleshooting and problem-solving skills, with a keen attention to detail to identify and resolve complex production issues
  • Deep understanding of cloud computing platforms (GCP) and containerization technologies (e.g., Docker, Kubernetes)
  • Solid experience with core Kubernetes concepts such as Pods, Workloads, Services, Ingress/Egress, Deployments, ConfigMaps, HPA, Liveliness Probe, and Secrets
  • Strong knowledge of infrastructure as code tools (e.g., Terraform, Ansible, ArgoCD) and CI/CD pipelines
  • Strong experience working with integration of code quality tool (SonarQube or Checkmarx) with CI/CD pipeline
  • Strong experience with monitoring, logging, and observability tools like, Splunk, GCP log, Dynatrace etc.
  • Ability to work independently and as part of a collaborative team, effectively communicating technical concepts to both technical and non-technical stakeholders
  • Must have proven written and verbal communication skills, including presentations using tools like PowerPoint
  • Must have ability to communicate and abstract the messaging to multiple target audiences including Sr business & IT leadership, technology and business teams

BONUS POINTS FOR:

  • Certifications such as Google Professional Cloud DevOps Engineer or AWS Certified DevOps Engineer

#LI-SS1

 

Brightspeed

Brightspeed

A company reimagining how people live, work, play and connect by providing fast, reliable internet connections and an awesome customer experience in twenty states throughout the Midwest and South.

Telecommunications

Other jobs at Brightspeed

 

 

 

 

 

 

 

 

View all Brightspeed jobs

Notifications about similar jobs

Get notifications to your inbox about new jobs that are similar to this one.

🇺🇸 United States
Site Reliability Engineer

No spam. No ads. Unsubscribe anytime.

Similar jobs