Canonical
Senior Site Reliability / GitOps Engineer
Posted
3 weeks ago
Experience
2+ Years
Deadline
Closed
Job Summary
The Senior Site Reliability / GitOps Engineer builds, scale, and automates Canonical's global IT production services and cloud architectures. Core daily duties include developing Python automation scripts, maintaining infrastructure as code templates across public and private clouds, managing Kubernetes cluster orchestration, configuring observability tools (Prometheus, Grafana, and Elasticsearch), handling time-critical technical escalations, and contributing code fixes back to open-source project Upstreams.
Foundational Engineering & Technical Prerequisites
- Educational Credentials: Bachelor's degree or higher, preferably in computer science or a related engineering field.
- Enterprise Linux Administration: Proven hands-on experience administering enterprise Linux servers, with a deep affinity for Linux storage types (such as Ceph or relational databases).
- GitOps & Declarative Code: Comprehensive exposure to managing and deploying cloud infrastructure with code, backed by a modern view of hosting architecture.
- Scripting & Software Development: Practical Python software development experience, specifically contributing to large-scale infrastructure or tool projects.
- Container Orchestration: Proven professional experience configuring, managing, and maintaining production clusters within Kubernetes or similar container orchestration platforms.
- Core Networking Literacy: Practical knowledge of Linux networking concepts, including complex routing tables, subnetting, and firewall rule configurations.
- Communication Matrix: Clear and effective communication skills in English across asynchronous platforms, including email, chat, video calls, and in-person sessions.
- Global Mobility Flex: Full willingness and geographical flexibility to travel internationally 2 to 4 times a year to participate in mandatory team engineering sprints.
Preferred Technical Multipliers
- Ecosystem Connection: A deep personal passion for and familiarity with open-source software, particularly the Ubuntu or Debian distribution environments.
- Full-Stack Troubleshooting Depth: An ability to troubleshoot technical failures across the entire software stack, moving cleanly from the low-level Linux kernel up to web applications.
- Distributed Systems Mentorship: Prior history mentoring remote engineers and leading collaborative design sessions within an asynchronous tech culture.
- Observability System Design: Direct experience designing, implementing, and maintaining automated alerts using Prometheus, Grafana, and Elasticsearch clusters.
Key Responsibilities
1. Embedded Technical Leadership & GitOps Evolution
- Drive the development of advanced automation and GitOps principles within your unit as an embedded team technical lead.
- Collaborate closely with the Information Systems (IS) architect to ensure all proposed infrastructure tools align with long-term company architecture targets.
- Design, build, and maintain services offered to the wider organization as structured, repeatable internal infrastructure products.
2. Infrastructure as Code (IaC) Optimization
- Apply your practical IaC knowledge to constantly increase internal automation levels and optimize deployment workflows.
- Automate software operations to ensure re-usability and configuration consistency across both private clouds and diverse public cloud nodes.
- Account for the intricacies and edge cases of highly distributed systems, building resilient fault domains into your declarative playbooks.
3. Operational Reliability & Observability Governance
- Maintain shared operational responsibility for all of Canonical’s core IT services, networks, and global web infrastructure.
- Set up, manage, and scale corporate monitoring and logging frameworks using Prometheus, Grafana, and Elasticsearch.
- Design and implement smart, noise-free monitoring rules and automated alerts to catch infrastructure issues before they impact end users.
- Drive advanced capacity planning, deep kernel-to-web troubleshooting, and performance investigation tracks.
4. Upstream Open Source Contribution & Escalation Management
- Identify scale-related bugs in Canonical's core products and upstream open-source packages, documenting findings and writing pull requests to patch code flaws.
- Share your technical know-how and structural best practices with team members via interactive design sessions and direct engineering mentorship.
- Handle final engineering responsibility for time-critical, high-priority system escalations to protect enterprise application availability.
Skills Required:
- Computer / Software / It / Data
Quick Actions
Share Vacancy