Senior Site Reliability Engineer
Summary
Senior SRE at Megaport (a global Network-as-a-Service leader) owning production reliability, incident response, and automation across Kubernetes, AWS, and Terraform-managed infrastructure for a globally distributed platform team.
About Megaport
Our Team Culture
The Role
What You Will Be Doing
- Improving production reliability and system resilience within an SRE scoped team
- Championing high standards of work and industry best practices
- Communicating with teams and stakeholders at all stages
- Bringing fresh ideas to the table and encouraging others
- Diving into complex technical problems with a can-do attitude
- Working across numerous technologies in a fast-changing industry
- Participating in on-call rotation, incident response, and blameless post-incident reviews
- Writing code, handling alerts, improving solutions, and supporting others
- Playing a crucial role in the success of your company and team
What We Are Looking For
- 5+ years administering Linux systems and related infrastructure in production environments
- A collaborative SRE mindset, with familiarity around SLIs/SLOs/SLAs, error budgets, blast radius, and blameless postmortems
- A focus on automation, reducing toil, and preventing problem recurrence
- A track record of writing runbooks that work for the broader team, not just yourself
- Strong Kubernetes and broader ecosystem fundamentals
- Cloud infrastructure experience; AWS strongly preferred and bare-metal is a bonus
- Strong tool development - Bash, plus either Python or Go preferred, or similar
- Infrastructure-as-code tooling experience - Terraform preferred
- CI/CD and version control, GitHub preferred
- Database experience - one of Postgres, Cassandra, or ClickHouse preferred
- Experience operating a production observability stack (metrics, logs, traces), with an eye for signal over noise
- Comfortable working on live production infrastructure, with strong troubleshooting instincts and ownership of incident response
- A history of continual professional development
- A self-directed style suited to an async, globally distributed team, and comfortable picking up adjacent work when the situation calls for it
What We Offer
- Flexible working environment – a remote-first culture with coworking options available.
- Generous leave plans – including 4 weeks of paid annual leave, parental leave, birthday leave, and a purchased annual leave program.
- Health and wellness support – through a wellness allowance and employee wellbeing initiatives.
- Comprehensive learning support – generous study and training allowance plus 5 days of paid study leave
- Creative, modern workspaces – designed to inspire when you're not working remotely
- Motivated, inclusive team – work alongside industry experts and fresh talent
- Recognition programs – celebrate achievements with our Legend and Kudos awards