Senior Site Reliability Engineer – AUS Region

Nucleus Security

Remote, Melbourne, Australia Remote Full-time 1 mo ago
Workplace
Remote
Location
Melbourne, Australia
Who can apply
Remote for people based in Australia

Check your CV against this job

Free · No signup · A 0–100 ATS match score and the keywords you're missing.

Check my CV free

$14.99/month, cancel anytime. Already have an account? Log in

About the role

Senior Site Reliability Engineer – AUS Region

Are you looking for more in life than just building another web app? Does upending cyber security resonate with you? We're a growth stage cyber security startup that is paving the way forward for how vulnerability management is run in large enterprise organizations. For our customers, vulnerability management has always been a game of catch up, with limited asset coverage and manual processes. Nucleus’ core goal is to build a fast and scalable platform that solves these problems and many more so that vulnerability   management isn't just possible, it's easy. We're looking for a passionate Senior Site Reliability Engineer to join our growing team of engineers AUS region.

What You Will Do

* Maintain Reliable, Secure, and AI-Assisted Production Operations  Keep production systems highly available, secure, patched, and performant. Use AI-assisted tooling to accelerate troubleshooting, identify risks, analyze incidents, and improve operational response.

* Build and Maintain Kubernetes, Cloud, and DevOps Infrastructure  Own and improve Kubernetes clusters, containerized workloads, Infrastructure as Code, CI/CD pipelines, and cloud infrastructure. Leverage AI-assisted development and automation tools to improve delivery speed, configuration quality, and operational consistency.

* Build Observability and Automation That Reduces Toil  Improve monitoring, alerting, logging, dashboards, and automated remediation to identify issues earlier and reduce repetitive operational work. Apply AI and intelligent automation to correlate signals, surface anomalies, assist with root-cause analysis, and automate common SRE workflows.

Expectations of Your Experience

Required Qualifications:

* 8+ years of experience in Site Reliability Engineering, DevOps, Cloud Engineering, Infrastructure Engineering, or related field.

* Strong hands-on experience with cloud platforms, including AWS, GCP, Azure, and/or OpenShift (OCP).

* Deep experience with Kubernetes, containers, and production container orchestration.

* Experience building and maintaining highly available, scalable, and secure production infrastructure.

* Strong experience with Infrastructure as Code, preferably Terraform, and configuration/automation tools such as Ansible.

* Strong scripting and automation skills using Python, Bash, or similar languages.

* Experience building and maintaining CI/CD pipelines using GitHub, GitLab, Bitbucket, or similar platforms.

* Strong experience with observability and monitoring platforms such as Prometheus, Grafana, Loki, CloudWatch, or equivalent tools.

* Experience with incident response, root-cause analysis, production troubleshooting, and reliability engineering practices.

* Experience using AI-assisted engineering tools to improve infrastructure automation, troubleshooting, documentation, code generation, or operational workflows.

* Ability to identify opportunities where AI and automation can reduce operational toil, improve signal detection, and accelerate incident investigation.

* Strong understanding of Linux, networking, security, cloud architecture, and distributed systems.

* Ability to provide technical leadership, mentor engineers, and help drive a culture of automation, reliability, and continuous improvement.

* Must be located in the AUS region and have citizenship in the country you currently reside.

Minimal Requirements

* Minimum 8 years of experience in SRE, DevOps, Cloud Engineering, Infrastructure Engineering, or a related field.

* Strong hands-on experience with cloud providers like AWS, GCP, Azure, and OpenShift (OCP).

* Strong experience with Kubernetes, Infrastructure as Code (terraform, tofu, CloudFormation), and automation (Python, bash, PowerShell).

* Proven experience supporting highly available production systems, including observability, incident response, troubleshooting, and reliability improvements.

Preferred Qualifications:

* Experience integrating LLMs or AI-enabled tools into engineering or operational workflows.

* Familiarity with AI-assisted log analysis, anomaly detection, incident summarization, or root-cause investigation.

* Experience building internal automation or tooling that combines APIs, scripting, infrastructure data, and AI models.

* Understanding of how to use AI safely in production engineering environments, including data security, access controls, validation, and human review.

Additional Information

At Nucleus we are committed to achieving excellence in our field by combining diversity, collaboration, teamwork, and pride in our work. All qualified applicants will receive consideration for employment without regard to race, sex, color, religion, sexual orientation, gender identity, national origin, protected veteran status, or disability.

Get matched & apply with FindAJobAI

Upload your resume once. We score every job against your profile, tailor your resume and cover letter, and autofill the application.