Site Reliability Engineer – Data Infrastructure & Distributed Systems - HFT - Sydney
Westbury Partners Sydney, AustralieSite Reliability Engineer – Data Infrastructure & Distributed Systems - HFT - Sydney
Join a high-impact SRE team operating large-scale data infrastructure, distributed systems and self-managed platforms, combining Linux expertise, automation, incident response and engineering to keep critical systems reliable.
What You'll Do:
Operate and evolve large-scale data infrastructure powering demanding users and workloads. You’ll work across distributed systems, storage, processing platforms and in-house technology, with reliability at the heart of everything you do.
Your responsibilities will include:
- Operate and monitor distributed data platforms, including Kafka, HDFS, Dremio and in-house pipelines.
- Build automation and CI/CD solutions for rapid infrastructure and software deployment.
- Respond to incidents, participate in on-call rotations and engineer lasting improvements that prevent recurring failures.
- Diagnose complex Linux, networking, storage, performance and infrastructure issues.
- Plan and execute upgrades, migrations, capacity improvements and technology rollouts.
- Work directly with traders, researchers and developers to troubleshoot problems and recommend effective data solutions.
- Develop infrastructure tooling and automation using Python and configuration-management technologies.
- Evaluate emerging technologies and introduce improvements to a continuously evolving data platform.
- Collaborate with systems and network engineers and colleagues across international offices.
Why Join Us:
- Own technology at serious scale: Work with multi-petabyte data infrastructure and millions of queries each day.
- Solve real engineering problems: Manage infrastructure where reliability, capacity and failure recovery genuinely matter.
- Build deep technical expertise: Develop hands-on experience with distributed systems, Linux, Kubernetes, Kafka, HDFS and automation.
- Learn from experienced engineers: Join a highly skilled team with structured development and dedicated technical ramp-up.
- Make a visible impact: Your engineering decisions directly influence the reliability and performance of critical data platforms.
- Work with demanding users: Partner directly with technical teams who need fast, practical solutions.
About You:
You’re an engineer who thinks like an SRE: you understand that keeping systems running is only the beginning. You investigate failures, understand their causes and build solutions that make the next incident less likely.
You’ll ideally bring:
- Hands-on production operations experience with self-managed infrastructure.
- Strong Linux fundamentals, including processes, filesystems, networking, memory and disk behaviour.
- Experience operating at least one complex infrastructure platform deeply, particularly Kafka, HDFS or Kubernetes.
- Experience with monitoring, alerting, incident response and production troubleshooting.
- A track record of infrastructure automation and CI/CD.
- Python experience, particularly for infrastructure tooling, deployment automation or operational workflows.
- Experience with configuration management or infrastructure-as-code such as Ansible, Puppet or Terraform.
- An appetite for understanding how complex systems work beneath the surface.
- Strong communication skills and the confidence to work directly with technical users.
- Curiosity, ownership and a willingness to develop breadth across a sophisticated data ecosystem.