Site Reliability Engineer | Weekend Warrior
You will manage the real-time production trading environment by monitoring and troubleshooting large-scale trading systems and exchange connectivity. You will build reliability tools, improve scalability and performance, coordinate deployments, manage incidents, reconcile trades, assess operational risk, document procedures, and mentor technical operations engineers. You will work a four-day schedule that includes one weekend day.
Responsibilities
- Own the production environment and drive performance, reliability, and operability improvements
- Monitor and troubleshoot large-scale trading systems and exchange connectivity
- Build and maintain site reliability tools for configuration management, process management, deployment, monitoring, data collection, and analysis
- Use firm-wide metrics to improve scalability and system performance
- Coordinate technology changes and deployments with traders, Risk Management, and Operational Trading Support teams
- Analyze and troubleshoot complex system problems
- Reconcile trades and position breaks with the Clearing team
- Assess and manage operational risk when deploying production changes
- Define and document processes and procedures
- Mentor and cross-train technical operations engineers
Requirements
- Degree in Computer Science, a related field, or equivalent professional experience
- 5+ years of relevant experience in IT operations, DevOps, SRE, Linux Systems Engineering, or Network Engineering
- 3+ years of experience with Python and shell scripting
- Linux operating system knowledge
- Networking knowledge including routing, multicast, LLDP, VLANs, and Ethernet
- Ability to handle shared operational and periodic on-call duties
- C++ experience is a plus
Benefits
- Private Medical, Vision and Dental Insurance
- Travel Medical Insurance
- Group Pension Scheme
- Group Life Assurance and Income Protection Schemes
- Paid Parental Leave
- Parking and Commuter Benefits