Site Reliability Engineer - Algorithmic Trading
You will provide first-line support for trading, testing, and research environments, monitor systems, automate observability and anomaly detection, maintain the test trading environment, standardize CI/CD and operational processes, improve deployment and release verification, perform stress tests, verify SLAs, and collaborate across engineering, trading, networking, and infrastructure functions.
Responsibilities
- Provide 24/7 first-line support for trading, testing, and research environments
- Monitor systems and identify issues before they impact trading
- Define and automate observability around application SLAs
- Build anomaly detectors that surface problems early
- Maintain the test trading environment
- Automate routine interventions for the test platform
- Standardize CI/CD and operational processes
- Improve deployment scripts and release verification processes
- Perform stress tests and verify SLAs under load
- Collaborate with traders, engineers, system administrators, data centers, networking, and shared services
Requirements
- Bachelor's degree in computer science or equivalent practical background
- 3+ years of hands-on software development or SRE experience
- Experience with logging, metrics, and tracing tools
- Development experience with Bash, Python, or Ruby
- Knowledge of network communications and the Linux TCP-IP stack
- Knowledge of multicast networking and network protocol interactions
- Diagnostic capabilities from the application layer through networking and low-level hardware
- Experience with network capture and time synchronization
- Experience with AI agents