Site Reliability Engineer

You will own the production environment and improve its performance, reliability, and operability. You will monitor and troubleshoot trading systems and exchange connectivity, build production operations tooling, coordinate changes and incidents, reconcile trades and position breaks, manage operational risk, document procedures, and mentor other technical operations SREs.

Responsibilities

  • Own the production environment
  • Monitor and troubleshoot large-scale trading systems and exchange connectivity
  • Build and maintain the DevOps toolkit
  • Improve scalability and system performance using firm-wide metrics
  • Analyze and troubleshoot complex system problems
  • Coordinate changes and manage incidents
  • Communicate technology changes with traders
  • Reconcile trades and position breaks
  • Assess operational risk of production changes
  • Define and document processes and procedures
  • Provide mentorship and cross-training

Requirements

  • Degree in Computer Science or a related field, or equivalent professional experience
  • At least 5+ years of relevant IT operations experience
  • At least 3+ years of experience with Python and shell scripting
  • Familiarity with C++
  • Linux operating system knowledge
  • Knowledge of network and system configuration
  • Knowledge of kernel internals, scheduling, and performance tuning
  • Networking knowledge including routing, multicast, LLDP, VLAN tagging, and Ethernet
  • Ability to handle shared operational and periodic on-call duties
  • Reliable and predictable availability

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available