SRE Engineer


Spatial Front, Inc. (SFI), a two-time USAToday Top Workplaces awardee and Washington Top Workplaces honoree, is seeking a SRE Engineer to support our growing team. The SRE Engineer will support the Infrastructure, Production, and Compliance Support (IPCS) team within the enabling rail of a large-scale federal enterprise program. This role is responsible for improving the reliability, availability, performance, observability, and operational resilience of mission-critical systems supporting a complex, multi-environment ecosystem across development, test, training, and production.


The SRE Engineer will help standardize and mature reliability engineering practices across a highly integrated environment that includes PeopleSoft-based enterprise applications, Oracle platforms, shared services, DevSecOps pipelines, and reporting/integration services operating in regulated NIPRNET and SIPRNET contexts. This position works closely with platform engineers, DevOps, release management, cybersecurity, test automation, and product teams to reduce operational toil, strengthen production readiness, improve incident response, and support continuous delivery without compromising stability or compliance.


As a valued member of the SFI team, you will play a critical role in delivering mission-critical capabilities to our Federal Government customers.


Work Environment: On-site


Key Responsibilities:

  • Define, implement, and maintain site reliability engineering practices for mission-critical applications and shared services, with emphasis on uptime, resiliency, recoverability, and operational excellence.
  • Establish and manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets for critical services and environments.
  • Implement and maintain monitoring, alerting, and observability solutions for production systems.
  • Support production and pre-production operations across development, test, training, staging, and production environments.
  • Lead incident response activities, conducting root cause analysis and implementing permanent fixes.
  • Support capacity planning, performance analysis, trend monitoring, and scalability planning for enterprise platforms and services.
  • Create and maintain runbooks, standard operating procedures, incident playbooks, operational dashboards, and knowledge articles.
  • Support high availability, disaster recovery, backup/restore validation, and business continuity activities.
  • Develop and implement automation to reduce manual operational toil and improve system reliability.
  • Contribute to post-deployment validation, smoke testing, rollback readiness, and environment health checks during releases and maintenance windows.
  • Collaborate with teams supporting Oracle/PeopleSoft platforms, integration services, reporting services, and shared enterprise tooling to improve reliability end to end.
  • Collaborate with development teams to improve system reliability through design reviews and reliability engineering practices.
  • Perform capacity planning and performance optimization for production systems.
  • Other duties as assigned.

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available