Staff Software Engineer, Site Reliability Engineering, Data Indexing
Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. SRE ensures that Google Cloud's services—both our internally critical and our externally-visible systems—have reliability, uptime appropriate to customer's needs and a fast rate of improvement. Additionally SRE’s will keep an ever-watchful eye on our systems capacity and performance.
Our team focuses on building, managing and evolving the indexing pipelines that enable search in many of Google's user-facing products, including Web Search, News, YouTube, Shopping and Lens. New areas of growth include indexing for embeddings/ML based retrieval to enable novel search use cases and indexing data sets and indexing data sets for model training/grounding.
This is a critical role to the team with strong technical leadership in system reliability, maintainability, and operational excellence.
Behind everything our users see online is the architecture built by the Technical Infrastructure team to keep it running. From developing and maintaining our data centers to building the next generation of Google platforms, we make Google's product portfolio possible. We're proud to be our engineers' engineers and love voiding warranties by taking things apart so we can rebuild them. We keep our networks up and running, ensuring our users have the best and fastest experience possible.
Individual pay is determined by factors including job-related skills, experience, and relevant education or training.US: $207000 - $300000 (USD) + 20% bonus target + equity + benefits
Learn more about benefits at Google.
- Develop the next generation of Indexing Engine platform, Web Data Service (for rapid prototyping and product-ionization of new slice-of-web verticals) and indexing for embeddings/ML based retrieval.
- Support data journeys built on the aforementioned systems for critical Google products, like Search, YouTube, Shopping and Lens.
- Lead key projects to ensure and improve reliability of the pipeline we supported across Data Indexing domain.
- Collaborate and manage stakeholders and development counterparts focusing on reliability, efficiency, and scalability.
- Perform effective production response, thorough follow ups, and comprehensive communications with users and clients.
Minimum qualifications:
- Bachelor’s degree in Computer Science, a related field, or equivalent practical experience.
- 8 years of experience with software development in one or more programming languages.
- 3 years of experience leading projects.
- 3 years of experience designing, analyzing, and troubleshooting distributed systems.
- 2 years of experience with distributed processing.
Preferred qualifications:
- Master's degree in Computer Science or Engineering, or a related field.
- Experience in troubleshooting and debugging of distributed systems.
- Experience with scaffolding based backend development.
- Experience with large-scale databases.
- Proficiency with programming in Golang or C++.