Member of Technical Staff - Compute Platform
You will build the platform software and infrastructure used to manage and monitor AI workloads. You will develop web interfaces, Python APIs and backend services, real-time debugging tools, distributed training infrastructure in Rust, automation pipelines, cloud resources, container orchestration, and scheduling systems for heterogeneous hardware.
Responsibilities
- Build web interfaces for AI workload management and monitoring
- Develop REST APIs and backend services in Python
- Create real-time monitoring and debugging tools
- Implement resource management and job control features
- Design distributed training infrastructure in Rust
- Build networking and coordination components
- Create Ansible infrastructure automation pipelines
- Manage cloud resources and container orchestration
- Implement scheduling systems for CPU, GPU, and TPU hardware
- Integrate backend features into existing infrastructure
Requirements
- Strong Python backend development with FastAPI and async
- Modern frontend development with TypeScript, React/Next.js, and Tailwind
- Developer tools and dashboard development
- RESTful API design and implementation
- Rust systems programming
- Ansible and Terraform automation
- Kubernetes
- GCP or other cloud platforms
- Prometheus and Grafana
- GPU computing or ML infrastructure experience