ML Compute SRE Lead: Scale, Uptime & Automation
Posted 11 days 5 hours ago by Google Inc.
Permanent
Full Time
Other
London, City Of Westminster, United Kingdom, NW1 4
Job Description
Google London is seeking a Systems Engineering Manager for Site Reliability Engineering in ML Compute. You will lead a multi-disciplinary team, own uptime, and shape reliability strategy for large-scale services.
You will mentor engineers, drive end-to-end availability, and collaborate with cross-functional partners across hardware and software domains. This role demands deep technical expertise and strong stakeholder influence.