Amazon Development Centre (London) Limited
Sr. SDM, Availability Eng, Prime Video
₹ Check with seller / month
✓ Actively Hiring
📍 London
💼 Full Time
🛡️ Verified Listing
⚡ Direct Apply — No Agent
🔒 Your Data is Safe
⭐ Trusted by 5 Lakh+ Jobseekers
Job at a Glance
- Category
- IT Engineer & Developer
- Location
- London, England, United Kingdom
- Salary
- Check with seller
- Job Type
- Full Time
- Company
- Amazon Development Centre (London) Limited
- Status
- Open & Active
Job Description
We are seeking an experienced Senior Software Development Manager to lead our Availability Engineering team within Prime Video. This team is responsible for developing and maintaining our observability platform, incident management systems, and resiliency programs.
Key job responsibilities
Manage a high-performing team of software engineers, program managers, data scientists, and incident responders focused on improving the availability and resilience of Prime Video
Oversee the development and evolution of our observability platform, which enables analysis of logs, traces, and other telemetry at scale to rapidly triage and resolve issues
Implement observability and incident management solutions, including the use of generative AI to assist developers in diagnosis and remediation
Establish and refine processes for effective incident management, including on-call rotations, escalation paths, and post-incident review
Drive initiatives to improve the overall resilience and fault-tolerance of the Prime Video platform
Partner closely with other engineering leaders to ensure availability and reliability goals are met
Hire, develop, and retain top technical talent for the Availability Engineering team
A day in the life
1. Team Management:
Hold 1-on-1 meetings with direct reports to discuss progress, challenges, and development goals
Lead daily/weekly team standups to align on priorities and unblock any issues
Facilitate team planning and retrospective sessions to continuously improve processes
Provide technical and career mentorship to team members
2. Observability Platform Oversight:
Review performance metrics and identify areas for improvement in the observability platform
Collaborate with applied scientists and engineers to enhance the platform's analytics capabilities, including the use of generative AI
Ensure the platform is scaling to meet the growing needs of the Prime Video development teams
Oversee the roadmap and backlog for new observability features and capabilities
3. Incident Management:
Oversee the incident management process, including establishing escalation paths and post-incident review
Analyze incident data to identify recurring issues and drive long-term reliability improvements
Key job responsibilities
Manage a high-performing team of software engineers, program managers, data scientists, and incident responders focused on improving the availability and resilience of Prime Video
Oversee the development and evolution of our observability platform, which enables analysis of logs, traces, and other telemetry at scale to rapidly triage and resolve issues
Implement observability and incident management solutions, including the use of generative AI to assist developers in diagnosis and remediation
Establish and refine processes for effective incident management, including on-call rotations, escalation paths, and post-incident review
Drive initiatives to improve the overall resilience and fault-tolerance of the Prime Video platform
Partner closely with other engineering leaders to ensure availability and reliability goals are met
Hire, develop, and retain top technical talent for the Availability Engineering team
A day in the life
1. Team Management:
Hold 1-on-1 meetings with direct reports to discuss progress, challenges, and development goals
Lead daily/weekly team standups to align on priorities and unblock any issues
Facilitate team planning and retrospective sessions to continuously improve processes
Provide technical and career mentorship to team members
2. Observability Platform Oversight:
Review performance metrics and identify areas for improvement in the observability platform
Collaborate with applied scientists and engineers to enhance the platform's analytics capabilities, including the use of generative AI
Ensure the platform is scaling to meet the growing needs of the Prime Video development teams
Oversee the roadmap and backlog for new observability features and capabilities
3. Incident Management:
Oversee the incident management process, including establishing escalation paths and post-incident review
Analyze incident data to identify recurring issues and drive long-term reliability improvements
Job Safety Alert
Real jobs on Jobsiya are always free. Never pay for an interview and never share bank or OTP details.
Report this job →
Similar Jobs:
Sdm Availability Jobs in London
—
IT Engineer & Developer Jobs Near You