ReVybe IT Recruitment Limited
Greater London / Global
You have blocked notifications
Oops! You have blocked notifications. Click here for more info
You have blocked notifications, please check your browser settings.
You're currently subscribed to job notifications
Subscribe to notifications
You will no longer receive notifications
Greater London / Global
Site Reliability EngineerCentral London – Hybrid (2/3 days a week in the office)Up to £85,000 + BenefitsBuild, Scale & Improve the Reliability of a Fast-Growing SaaS PlatformWe're partnering with a fast-growing SaaS company that's going through an exciting period of growth and investing heavily in its engineering and platform capabilities.They're looking for an experienced Site Reliability Engineer (SRE) to join the team and play a key role in building highly reliable, scalable, and observable infrastructure.This is a hands-on role focused on AWS, Kubernetes, Terraform, observability, monitoring, and automation, working closely with software engineering teams to improve platform reliability and developer experience.You'll have genuine ownership and the opportunity to influence how the platform evolves as the business continues to scale.What You'll Be DoingDesign, build, and maintain highly available and scalable AWS infrastructureManage and optimise Kubernetes environments and containerised workloadsBuild and maintain infrastructure using Terraform and Infrastructure as Code principlesDevelop and optimise CI/CD pipelines using GitHub ActionsBuild and improve comprehensive monitoring and observability across the platformImplement and maintain effective logging, metrics, tracing, alerting, and dashboardsDefine and improve SLIs, SLOs, and reliability metricsProactively identify and resolve performance, availability, and reliability issuesLead and contribute to incident response, troubleshooting, and root cause analysisAutomate operational processes and eliminate repetitive manual tasksWork closely with software engineers to improve deployment processes, system reliability, and developer experienceHelp improve platform resilience, scalability, and disaster recovery capabilitiesContribute to capacity planning and performance optimisation as the platform scalesEstablish and champion SRE best practices across the wider engineering functionWhat We're Looking ForProven commercial experience working as an SRE, DevOps Engineer, Platform Engineer, or similarStrong hands-on experience with AWSStrong experience working with KubernetesExcellent experience with Terraform and Infrastructure as CodeStrong experience building and managing GitHub Actions CI/CD pipelinesSolid experience with monitoring and observability toolingStrong understanding of metrics, logging, tracing, alerting, and system healthExperience troubleshooting complex production environmentsUnderstanding of SLIs, SLOs, SLAs, and error budgetsExperience with incident management and root cause analysisGood understanding of cloud networking, security, and infrastructure fundamentalsStrong scripting/automation skillsA strong understanding of reliability, scalability, performance, and availabilityExcellent communication skills and the ability to work closely with software engineering teamsA proactive mindset and genuine passion for automation and continuous improvementDon't Tick Every Box?That's okay.The company is open to speaking with engineers who may not have experience across every technology listed above.If you have strong foundations in AWS, Kubernetes, Terraform, and cloud infrastructure, along with a genuine interest in reliability and observability, we'd still love to hear from you.Why Join?Join a fast-growing SaaS company at an exciting stage of its journeyWork with a modern AWS and Kubernetes environmentTake ownership of reliability, automation, and platform performanceWork with modern observability and monitoring technologiesHave genuine influence over engineering and platform decisionsWork closely with talented software engineering teamsClear opportunities to progress as the business continues to scaleHelp shape and mature the company's SRE practicesHybrid working from Central London, 2/3 days per weekIf you're an experienced SRE, Platform Engineer or DevOps Engineer who enjoys solving complex reliability challenges and wants to have a real impact within a rapidly growing SaaS business, we'd love to hear from you.Site Reliability EngineerCentral London | Hybrid – 2/3 days per weekUp to £85,000 + Bonus + BenefitsAWS | Kubernetes | Terraform | GitHub Actions | SRE | Observability | Monitoring | CI/CD | Infrastructure as Code | Reliability | Automation | Cloud
#J-18808-Ljbffr
Greater London / Global
Greater London / Global
Greater London / Global
Greater London / Global
Greater London / Global
Greater London / Global