Jobgether SRL→
Site Reliablity Engineer at Jobgether SRL in Birmingham, England, GB
Skills
Job Description
Accountabilities:
- Deploy, manage, and scale distributed systems across multi-region cloud environments with a focus on high availability and performance.
- Design, maintain, and optimize Kubernetes-based infrastructure for large-scale production workloads.
- Build and manage Helm charts to enable consistent, automated, and repeatable deployments.
- Implement and maintain infrastructure-as-code solutions using tools such as Terraform and related automation frameworks.
- Monitor system health using observability tools such as Grafana, Prometheus, and logging stacks, ensuring proactive issue detection and resolution.
- Collaborate with development teams to improve CI/CD pipelines, deployment reliability, and production readiness.
- Lead incident response, troubleshooting, and root cause analysis for production issues.
- Develop and maintain runbooks, operational documentation, and best practices for system reliability.
- Drive continuous improvements in scalability, performance, automation, and cloud infrastructure efficiency.
- Support multi-region deployment strategies and global infrastructure optimization initiatives.
Requirements:
- 4–5 years of experience in Site Reliability Engineering, DevOps, or similar infrastructure-focused roles.
- Strong hands-on experience with Kubernetes in production environments.
- Experience with cloud platforms, especially AWS (EKS, VPC, S3, IAM, ECR, and related services).
- Solid understanding of infrastructure-as-code tools such as Terraform and Git-based workflows.
- Experience building and managing CI/CD pipelines and automation systems.
- Proficiency in Helm charts and containerized deployment strategies.
- Strong scripting skills in Bash and familiarity with at least one programming language (Go or Python preferred).
- Experience working with distributed systems, microservices architectures, and cloud-native ecosystems.
- Strong debugging, troubleshooting, and problem-solving skills in production environments.
- Experience with observability tools such as Grafana, Prometheus, Loki, or similar stacks.
Benefits:
- Opportunity to work on global-scale, mission-critical cybersecurity infrastructure
- Exposure to advanced cloud-native technologies and distributed systems architecture
- Flexible and remote-friendly work environment
- Strong focus on engineering excellence, ownership, and continuous learning
- Global collaboration with high-performing engineering teams
- Career growth in SRE, platform engineering, and cloud infrastructure domains
- Inclusive, innovation-driven culture focused on impact and reliability.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1