Radiant
US Infrastructure & Operations Technical Lead at Radiant · Location…
ExperiencedOn-siteFull-timeNot specified
Job Description
ROLE OVERVIEW As the US Infrastructure & Operations Technical Lead at Radiant, you will serve as the senior technical and operational leader for our growing US-based infrastructure team. This is a hands-on player-manager role that bridges deep technical execution with day-to-day team leadership. You will work closely with the UK Infrastructure Operations Manager during overlapping morning hours (US Eastern time), participating in cross-regional planning, incident reviews, and strategic alignment. In the US afternoon, you will lead and develop the local team, currently composed of three engineers, with a roadmap to grow the team in the future. The role requires a strong technical background in at least one of: HPC compute/storage infrastructure or high-performance networking, combined with the management skills to guide a small but growing team in a fast-paced, AI-native GPU Cloud environment. KEY RESPONSIBILITIES CROSS-REGIONAL COLLABORATION - Work with the UK Infrastructure Operations Manager during overlapping morning hours to align on priorities, incidents, and deployments. - Participate in global planning sessions, capacity reviews, and cross-functional engineering discussions. - Serve as the primary US point of contact for infrastructure operations, escalating and coordinating with UK leadership as needed. TEAM LEADERSHIP & MANAGEMENT - Lead, mentor, and grow a US-based team of 3 infrastructure engineers and grow the team as the business requires to support our US presence. - Set clear team objectives, priorities, and KPIs aligned with platform reliability, delivery velocity, and operational excellence. - Foster a collaborative, accountable, and continuously improving team culture. - Conduct regular 1:1s, performance reviews, and career development conversations. TECHNICAL LEADERSHIP - Act as a hands-on technical lead, contributing directly to infrastructure design, implementation, and troubleshooting. - Champion best practices in Infrastructure as Code (IaC), observability, automation, and incident management. - Lead or support major incident response for US-region infrastructure, including root cause analysis and corrective action. - Drive reliability improvements through SRE practices: SLO/SLI tracking, proactive monitoring, and automation. INFRASTRUCTURE OPERATIONS - Oversee cloud and data centre operations for US-region infrastructure, including HPC/AI hardware deployment and maintenance. - Contribute to network architecture and implementation supporting low-latency, high-throughput workloads. - Ensure robust capacity planning and resource allocation for US infrastructure footprint. - Coordinate 24/7 on-call coverage within the US team and ensure handoff processes with the UK team are seamless. KEY OBJECTIVES - Establish and grow a high-performing US infrastructure team aligned with Radiant’s global operational standards. - Ensure 99.9%+ platform uptime across US-region services. - Enable rapid and predictable infrastructure deployments through automation and operational maturity. - Build effective cross-regional collaboration and follow-the-sun support coverage with the UK team. - Deliver technical excellence in HPC compute/storage or networking for AI/GPU workloads. KEY METRICS - MTTR, MTBF, and overall system uptime for US-region infrastructure. - Achievement of SLOs and tracking of SLIs. - Infrastructure deployment velocity and lead time. - Team growth, retention, and development milestones. - Operational cost efficiency and capacity utilisation. QUALIFICATIONS AND EXPERIENCE EDUCATION - Bachelor’s or Master’s degree in Computer Science, Information Technology, Engineering, or a related field. CERTIFICATIONS (DESIRABLE) - Relevant cloud or infrastructure certifications (AWS, GCP, Azure, CCIE, JNCIS, etc.). - PMP, ITIL, or equivalent project/operations management certification. EXPERIENCE - 6+ years of experience in infrastructure or platform operations. - 2+ years in a technical lead or management role, with direct reports. - Hands-on experience deploying and operating large-scale HPC, AI/GPU, or cloud infrastructure. - Demonstrated experience in at least one of: HPC compute/storage systems or high-performance networking environments. - Background in SRE or DevOps practices including observability, automation, and incident management. TECHNICAL SKILLS - HPC compute/storage: bare-metal server deployment, GPU cluster management, storage fabric (NFS, Lustre, GPFS, or similar). - High-performance networking: InfiniBand, RoCE, high-throughput Ethernet, low-latency network design, LAN and WAN. - SRE and DevOps tooling: Terraform, Ansible, Kubernetes, Prometheus, Grafana, ELK stack. - Scripting and automation: Python, Bash. - Cloud infrastructure management (AWS, GCP, or Azure). SOFT SKILLS - Comfortable operating as both a hands-on technical contributor and a people manager. - Excellent communicator with strong cross-functional and cross-regional collaboration skills. - Proven ability to lead through ambiguity and prioritise effectively in a fast-moving environment. - Analytical and outcomes-focused, with the ability to translate technical complexity for non-technical stakeholders.