Finthrive→
Site Reliability Engineer at Finthrive in Hybrid - Gurugram
ExperiencedHybridFull-timeHybrid - Gurugram
Skills
GitHub CopilotLog Analytics WorkspaceTerraformApp ServicesAzure Application Gateway (AGW)ARM TemplatesAzure Front DoorSite 24x7
Job Description
Role & responsibilities
Site Reliability Engineer / Cloud Engineer
SRE & Reliability Engineering
- Managed production environments ensuring high availability and reliability of cloud-hosted applications
- Led incident response, performed deep root cause analysis, and implemented preventive measures to reduce recurrence
- Improved system resilience through proactive monitoring and performance tuning strategies
Azure Application & Platform Engineering
- Designed and supported application architectures using:
- Azure App Services and App Service Plans
- Azure App Service Environment v3 (ASEv3) for isolated, high-scale workloads
- Azure Application Gateway (WAF-enabled) for L7 traffic management
- Azure Front Door for global traffic routing and failover
- Implemented secure and scalable cloud networking patterns, optimizing latency and throughput
Automation & Toil Reduction
- Identified repetitive operational tasks and reduced manual effort through automation-first solutions
- Developed automation using:
- Terraform / Bicep / ARM templates
- Azure Automation (Hybrid Workers)
- Azure Functions for event-driven workflows
- Leveraged AI-assisted tools (GitHub Copilot, Copilot) to accelerate scripting and automation development, while ensuring strict validation for enterprise use
Observability & Monitoring
- Built and enhanced observability using:
- Azure Monitor, Application Insights, Log Analytics
- Created KQL-based queries and dashboards for proactive issue detection
- Reduced false alerts by optimizing alert thresholds and improving signal quality
Performance & System Optimization
- Analyzed application performance across distributed systems to identify bottlenecks
- Implemented improvements through:
- Scaling strategies (horizontal & vertical)
- Network optimization (AGW / Front Door tuning)
- Backend service improvements
Collaboration & Engineering Enablement
- Partnered with SRE, CloudOps, and development teams to design resilient systems
- Contributed to runbooks, documentation, and operational standards
- Enabled engineering teams by improving platform reliability and deployment pipelines
Key Achievements
- Reduced manual operational effort by X% through automation initiatives
- Improved system availability to 99.X% by strengthening monitoring and failure handling mechanisms
- Decreased incident resolution time by X% via enhanced observability and streamlined runbooks
- Optimized application performance using Front Door and AGW tuning, reducing latency by X%
Preferred candidate profile
Cloud & Platform Engineering
- Microsoft Azure (Preferred)
- Understanding and experience in developing Azure function Apps, Azure logic Apps
- Understanding of event triggers, event hub, service bus.
- Azure Landing zones, Azure Cloud Adoption Framework, Azure Well Architectured Framework
- Application Hosting: App Services, App Service Plans, ASEv3
- Networking: Azure Application Gateway (AGW), Azure Front Door, VNet, NSGs, Load Balancing
- Cloud Architecture: High Availability, Fault Tolerance, Scalability Patterns
Incident Management and RCA
- Incident Management, P1 troubleshooting, Change Management
- Experienced in leading RCA and representing on the weekly call
- SLA / SLO / Error Budget concepts
- System Performance Optimization & Capacity Planning
- Toil Reduction through Automation
Infrastructure as Code & Automation
- Terraform, Azure Bicep, ARM Templates
- Azure Automation (Hybrid Workers)
- Azure Functions (Serverless automation)
- API-based automation and orchestration
Observability & Monitoring
- Azure Monitor, Log Analytics Workspace, Grafana, Site 24x7 (or similar SaaS based synthetic monitoring tool)
- Application Insights
- KQL (Kusto Query Language)
- Alert tuning and signal-to-noise optimization
AI-Enabled Productivity (Not as Skill)
- Leveraging GitHub Copilot / Microsoft Copilot for:
- Code acceleration and script generation
- Automation development support
- Troubleshooting and log analysis assistance
- Proven track record of workforce optimization leveraging AI tools.
- Applying validation frameworks to ensure secure, accurate, and production-grade outputs
DevOps & Integration
- CI/CD using Azure DevOps
- Deep understanding on version control
- API integrations (REST, Postman, SoapUI)
- Source control and release management
- Bachelors Degree in Computer Science / Engineering or related field
Preferred/Additional Experience
- Experience with microservices and distributed architectures
- Exposure to low-code automation platforms
- Working knowledge of AWS cloud services
Preferred/Additional Certifications
- AZ-104 Azure Administrator
- AZ-700 Designing and Implementing Microsoft Azure Networking Solutions
- AZ-400 Microsoft Certified: DevOps Engineer Expert