- Collaborate with development teams to design and implement monitoring, alerting, dashboards, and APM instrumentation across applications and services.
- Lead the implementation, configuration, and optimization of Application Performance Monitoring (APM) solutions.
- Apply observability best practices using tools such as Azure Monitor, Application Insights, New Relic, and Log Analytics (KQL).
- Enable code-level instrumentation, distributed tracing, and structured logging to improve application visibility and reliability.
- Design and maintain application-level monitoring dashboards and operational health metrics.
- Define and implement Service Level Indicators (SLIs), Service Level Objectives (SLOs), and effective alerting strategies based on latency, error rates, traffic, and resource saturation.
- Continuously improve monitoring and alerting mechanisms through production insights and incident learnings.
- Participate in production readiness reviews, identifying operational risks, observability gaps, and potential failure scenarios before deployment.
- Support incident analysis and post-incident improvements through enhanced telemetry and monitoring practices.
- Partner with engineering teams to ensure applications are reliable, scalable, and production-ready.
- Strong experience supporting and operating applications in Microsoft Azure IaaS environments.
- Hands-on experience with application observability, monitoring, and reliability engineering practices.
- Mandatory experience with DBT, Databricks, and SQL (minimum 1 year of experience).
- Experience implementing and managing APM solutions such as Application Insights, New Relic, or similar platforms.
- Experience designing dashboards and monitoring solutions using Azure Monitor, Application Insights, and Log Analytics (KQL).
- Familiarity with CI/CD environments including Azure DevOps and GitHub Actions.
- Solid understanding of cloud-native architectures and distributed application systems.
- Practical SRE mindset with experience in incident analysis, root cause investigation, and proactive problem prevention.
- Strong verbal and written English communication skills, with the ability to collaborate effectively with global teams.
- Experience with scripting and automation using PowerShell and/or Bash.
- Knowledge of scalability, availability, and resilience patterns in modern cloud environments.
- Experience driving production readiness and operational excellence initiatives.
- Exposure to reliability engineering best practices in enterprise-scale environments.
Skills Required
- 5+ years of experience in Site Reliability Engineering, Cloud Operations, or related roles
- At least 1 year of hands-on experience with DBT, Databricks, and SQL
- Strong experience supporting and operating applications in Microsoft Azure IaaS environments
- Hands-on experience with application observability, monitoring, and reliability engineering practices
- Experience implementing and managing APM solutions such as Application Insights, New Relic, or similar platforms
- Experience designing dashboards and monitoring solutions using Azure Monitor, Application Insights, and Log Analytics KQL
- Familiarity with CI/CD environments including Azure DevOps and GitHub Actions
- Understanding of cloud-native architectures and distributed application systems
- Experience with incident analysis, root-cause investigation, and proactive problem prevention
- Strong verbal and written English communication skills
- Experience with scripting and automation using PowerShell and/or Bash
- Knowledge of scalability, availability, and resilience patterns in modern cloud environments
- Experience driving production readiness and operational excellence initiatives
- Exposure to reliability engineering best practices in enterprise-scale environments
Encora Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Encora and has not been reviewed or approved by Encora.
-
Healthcare Strength — Health coverage is described as employer-provided in multiple locations, with private plans and family coverage highlighted in Spain and Mexico. Medical insurance quality is presented as a recurring bright spot alongside standard coverage.
-
Leave & Time Off Breadth — Time off includes paid holidays and PTO, with regional materials indicating additional leave provisions in certain countries. Leave is generally portrayed as conventional to generous depending on location.
-
Flexible Benefits — Work-from-home flexibility is frequently highlighted as a plus, though it varies by role and client needs. Remote and hybrid options are positioned as part of the overall package.
Encora Insights
What We Do
Headquartered in Santa Clara, California, and backed by renowned private equity firms Advent International and Warburg Pincus, Encora is the preferred technology modernization and innovation partner to some of the world’s leading enterprise companies. It provides award-winning digital engineering services including Product Engineering & Development, Cloud Services, Quality Engineering, DevSecOps, Data & Analytics, Digital Experience, Cybersecurity, and AI & LLM Engineering. Encora's deep cluster vertical capabilities extend across diverse industries, including HiTech, Healthcare & Life Sciences, Retail & CPG, Energy & Utilities, Banking Financial Services & Insurance, Travel, Hospitality & Logistics, Telecom & Media, Automotive, and other specialized industries. With over 9,000 associates in 47+ offices and delivery centers across the U.S., Canada, Latin America, Europe, India, and Southeast Asia, Encora delivers nearshore agility to clients anywhere in the world, coupled with expertise at scale in India. Encora’s Cloud-first, Data-first, AI-first approach enables clients to create differentiated enterprise value through technology








