Job Responsibilities:
• Develop and maintain infrastructure and configuration as code using CloudFormation, Terraform, Ansible, and related automation tools.
• Administer and optimize AWS environments, including core services, networking, security, and architecture for availability, performance, and cost.
• Manage and support Kubernetes clusters and containerized workloads, including configuration, scaling, and upgrades.
• Design, implement, and evolve end‑to‑end monitoring and observability frameworks using tools such as Open Telemetry, Groundcover, CloudWatch, Datadog, Prometheus, New Relic, or similar platforms.
• Create and maintain dashboards, logs, traces, SLIs/SLOs, and automated alerting systems to ensure reliability and rapid detection of anomalies.
• Embed observability, CI/CD best practices, and operational readiness into all stages of the software development lifecycle in partnership with engineering teams.
• Lead or participate in incident response, troubleshooting, and root cause analysis for production incidents, using observability data to drive fast resolution.
• Automate operational tasks, runbooks, and incident remediation workflows to reduce toil and improve service reliability.
• Contribute to risk mitigation, backup, and disaster recovery strategies, including periodic testing and continuous improvement.
• Participate in shared after hours support and project work as needed.
Job Qualification:
• 2-4 years of experience designing, implementing, and maintaining CI/CD pipelines (e.g., Harness, GitHub Actions, ArgoCD or similar tools).
• Hands‑on experience with automation tools and Infrastructure as Code / Configuration as Code (CloudFormation, Terraform, Ansible).
• Strong understanding of Infrastructure as Code and Configuration as Code principles and patterns.
• Solid grasp of the software development lifecycle and modern SRE/DevOps practices.
• AWS administration and architecture experience, including networking, security, IAM, and core services.
• Experience operating Kubernetes clusters (EKS or other distributions) and containerized workloads.
• Deep experience with monitoring and observability tools such as OpenTelemetry, Groundcover, CloudWatch, Datadog, Prometheus, New Relic, or equivalent, including metrics, logs, and traces.
• Ability to define and track SLIs/SLOs and use them to guide reliability improvements.
• Proficiency in Linux administration, including system configuration, troubleshooting, and performance tuning.
• Programming/scripting skills in at least one language such as Python, Go, or Rust for automation, tooling, and observability integrations.
• Solid understanding of networking, load balancing, and performance tuning.
• Experience troubleshooting complex distributed systems, supporting incident response, and driving root cause analysis.
• Familiarity with risk mitigation, backup, and disaster recovery concepts.
Preferred
• Experience building unified observability platforms or standardized dashboards for multiple services/teams.
• Experience with GitOps workflows and tools for declarative infrastructure and application delivery.
• Background in incident command and post‑mortem frameworks.
• Experience integrating observability and reliability practices into microservices and/or serverless architectures.
• Experience integrating testing, security and compliance checks into CI/CD pipelines.
Skills Required
- 2-4 years designing, implementing, and maintaining CI/CD pipelines (Harness, GitHub Actions, ArgoCD or similar)
- Hands-on experience with CloudFormation, Terraform, Ansible (Infrastructure as Code / Configuration as Code)
- Strong understanding of Infrastructure as Code and Configuration as Code principles
- AWS administration and architecture experience, including networking, security, IAM, and core services
- Experience operating Kubernetes clusters (EKS or other distributions) and containerized workloads
- Deep experience with monitoring and observability tools (OpenTelemetry, Groundcover, CloudWatch, Datadog, Prometheus, New Relic or equivalent)
- Ability to define and track SLIs/SLOs and use them to guide reliability improvements
- Proficiency in Linux administration, system configuration, troubleshooting, and performance tuning
- Programming/scripting skills in at least one language such as Python, Go, or Rust for automation and tooling
- Solid understanding of networking, load balancing, and performance tuning
- Experience troubleshooting complex distributed systems, incident response, and root cause analysis
- Familiarity with risk mitigation, backup, and disaster recovery concepts
- Participation in shared after-hours support and on-call rotations as needed
- Experience building unified observability platforms or standardized dashboards for multiple services/teams
- Experience with GitOps workflows and tools for declarative infrastructure and application delivery
- Background in incident command and post-mortem frameworks
- Experience integrating observability and reliability practices into microservices and/or serverless architectures
- Experience integrating testing, security, and compliance checks into CI/CD pipelines
What We Do
Model N enables life sciences and high tech companies to drive growth and market share, minimizing revenue leakage throughout the revenue lifecycle. With deep industry expertise and solutions purpose-built for these industries, Model N delivers comprehensive visibility, insight and control over the complexities of commercial operations and compliance. Our integrated cloud solution is proven to automate pricing, incentive and contract decisions to scale business profitably and grow revenue. Model N is trusted across more than 120 countries by the world’s leading pharmaceutical, medical technology, semiconductor, and high tech companies, including Johnson & Johnson, AstraZeneca, Stryker, Seagate Technology, Broadcom and Microchip Technology. For more information, visit www.modeln.com.






