Responsibilities
Observability & Alerting: Design and improve monitoring, logging, tracing, dashboards, and SLI/SLOs for production visibility.
Automation & CI/CD: Build internal tools and delivery pipelines to boost efficiency, eliminate toil, and ensure reliable deployments.
Reliability & System Quality: Drive continuous improvements in system performance, scalability, and security to mitigate customer impact.
Incident Management: Lead production incident response from triage to resolution, turning lessons into durable runbooks and preventatives.
Technical Leadership: Guide major technical initiatives, align cross-team engineering efforts, and manage technical risks.
Mentorship: Elevate SRE team capability through design reviews, paired problem-solving, and incident post-mortems.
Operations & On-Call: Join the on-call rotation while leading capacity planning and disaster recovery readiness.
AI/ML Operationalization: Leverage operational AI/ML tools to forecast capacity risks, detect patterns, and automate system health.
Qualifications
Experience: 8+ years in software/platform engineering or DevOps, including 5+ years dedicated to Site Reliability Engineering.
Infrastructure & Observability: Expertise in cloud platforms (AWS), Kubernetes, IaC, and full-stack observability (tracing, logging, SLI/SLOs). Grafana, Datadog, NewRelic
Automation & Scripting: Proficient in Python, Go, or Bash for building CI/CD pipelines, production tools, and toil-reducing automation.
Incident Leadership: Proven track record in root cause analysis, high-severity incident response, and long-term reliability engineering.
Leadership & Mentorship: Strong communication skills with a history of mentoring engineers, leading cross-functional projects, and setting technical strategy.
Operational AI/ML: Hands-on experience using AI/ML on telemetry data to predict capacity risks, spot anomalies, and optimize system workflows.
Skills Required
- 8+ years hands-on experience in software engineering, cloud infrastructure, platform engineering, DevOps, or related roles, including at least 5 years in SRE or reliability-focused role
- Strong knowledge of distributed systems and hands-on experience operating Kubernetes workloads
- Hands-on experience with cloud infrastructure in AWS or a comparable platform
- Proficiency in Infrastructure as Code
- Experience with monitoring, logging, alerting, distributed tracing, SLIs, and SLOs
- Strong proficiency with Python, Go, Bash, or a similar language and experience building production tooling and automation
- Experience building and maintaining CI/CD pipelines and deployment systems
- Proven ability to lead troubleshooting, incident response, root cause analysis, and long-term reliability improvements
- Proven ability to mentor Site Reliability Engineers and lead cross-team technical initiatives
- Demonstrated experience applying AI and machine learning to operational data to forecast risks and improve reliability
Filevine Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Filevine and has not been reviewed or approved by Filevine.
-
Healthcare Strength — Health coverage is described as covering the major bases (medical, dental, vision) and is often framed as decent quality. In some cases, premiums and copays are portrayed as relatively favorable, suggesting tangible value from the plans.
-
Parental & Family Support — Paid parental leave is positioned as a standard, clearly offered benefit. The presence of parental leave alongside disability coverage signals baseline family-support provisions typical of growth-stage tech employers.
-
Fair & Transparent Compensation — Compensation is sometimes framed as fair or reasonable relative to role expectations, with technical roles in particular appearing closer to market-aligned ranges. This creates pockets where pay is perceived as competitive even if not consistently top-of-market across the company.
Filevine Insights
What We Do
Filevine is case management software built for and inspired by real attorneys. As a fully-featured suite of tools, it comes ready to manage every part of a moving case. Assign tasks, upload files or images, monitor staff productivity, and communicate with your client directly from within their case file. Our software is built on the truth that every law firm functions differently. That’s why Filevine is so customizable. Build new case-type templates, design automatic workflows, and receive customized reports on a schedule that fits your needs. Accessing your information is never a problem, because Filevine is hosted on The Cloud. To ensure security, your law firm’s data is protected through state-of-the-art encryption on redundant servers. All you need to get started is an internet connection and your favorite web browser. Learn more at filevine.com.
Gallery
.png)






