Role Summary
and a strategic partner to engineering leadership. You don’t just maintain systems — you shape
engineering culture, define the technical standard for how Filevine runs in production, and bridge the
gap between high-level business goals and robust, internet-scale technical execution. You bring a
forward-looking perspective — actively shaping how AI and machine learning drive the future of
reliability practice.
You own the roadmap across two critical SRE domains — Observability & Alerting and Platform
Infrastructure — and are accountable for ensuring the team solves reliability problems permanently
rather than absorbing them as toil. You operate as the senior IC counterpart to the Engineering
Manager: technical correctness lives with you. You partner with the Reliability Architect and engineering
leadership on significant technical decisions, mentor engineers across experience levels, and influence
reliability strategy across the broader organization. Reliability at Filevine protects revenue. You are the
senior technical voice responsible for ensuring that uptime, incident response, and every production
change meet the operational standard the business demands.
This role does not participate in on-call rotation, but you are deeply invested in the engineers who do —
shaping the on-call strategy, tooling, and culture that make production support sustainable and
effective.
Who You Are
• Master of the Craft: You bring deep expertise in distributed systems, cloud infrastructure,
observability, and reliability engineering. You raise the technical standard for every engineer
around you and thrive where the challenges are complex and the stakes are real.
• Technical Leader and Mentor: You are passionate about mentoring engineers and investing in
their growth. You influence technical direction and communicate production risk clearly across
engineering, product, and executive audiences.
• Forward-Thinking & AI/ML Fluent: You bring deep knowledge of AIOps and drive the use of
AI and machine learning in observability, anomaly detection, incident response, automated
remediation, and resource optimization.
• Production-Scale Problem Solver: You turn ambiguous, complex reliability challenges into
durable solutions for systems where availability, performance, and production changes carry
meaningful business impact.
• Software-Minded Builder: You use software, automation, Infrastructure as Code, and platform
capabilities to eliminate toil and make systems safer, more scalable, and easier to operate.
What you will do
and operational excellence.
• Lead the evolution of reliable, scalable, secure, and efficient cloud platforms and distributed
systems.
• Champion SLIs, SLOs, error budgets, capacity planning, operational readiness, and automation
across the service lifecycle.
• Lead the organization through complex production incidents and turn post-incident learning into
permanent engineering improvements.
• Build self-service platform capabilities that reduce toil, improve engineering safety and velocity,
and make every team more capable of owning their own reliability.
• Mentor engineers and serve as a trusted technical authority for long-term reliability and platform
direction.
Qualifications
including 6+ years in SRE and 3+ years leading complex, cross-functional technical initiatives
for distributed production systems.
• Expert-level depth in observability and platform infrastructure, with broad expertise in incident
response, capacity planning, automation, and reliability engineering.
• Advanced experience with a major container-orchestration platform, preferably Kubernetes, and
an observability platform such as New Relic, Datadog, or equivalent.
• Strong software-engineering ability in Python, Go, Bash, or another general-purpose language,
with experience building production tooling, automation, or platform capabilities.
• Proven ability to mentor engineers and communicate technical risk clearly to engineering,
product, and executive audiences.
• Experience in a regulated environment such as FedRAMP, CJIS, HIPAA, SOC 2, or PCI is
strongly preferred.
Skills Required
- 12+ years of experience in software engineering, infrastructure, platform engineering, or SRE
- 6+ years in SRE
- 3+ years leading complex, cross-functional technical initiatives for distributed production systems
- Advanced experience with a major container-orchestration platform (preferably Kubernetes)
- Experience with an observability platform (e.g., New Relic, Datadog, or equivalent)
- Expert-level knowledge of observability, incident response, capacity planning, automation, SLIs/SLOs, and reliability engineering
- Strong software-engineering ability in Python, Go, Bash, or another general-purpose language (building production tooling/automation/platform capabilities)
- Experience designing and delivering self-service platform capabilities and Infrastructure as Code to reduce toil
- Proven ability to mentor engineers and communicate technical risk to engineering, product, and executive audiences
- Experience in a regulated environment (FedRAMP, CJIS, HIPAA, SOC 2, or PCI)
Filevine Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Filevine and has not been reviewed or approved by Filevine.
-
Healthcare Strength — Health coverage is described as covering the major bases (medical, dental, vision) and is often framed as decent quality. In some cases, premiums and copays are portrayed as relatively favorable, suggesting tangible value from the plans.
-
Parental & Family Support — Paid parental leave is positioned as a standard, clearly offered benefit. The presence of parental leave alongside disability coverage signals baseline family-support provisions typical of growth-stage tech employers.
-
Fair & Transparent Compensation — Compensation is sometimes framed as fair or reasonable relative to role expectations, with technical roles in particular appearing closer to market-aligned ranges. This creates pockets where pay is perceived as competitive even if not consistently top-of-market across the company.
Filevine Insights
What We Do
Filevine is case management software built for and inspired by real attorneys. As a fully-featured suite of tools, it comes ready to manage every part of a moving case. Assign tasks, upload files or images, monitor staff productivity, and communicate with your client directly from within their case file. Our software is built on the truth that every law firm functions differently. That’s why Filevine is so customizable. Build new case-type templates, design automatic workflows, and receive customized reports on a schedule that fits your needs. Accessing your information is never a problem, because Filevine is hosted on The Cloud. To ensure security, your law firm’s data is protected through state-of-the-art encryption on redundant servers. All you need to get started is an internet connection and your favorite web browser. Learn more at filevine.com.
Gallery









