The Role
Lead application observability discovery and implementation for a visual effects studio. Document application teams’ monitoring, logging, and APM practices; present findings and timelines; align application monitoring with existing SRE incident-response processes; recommend platforms and estimate costs; and support proof-of-concept or implementation work.
Summary Generated by Built In
Senior SRE – Application Performance Monitoring (Contract)About the opportunityGPL Technologies is looking for a senior Site Reliability Engineer for a 3–6 month remote contract with a large, well-known visual effects studio, with a possible extension. You'll lead the next phase of the studio's observability program: extending a proven incident-response process from infrastructure to production applications and services.
The infrastructure phase is complete. Your job is to learn how application teams monitor their services today, then recommend how to streamline monitoring and fold application incidents into the existing response workflow.What you'll do
The infrastructure phase is complete. Your job is to learn how application teams monitor their services today, then recommend how to streamline monitoring and fold application incidents into the existing response workflow.What you'll do
- Meet with each application team to document which monitoring, logging and APM tools they use and how they use them.
- Present discovery findings and timeline updates to stakeholders in weekly check-ins.
- Work alongside the studio's infrastructure SRE to plan how application monitoring fits the current incident-response process and culture.
- Recommend solutions for a proof-of-concept phase, whether one platform or several, with estimated platform and service costs.
- Support the proof-of-concept or implementation phase that follows, as scope allows.
- Senior-level experience in Site Reliability Engineering and/or DevOps, with a heavy focus on application performance monitoring, centralized logging and tracing.
- Hands-on work with SaaS or on-prem platforms such as Datadog, Dynatrace, Elastic/ELK, Logz.io, Grafana or Prometheus.
- Experience connecting application alerts to an incident-response process.
- Clear, responsive communication with technical and non-technical teams.
- Friendly but assertive follow-through: you can get information and move actions forward across many teams in a large organization.
- Zabbix and/or PagerDuty experience
- Visual effects, animation or media production background
- Experience running discovery or assessment engagements as a consultant
Skills Required
- Senior-level experience in Site Reliability Engineering and/or DevOps
- Heavy focus on application performance monitoring, centralized logging, and tracing
- Hands-on experience with platforms such as Datadog, Dynatrace, Elastic/ELK, Logz.io, Grafana, or Prometheus
- Experience connecting application alerts to an incident-response process
- Clear, responsive communication with technical and non-technical teams
- Ability to follow through and move actions forward across multiple teams
- Zabbix and/or PagerDuty experience
- Visual effects, animation, or media production background
- Experience running discovery or assessment engagements as a consultant
Am I A Good Fit?
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.
Success! Refresh the page to see how your skills align with this role.
The Company
What We Do
Since our founding in 2003 in Los Angeles, California, we’ve spent the last twenty years building on our expertise as trusted technology advisers and vendors for the media & entertainment, and architecture industries. Our staff of technology specialists and VFX pipeline engineers are known for their innovative approach to problem-solving and trusted for their expertise in architecting client-centric production IT solutions that help your studio achieve its unique technology goals.






