Senior Cloud Engineer

Posted 4 Days Ago
Be an Early Applicant
San Francisco, CA, USA
In-Office
122K-204K Annually
Senior level
Digital Media
The Role
Designs and scales highly available cloud, edge, and routing infrastructure across on-premises datacenters, AWS, and GCP. Manages Kubernetes, Terraform, Chef, Cloudflare, CI/CD, MySQL replication, disaster recovery, monitoring, security, and automation. Participates in on-call incident response, root-cause analysis, operational planning, and mentoring. Develops internal tools and improves reliability through SLOs, instrumentation, performance tuning, and centralized logging.
Summary Generated by Built In
About this Role

Fandom is growing! We’re looking for a Senior Cloud Engineer to help evolve and support the infrastructure that powers our platform for over 300 million fans around the world. This is a hands-on role focused on building reliable, scalable systems in a Linux and Kubernetes-based environment.

As part of the TechOps team, you’ll report to the Manager of TechOps and work closely with developers, product engineers, and other infrastructure teams. You’ll contribute to our CI/CD, monitoring, automation, and cloud efforts — helping ensure Fandom’s platform remains fast, stable, and secure as we grow.

This is a great opportunity for someone who enjoys solving complex infrastructure challenges, improving deployment systems, and enabling engineering teams to move faster and safer.

You Will...
  • Design, architect, and scale high-availability routing architectures to seamlessly balance and secure global user traffic across a hybrid footprint of on-premise datacenters, AWS, and GCP.
  • Manage and automate cloud and edge infrastructure as code (IaC) using Terraform and Chef, ensuring consistent configurations for Kubernetes, global Cloudflare services, and CI/CD pipelines.
  • Maintain and optimize large-scale production environments, orchestrating high-availability MySQL replication, automated failover, robust disaster recovery, and system monitoring.
  • Develop internal tools for operational automation, system/data backups, performance tuning, and comprehensive security monitoring.
  • Drive operational excellence by leading planning meetings, retrospectives, and RCAs, while continuously evaluating systems against industry best practices.
  • Participate in an on-call rotation to maintain production stability, handle incident responses, and collaborate with/mentor cross-functional engineering teams.
You Have...
  • 5+ years of experience in Technical/Network Operations, DevOps, or SRE roles managing large-scale production platforms (e.g., 10M+ monthly active users).
  • Deep proficiency in Linux systems administration, networking protocols (TCP/IP, routing), secure systems practices, and scripting/programming (Go, Python, or Bash).
  • Proven hands-on experience with core infrastructure tech: Kubernetes/container orchestration, CI/CD pipelines (GitHub Actions, Jenkins), and monitoring/reliability systems (e.g., Prometheus).
  • Practical experience managing production MySQL database environments, including deep familiarity with replication topologies, failover mechanisms, and performance tuning.
  • Demonstrated capability using generative AI tools (e.g., Gemini, NotebookLM) to enhance productivity, paired with the ability to critically audit and verify outputs for accuracy, security, and context.
Bonus Points...
  • Advanced Cloudflare expertise, including CDN optimization, WAF security, DNS management, edge performance tuning, and Cloudflare Tunnels.
  • Strong understanding of distributed systems architecture, edge caching, and centralized log management using the ELK stack (Elasticsearch, Logstash, Kibana).
  • Experience defining SLOs and instrumentation, implementing meaningful metrics, logs, and traces to reduce alert noise and drive postmortem action items.
Benefits & Perks
  • Salary Range = $122k - $204k (Actual salary available will vary based on location and market factors.)
  • Vibrant team culture
  • Comprehensive Medical, Dental, Vision
  • Training (unlimited Udemy + more)
  • Flexible working hours and time off
  • Equity & Retirement Programs including 401K match
  • Paid Parental Leave
  • International work environment with start-up culture 
About Fandom

Fandom is the world’s largest fan platform where fans immerse themselves in imagined worlds across entertainment and gaming. Reaching more than 350 million unique visitors per month and hosting more than 250,000 wikis, Fandom is the #1 source for in-depth information on pop culture, gaming, TV and film, where fans learn about and celebrate their favorite fandoms. Fandom’s Gaming division manages the online video game retailer Fanatical. Fandom Productions, the content arm of Fandom, enhances the fan experience through curated editorial coverage and branded content from trusted and established publishing brands Gamespot, TV Guide and Metacritic, along with its Emmy-nominated Honest Trailers and the weekly video news program The Loop. For more information follow @getfandom or visit: www.fandom.com.

Fandom is an equal opportunity employer. Fandom values diversity, and all employment decisions are made on the basis of job requirements and individual qualifications.

#LI-TM1

Skills Required

  • 5+ years of experience in Technical Operations, Network Operations, DevOps, or SRE roles managing large-scale production platforms
  • Deep proficiency in Linux systems administration
  • Deep proficiency in networking protocols, including TCP/IP and routing
  • Experience with secure systems practices
  • Scripting or programming experience with Go, Python, or Bash
  • Hands-on experience with Kubernetes or container orchestration
  • Experience with CI/CD pipelines, including GitHub Actions or Jenkins
  • Experience with monitoring and reliability systems such as Prometheus
  • Production MySQL database experience, including replication topologies, failover mechanisms, and performance tuning
  • Demonstrated capability using generative AI tools such as Gemini or NotebookLM, with ability to audit outputs for accuracy, security, and context
  • Advanced Cloudflare expertise, including CDN optimization, WAF security, DNS management, edge performance tuning, and Cloudflare Tunnels
  • Understanding of distributed systems architecture and edge caching
  • Experience with centralized log management using the ELK stack
  • Experience defining SLOs and instrumentation, including meaningful metrics, logs, and traces

Fandom Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Fandom and has not been reviewed or approved by Fandom.

  • Healthcare Strength Health coverage is described as robust, including medical, dental, vision, disability, and mental-health support. Employee-facing materials and third-party listings consistently highlight comprehensive healthcare.
  • Leave & Time Off Breadth Time-off offerings span flexible time off, paid holidays and sick time, paid parental leave, and paid volunteer time. Public materials and listings also reference generous or unlimited PTO in recent periods.
  • Retirement Support Retirement programs include a 401(k) with company match. Additional financial benefits such as an employee stock purchase plan are also cited on company and aggregator pages.

Fandom Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, CA
505 Employees
Year Founded: 2006

What We Do

Fandom is a global entertainment media brand powered by fan passion. The fan-trusted source in entertainment, Fandom provides a home to explore, contribute to, and celebrate the world of pop culture. Whether looking for in-depth information on favorite fandoms or what’s buzzing in entertainment, Fandom has your pop culture curiosities covered through fan-expert knowledge and carefully curated and fun, original multi-platform content. Fandom has a global audience of 200 million monthly uniques and encompasses over 400,000 fan communities. We currently feature more than 55 million pages of content, inclusive of video.

Similar Jobs

BlackLine Logo BlackLine

Senior Cloud Engineer

Cloud • Fintech • Information Technology • Machine Learning • Software • App development • Generative AI
Hybrid
Pleasanton, CA, USA
1810 Employees
136K-170K Annually

Federal Reserve Bank of Boston Logo Federal Reserve Bank of Boston

Data Engineer

Fintech • Information Technology • Payments • Sharing Economy • Financial Services • Cryptocurrency
In-Office
12 Locations
1200 Employees
140K-210K Annually

PwC Logo PwC

Site Reliability Engineer

Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Hybrid
9 Locations
370000 Employees
124K-280K Annually

CrowdStrike Logo CrowdStrike

Senior Engineer

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Hybrid
5 Locations
11000 Employees
140K-215K Annually

Similar Companies Hiring

Hedra Thumbnail
Software • News + Entertainment • Marketing Tech • Generative AI • Enterprise Web • Digital Media • Consumer Web
San Francisco, CA
14 Employees
Bankrate Thumbnail
Artificial Intelligence • Consumer Web • Digital Media • Fintech • Marketing Tech • Software • Financial Services
US
160 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account