DevOps Engineer, Cloud Infrastructure & Live Games

Posted 5 Days Ago
Be an Early Applicant
Toronto, ON, CAN
In-Office
95K-115K Annually
Senior level
Gaming • Mobile
The Role
Design, maintain, and modernize cloud infrastructure and automation for live-service games. Improve CI/CD, IaC, observability, secrets management, data pipeline reliability, and incident response while balancing uptime with modernization. Support AI-enabled tooling, cost management, and create runbooks and SOPs for production ownership.
Summary Generated by Built In
About Big Viking Games

Big Viking Games is a Canadian gaming company focused on building, operating, and growing long-standing online game communities. Our games have entertained players for years, supported by loyal audiences, live operations, evolving content systems, product innovation, and deep player-driven economies.

Our flagship titles, YoWorld and FishWorld, have served millions of players over their lifetime. These are enduring live-service virtual worlds with rich in-game economies, virtual goods, social interaction, and long-term player engagement at their core.

We are entering a new phase of modernization and growth, with a focus on stronger infrastructure, better automation, practical AI adoption, improved reliability, stronger security practices, and scalable systems that help our games and teams perform at a higher level.

About the Role

Big Viking Games is hiring a Senior DevOps Engineer to help design, maintain, secure, and modernize the infrastructure that supports our live-service games and internal development workflows.

This is a hands-on role for someone who understands cloud infrastructure, automation, CI/CD, containers, monitoring, uptime, and production reliability — and who is comfortable working with legacy production systems alongside modern infrastructure patterns. Our games have been running for over a decade; the infrastructure reflects that history, and the right person sees that as an interesting challenge rather than a dealbreaker.

You will work closely with engineering, product, QA, data, and live operations teams to improve how we build, deploy, monitor, and operate our systems.

The right person is practical, security-aware, automation-minded, and able to balance speed with reliability. They can own infrastructure in a live production environment, improve DevOps processes, reduce manual work, and help development teams ship safely and efficiently.

This is a hybrid role based in Toronto, with an expectation of working in office three days per week. Live-service games require operational awareness outside regular business hours, including periodic on-call and incident response availability.


RequirementsWhat You'll Do
  • Monitor, maintain, and improve cloud infrastructure across AWS, Netlify, Vercel, and related platforms that support our live games, data systems, internal tools, and AI-powered operational workflows.
  • Drive infrastructure modernization while maintaining uptime for live games with active player communities — every improvement ships while the plane is flying.
  • Build, maintain, and improve automation for deployments, environment management, provisioning, secrets rotation, and operational workflows — reducing manual toil and human error.
  • Implement and maintain Infrastructure as Code using tools such as Terraform, CloudFormation, CDK, or similar technologies.
  • Maintain and monitor data pipelines between game source databases (MariaDB), the Snowflake data warehouse, and downstream analytics and reporting systems — ensuring pipeline health, freshness, and alerting when data stops flowing.
  • Improve CI/CD pipelines, release workflows, and deployment reliability so development teams can ship safely and frequently.
  • Own secrets and credential lifecycle management across platforms — including API key rotation, access controls, environment variable governance, and least-privilege practices.
  • Support and improve the infrastructure that powers AI and automation tooling, including API integrations, MCP servers, serverless functions, webhook reliability, and orchestration platforms.
  • Improve observability across the stack: logging, metrics, alerting, dashboards, and operational visibility — with particular attention to early detection of silent failures in data pipelines and production systems.
  • Support incident response, root cause analysis, remediation planning, and post-incident improvements.
  • Help manage cloud spend, infrastructure usage, resource tagging, and environment efficiency.
  • Create clear documentation, runbooks, SOPs, and repeatable processes for infrastructure and DevOps workflows.
What You Bring
  • 5+ years of experience in DevOps, infrastructure engineering, cloud engineering, site reliability engineering, or a similar role.
  • Strong hands-on experience with AWS or similar cloud platforms.
  • Experience designing, maintaining, and improving production infrastructure — including comfort with legacy systems that predate modern cloud-native patterns.
  • Proficiency with Infrastructure as Code tools such as Terraform, CloudFormation, CDK, Pulumi, or similar.
  • Experience with containerized applications, especially Docker.
  • Experience with CI/CD tools, version control, deployment automation, and modern release workflows.
  • Strong understanding of Linux systems, networking, cloud security, monitoring, logging, and operational troubleshooting.
  • Experience supporting production systems where uptime, reliability, and performance matter — especially systems that cannot tolerate extended downtime.
  • Experience with relational databases (MariaDB, MySQL, Postgres) and comfort working adjacent to data pipelines and ETL processes.
  • Security-aware mindset with practical experience in secrets management, credential rotation, access control, vulnerability reduction, and least-privilege practices.
  • Strong problem-solving skills and the ability to investigate complex infrastructure or production issues, including silent failures and data pipeline outages.
  • Ability to work closely with software engineers to improve build, deploy, and operational workflows.
  • Comfort creating documentation, runbooks, and repeatable operating processes.
  • Strong communication skills with both technical and non-technical stakeholders.
  • Practical ownership mindset with the ability to prioritize, execute, and close loops.
  • Nice to Have
    • Experience in gaming, live-service products, SaaS, digital products, or other high-availability consumer platforms.
    • Experience supporting live games, virtual worlds, multiplayer systems, or real-time online products.
    • Experience with GitHub Actions, GitLab CI, Jenkins, CircleCI, Buildkite, or similar CI/CD tools.
    • Experience with Datadog, Grafana, Prometheus, CloudWatch, ELK, OpenTelemetry, or similar observability tools.
    • Experience with Redis, Memcached, queues, workers, or event-driven systems.
    • Experience with Snowflake, data warehouse connectivity, ETL monitoring, or data pipeline reliability.
    • Experience with serverless platforms (Netlify Functions, Vercel, AWS Lambda) and multi-platform hosting environments.
    • Experience with disaster recovery, backup strategies, incident management, load testing, and performance tuning.
    • Experience improving cloud cost management, tagging, resource optimization, or infrastructure governance.
    • Experience with container orchestration platforms such as Kubernetes, ECS, EKS, or Nomad.
    • Experience operating infrastructure that supports AI/ML workflows, API integrations, or automation platforms (Make.com, webhook-driven orchestration, MCP servers).
    • Experience using AI tools such as Claude, ChatGPT, Gemini, or similar platforms to improve DevOps workflows, documentation, troubleshooting, and automation.
    • Experience working in small, high-leverage engineering teams where infrastructure ownership is broad and hands-on.
  • Ideal Candidate Profile
  • The ideal candidate is a practical infrastructure engineer who can keep live systems stable while helping modernize how the company builds, deploys, secures, and operates technology.
  • They are not only focused on tools. They understand uptime, developer experience, production risk, cloud costs, security, release quality, and operational discipline. They recognize that modernizing a decade-old live game requires patience, pragmatism, and the ability to improve systems incrementally without disrupting what's working.
  • They are comfortable operating across a mix of legacy and modern infrastructure, managing credentials and secrets lifecycle across multiple hosting platforms, and ensuring data pipelines are healthy and alerting properly. They can work independently, collaborate with engineers, and create systems that reduce friction instead of adding process for its own sake.
  • This role is best suited for someone who wants meaningful ownership over production infrastructure, cloud health, automation, and DevOps modernization inside a live-service gaming company.

BenefitsCompensation

Compensation range: $95,000 to $115,000 determined based on experience, technical depth, infrastructure ownership breadth, and overall fit. 

Benefits
  • Group Retirement Savings Plan matching and participation.
  • Comprehensive benefits package, including health, dental, and vision coverage.
  • Health and Wellness spending account.
  • Generous time off policies.
  • Opportunity to support long-running live-service games with established player communities.
  • Exposure to cloud modernization, DevOps automation, security improvement, and AI-enabled infrastructure workflows.
  • A high-impact role with meaningful ownership over reliability, performance, and engineering operations.
Accessibility and Accommodation

Big Viking Games is committed to creating an inclusive and accessible environment for all candidates. We welcome applications from individuals of all abilities and will provide accommodations throughout the hiring process as needed.

If you require accommodation during the hiring process, please contact [email protected] so we can work with you to support your needs.

Skills Required

  • 5+ years of experience in DevOps, infrastructure engineering, cloud engineering, site reliability engineering, or a similar role.
  • Strong hands-on experience with AWS or similar cloud platforms.
  • Experience designing, maintaining, and improving production infrastructure, including legacy systems.
  • Proficiency with Infrastructure as Code tools such as Terraform, CloudFormation, CDK, or Pulumi.
  • Experience with containerized applications, especially Docker.
  • Experience with CI/CD tools, version control, deployment automation, and modern release workflows.
  • Strong understanding of Linux systems, networking, cloud security, monitoring, logging, and operational troubleshooting.
  • Experience supporting production systems where uptime, reliability, and performance matter.
  • Experience with relational databases (MariaDB, MySQL, Postgres) and working adjacent to data pipelines and ETL processes.
  • Practical security experience: secrets management, credential rotation, access control, and least-privilege practices.
  • Ability to create documentation, runbooks, SOPs, and repeatable operating processes.
  • Strong communication skills with technical and non-technical stakeholders.
  • Ability to work hybrid in Toronto (expected in-office three days per week) and participate in periodic on-call/incident response.
  • Experience in gaming, live-service products, multiplayer systems, or high-availability consumer platforms.
  • Experience with GitHub Actions, GitLab CI, Jenkins, CircleCI, Buildkite, or similar CI/CD tools.
  • Experience with Datadog, Grafana, Prometheus, CloudWatch, ELK, or OpenTelemetry.
  • Experience with Redis, Memcached, queues, workers, or event-driven systems.
  • Experience with Snowflake, data warehouse connectivity, and ETL monitoring.
  • Experience with serverless platforms (Netlify Functions, Vercel, AWS Lambda) and multi-platform hosting.
  • Experience with disaster recovery, backup strategies, incident management, load testing, and performance tuning.
  • Experience improving cloud cost management, tagging, resource optimization, or infrastructure governance.
  • Experience with container orchestration platforms such as Kubernetes, ECS, EKS, or Nomad.
  • Experience operating infrastructure that supports AI/ML workflows, API integrations, or automation platforms (Make.com, MCP servers).
  • Experience using AI tools (Claude, ChatGPT, Gemini) to improve DevOps workflows, documentation, and automation.
  • Experience working in small, high-leverage engineering teams with broad infrastructure ownership.
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: London, Ontario
80 Employees
Year Founded: 2011

What We Do

Since our humble beginnings in 2011, these two words have driven Big Viking Games to become the successful company it is today. We are focused on making our mark by building awesome free-to-play mobile and social games. We’re one of the largest independent mobile and social game studios in Canada and a pioneer in mobile HTML5 games! The company has grown profitably to a team of over 70+ Vikings across two studios in Toronto and London, ON. Prior to 2020, we had 2 studios, but we’ve now converted to a 100% remote company! From our beginnings with hits like FishWorld and YoWorld, we strive to become a leader in live operations. Our titles are played by millions of people on iOS, Android, Facebook, and the web.

Similar Jobs

TransUnion Logo TransUnion

Business Development Executive, Public Sector

Big Data • Fintech • Information Technology • Business Intelligence • Financial Services • Cybersecurity • Big Data Analytics
Hybrid
Burlington, ON, CAN
13000 Employees

Mastercard Logo Mastercard

Director, Technical Program Management

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Hybrid
Toronto, ON, CAN
38800 Employees
150K-234K Annually

Mastercard Logo Mastercard

Director, Product Management:Go-to-Market

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Hybrid
Toronto, ON, CAN
38800 Employees
150K-234K Annually

Mastercard Logo Mastercard

Director - New Product Development

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Hybrid
Toronto, ON, CAN
38800 Employees
150K-234K Annually

Similar Companies Hiring

Prolaio Thumbnail
Artificial Intelligence • Big Data • Healthtech • Mobile • Wearables • Analytics
Chicago, IL
82 Employees
ARB Interactive Thumbnail
Gaming • Software
Miami, Florida
175 Employees
Granted Thumbnail
Mobile • Insurance • Healthtech • Financial Services • Artificial Intelligence
New York, New York
23 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account