SRE / High-Middle DevOps

Posted Yesterday
Be an Early Applicant
2 Locations
In-Office
Mid level
Fintech • Gaming • Payments • Software
The Role
Operate and improve a high-load production platform on GCP/GKE. Implement infrastructure changes with Terraform and Helm, improve CI/CD (GitHub Actions), build Datadog observability, participate in incident response and postmortems, support security and cost optimization, and reduce operational toil.
Summary Generated by Built In

Aghanim is an integrated commerce, liveops automation, community engagement, and payments platform for video games.

 

Mobile games have traditionally depended on app stores for distribution, payments, and player relationships. We believe there is a better way. Aghanim helps game studios build direct relationships with players, sell directly, and build their future on their own terms. Today, more than 100 games worldwide are already building this future with Aghanim.

 

Our team brings together people across Los Angeles, New York, Seoul, Beijing, London, Lisbon, Belgrade and other locations around the globe, with deep expertise in gaming, fintech and technology. We move quickly, keep communication direct, and focus on getting things done. We believe the best people thrive when they have autonomy, ownership, and a stake in the company's success.

We’re looking for a Middle/High-Middle DevOps / SRE Engineer to help run and improve our production platform in GCP + GKE, fronted by Cloudflare, with observability in Datadog and CI/CD in GitHub Actions.

You’ll work closely with Senior/Principal engineers, implementing reliability improvements, expanding monitoring coverage, and reducing operational toil - especially important in a highload system with sudden traffic spikes.

Key Responsibilities

Platform Operations

  • Operate and improve production systems on GCP, GKE, and related managed services

  • Contribute to platform reliability, scalability, and operational improvements alongside Senior and Principal engineers

Infrastructure & Delivery

  • Implement infrastructure changes using Terraform and maintain Kubernetes configurations, Helm charts, and deployment tooling

  • Improve CI/CD automation and deployment reliability

Observability & Incident Management

  • Build and maintain monitoring, alerting, and observability in Datadog

  • Participate in incident response, troubleshooting, and postmortem activities

Security & Cost Optimization

  • Support security tooling, vulnerability remediation, and secure platform practices

  • Identify and implement cost optimization opportunities without compromising reliability

Required Qualifications
  • Hands-on production experience with Kubernetes (ideally GKE) and basic cluster operations.

  • Working experience with Terraform and Helm in PR-based workflows.

  • Familiarity with GCP services used in SaaS operations (e.g., Cloud SQL, BigQuery, BigTable, Pub/Sub, Cloud Run, Memorystore).

  • Monitoring/alerting and troubleshooting skills (preferably Datadog).

  • Strong scripting/automation mindset to reduce manual work and prevent repetitive incidents.

  • Reliability awareness: understanding how changes affect availability/latency and how to operate under SLA constraints.

Preferred Qualifications
  • Cloudflare basics (WAF/DNS, edge concepts; Workers/CDN is a plus).

  • Experience writing/maintaining runbooks and participating in postmortems.

  • Exposure to SOC 2 / PCI-DSS requirements or willingness to learn.

  • Experience in high-load consumer products or game dev.

Why Join Us
  • World-class team – work alongside experienced professionals from around the globe who have built products used by millions of players

  • High growth, high impact – be part of a fast-growing company where ideas turn into products and reach customers in days, not months

  • Autonomy and ownership – we trust people to make decisions, take initiative, and drive results

  • Modern tools and technology – use AI, automation, and modern tools as part of your everyday work

  • Equity – participate in the company's growth and long-term success

Skills Required

  • Hands-on production experience with Kubernetes (ideally GKE) and basic cluster operations.
  • Working experience with Terraform and Helm in PR-based workflows.
  • Familiarity with GCP services used in SaaS operations (Cloud SQL, BigQuery, BigTable, Pub/Sub, Cloud Run, Memorystore).
  • Monitoring/alerting and troubleshooting skills (preferably Datadog).
  • Strong scripting/automation mindset to reduce manual work and prevent repetitive incidents.
  • Reliability awareness: understanding how changes affect availability/latency and operating under SLA constraints.
  • Cloudflare basics (WAF/DNS, edge concepts; Workers/CDN is a plus).
  • Experience writing/maintaining runbooks and participating in postmortems.
  • Exposure to SOC 2 / PCI-DSS requirements or willingness to learn.
  • Experience in high-load consumer products or game development.
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
16 Employees
Year Founded: 2023

What We Do

Aghanim is a direct-to-consumer commerce and payments platform for mobile games. It helps game studios sell directly outside app stores, build direct player relationships, and manage monetization, live operations, community engagement, and payment compliance. Its tools include web-based game hubs, virtual goods and subscriptions, segmentation, automated campaigns, predictive analytics, incentives, and merchant-of-record services, supporting developers’ financial and creative independence worldwide.

Similar Jobs

Datadog Logo Datadog

Senior Software Engineer

Artificial Intelligence • Cloud • Security • Software • Cybersecurity
Easy Apply
Hybrid
Lisbon, PRT
6500 Employees

Deepgram Logo Deepgram

Sales Development Representative

Artificial Intelligence • Machine Learning • Natural Language Processing • Software • Conversational AI
In-Office or Remote
28 Locations
150 Employees

Riskified Logo Riskified

Senior Full-stack Engineer

Big Data • eCommerce • Fintech • Machine Learning • Payments • Software
Hybrid
Lisbon, PRT
680 Employees

Riskified Logo Riskified

Full-stack Engineer

Big Data • eCommerce • Fintech • Machine Learning • Payments • Software
Hybrid
Lisbon, PRT
680 Employees

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account