Senior SRE

Posted 23 Days Ago
Be an Early Applicant
Ulaanbaatar, Sukhbaatar, Ulaanbaatar, MNG
In-Office
Senior level
Fintech • Software • Financial Services
The Role
Design, build, and maintain highly available, secure, and scalable cloud and Kubernetes infrastructure. Monitor production, automate provisioning and CI/CD, lead incident response and postmortems, implement SLOs/SLIs, optimize capacity and costs, build observability and runbooks, and mentor junior engineers to improve reliability and deployment processes.
Summary Generated by Built In

As a Senior Site Reliability Engineer (SRE), you will play a key role in designing, building, and maintaining highly available, secure, and scalable infrastructure across AND Global and its subsidiaries. You will collaborate closely with engineering teams to deliver reliable platform solutions, automate operational processes, and ensure system stability while supporting new business initiatives.

You'll join a collaborative team of four experienced engineers and contribute to shaping infrastructure best practices, operational excellence, and continuous improvement.

Key Responsibilities:

  • Monitor production systems, application availability, performance, and overall system - health.
  • Build and maintain reliable cloud and Kubernetes infrastructure.
  • Automate infrastructure provisioning, deployment, monitoring, and recovery processes.
  • Investigate production incidents, identify root causes, and implement permanent fixes.
  • Lead incident response and prepare post-incident reports.
  • Work with development teams to improve application reliability and release processes.
  • Define and maintain service-level indicators, service-level objectives, and availability targets.
  • Improve CI/CD pipelines and deployment procedures.
  • Perform capacity planning, performance tuning, and cost optimization.
  • Build dashboards, alerts, monitoring standards, and operational runbooks.
  • Maintain backup, disaster recovery, and high-availability procedures.
  • Identify technical risks, bottlenecks, and areas for continuous improvement.
  • Mentor junior engineers and support engineering best practices.

Requirements:

  • Bachelor’s degree in computer science, information technology, or equivalent practical experience.
  • At least 5 years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Cloud Engineering, or system administration.
  • Strong Linux administration and troubleshooting skills.
  • Production experience with AWS, Microsoft Azure, or Google Cloud.
  • Strong experience with Kubernetes and Docker.
  • Experience with infrastructure-as-code tools such as Terraform or OpenTofu.
  • Experience with CI/CD tools such as GitLab CI, GitHub Actions, Jenkins, or Argo CD.
  • Experience with monitoring and observability tools such as Prometheus, Grafana, Loki, ELK, OpenSearch, Datadog, or OpenTelemetry.
  • Ability to automate tasks using Python, Go, Bash, or another programming language.
  • Good understanding of networking concepts, including DNS, TCP/IP, HTTP/HTTPS, load balancers, VPNs, firewalls, and routing.
  • Experience supporting distributed applications, databases, storage, and messaging systems.
  • Experience with incident management, root-cause analysis, and postmortems.
  • Strong problem-solving, communication, and documentation skills.

Skills Required

  • Bachelor's degree in computer science, information technology, or equivalent practical experience
  • At least 5 years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Cloud Engineering, or system administration
  • Strong Linux administration and troubleshooting skills
  • Production experience with AWS, Microsoft Azure, or Google Cloud
  • Strong experience with Kubernetes and Docker
  • Experience with infrastructure-as-code tools such as Terraform or OpenTofu
  • Experience with CI/CD tools such as GitLab CI, GitHub Actions, Jenkins, or Argo CD
  • Experience with monitoring and observability tools such as Prometheus, Grafana, Loki, ELK, OpenSearch, Datadog, or OpenTelemetry
  • Ability to automate tasks using Python, Go, Bash, or another programming language
  • Good understanding of networking concepts including DNS, TCP/IP, HTTP/HTTPS, load balancers, VPNs, firewalls, and routing
  • Experience supporting distributed applications, databases, storage, and messaging systems
  • Experience with incident management, root-cause analysis, and postmortems
  • Strong problem-solving, communication, and documentation skills
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Ulaanbaatar
29 Employees
Year Founded: 2015

What We Do

AND Global Pte. is a Singapore based fintech company that manages startups developing AI powered financial products. The company’s proprietary platform provides instant credit scoring and unlocks access to unsecured personal loans and payment options instantly to customers on their mobile devices. AND Global's first and so far most successful product, a smartphone-based personal loan platform LendMN was launched in Mongolia in 2016. This AI-driven product assesses credit risk and reduces cost for lending, while providing access to financing for people who are under-banked. AND Systems, the R&D subsidiary of AND Global, is located in Ulaanbaatar, Mongolia.

Similar Jobs

DBS Bank Ltd Logo DBS Bank Ltd

Full-stack Engineer

Fintech • Information Technology • Software • Financial Services
In-Office or Remote
17 Locations
41000 Employees

Binance Logo Binance

Traditional Finance PMO (CEO Office)

Blockchain • Fintech • Software • Cryptocurrency • Metaverse
In-Office or Remote
17 Locations
7696 Employees

Binance Logo Binance

Market Data Lead

Blockchain • Fintech • Software • Cryptocurrency • Metaverse
In-Office or Remote
17 Locations
7696 Employees

Kintsugi AI, Inc. Logo Kintsugi AI, Inc.

Technical Lead

Artificial Intelligence • Fintech • Software • Automation
In-Office or Remote
24 Locations

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account