Site Reliability Engineer (India)

Posted 4 Days Ago
Be an Early Applicant
New Delhi, Delhi, IND
In-Office
Entry level
Database • Analytics
The Role
Build and maintain reliable, scalable cloud infrastructure for Kubernetes-based distributed database systems on AWS and Google Cloud. Develop Golang automation, optimize CI/CD pipelines, implement monitoring and observability, troubleshoot complex infrastructure issues, and design disaster recovery and high-availability strategies. Participate in on-call support, incident response, and cross-functional efforts to improve system reliability and resolve customer issues.
Summary Generated by Built In
                                                                         Site Reliability Engineer 
About Arango:
Arango delivers a unified, natively multimodel contextual data platform that powers AI agents, assistants, and applications with the unified, current, and trusted business context needed to reason, decide, and act at scale.
The Arango Contextual Data Platform connects fragmented enterprise data with LLMs, copilots, and AI agents through a simplified architecture delivered out of the box. By combining graph, vector, document, key-value, and search capabilities in a single platform, Arango eliminates the complex stacks many organizations build to operationalize enterprise AI.
Trusted by organizations including NVIDIA, HPE, the London Stock Exchange, PSI CRO, the U.S. Air Force, NIH, Siemens, Transient.AI, Matpriskollen, and Articul8, Arango helps enterprises move from AI pilots to reliable production systems faster while lowering infrastructure complexity and total cost of ownership. Arango is a proud member of the NVIDIA Inception Program and the AWS ISV Accelerate Program. Learn more at arango.ai, LinkedIn, and G2.
Location: (Remote)- India
Job Overview: 
At ArangoDB, we are building a robust, cloud-native infrastructure to support our distributed database systems, which power mission-critical applications for a wide range of industries. We are searching for a Site Reliability Engineer (SRE) to ensure the reliability, scalability, and performance of our infrastructure and applications, with a focus on automation, monitoring, and optimizing cloud environments. 
As a Site Reliability Engineer (SRE), you will be responsible for maintaining and improving the reliability of our distributed database systems running on Kubernetes and cloud environments (AWS, Google Cloud). You will design, implement, and maintain scalable infrastructure solutions, improve and expand observability into these solutions, and troubleshoot complex system issues. It is expected that you will come to work to write clean and efficient code in Golang, working closely with development teams
Your goal is to ensure high availability and performance of our cloud-based systems, automating repetitive tasks, and enhancing our CI/CD pipelines. If you're passionate about building resilient systems, managing cloud infrastructure, and using Golang to create scalable solutions (or willingness to learn Golang), we want to hear from you! 
Key Responsibilities: 
● Design, implement, and maintain cloud infrastructure on AWS and Google Cloud platforms. 
● Ensure the scalability, performance, and reliability of our Kubernetes-based distributed database systems. 
● Collaborate with developers to write efficient, production-grade code in Golang to automate infrastructure management and improve system operations. 
● Optimize and automate CI/CD pipelines, deployment processes, and monitoring systems to support our production environment. 
● Develop strategies for disaster recovery, high availability, and fault tolerance. 
● Proactively identify system bottlenecks, troubleshoot, and resolve issues across the stack (network, OS, cloud infrastructure). 
● Implement monitoring, logging, and alerting systems to ensure visibility into system health and performance. 
● Participate in on-call rotations to support critical production systems and respond to incidents. 
● Collaborate with cross-functional teams to improve overall system reliability and scalability. 
● Collaborate with the Customer Success team to resolve customer issues. 
Required Skills and Qualifications: 
● Experience: 4-7 years of experience: SRE or DevOps Engineer background in cloud-native environments. Self-organized, autonomous remote team player with strong communication skills.
● Cloud & Infrastructure: AWS and GCP; advanced Linux internals (processes, environment variables); containerization and orchestration (Docker, Kubernetes at scale).
● CI/CD & Observability: CI/CD pipelines (Jenkins, CircleCI); monitoring, alerting, and logging (Prometheus, Grafana, ELK stack); Git version control.
● Networking & Security: Core networking, security best practices, and systematic troubleshooting of complex infrastructure issues.
● Development: Programming proficiency in Golang or Python.
Nice-to-Have: 
● Experience managing distributed databases or large-scale data storage systems.
●Knowledge of security best practices in cloud environments. 
● Experience with scripting languages like Python or Bash. 
● Experience with Infrastructure-as-Code (IaC) tools like Terraform is a plus.
●Experience working with GitOps 
● Strong programming skills in Golang, with experience in developing automation tools, scripts, or services. 

 What Makes Arango Special?At Arango, we believe that AI is only as powerful as the data foundation. Our mission is to help organizations build AI systems that can reason, decide and act based on unified, current, and trusted business context at scale. We are helping define a new category of infrastructure: the contextual data layer for AI.
Working at Arango means:
Contributing to cutting-edge AI and data infrastructure
Collaborating with experienced engineers, marketers, and product leaders
Helping shape how enterprises build AI-powered applications
If you're excited about the intersection of AI, data, and social media, we’d love to hear from you.

Skills Required

  • Experience in an SRE or DevOps Engineer role within cloud-native environments
  • Ability to work autonomously and communicate effectively on a remote team
  • Experience with AWS and Google Cloud
  • Advanced Linux internals knowledge, including processes and environment variables
  • Experience with Docker and Kubernetes at scale
  • Experience with CI/CD pipelines, including Jenkins or CircleCI
  • Experience with monitoring, alerting, and logging tools, including Prometheus, Grafana, or ELK Stack
  • Git version control experience
  • Knowledge of core networking, security best practices, and systematic infrastructure troubleshooting
  • Programming proficiency in Golang or Python
  • Experience managing distributed databases or large-scale data storage systems
  • Experience with cloud security best practices
  • Experience with Python or Bash scripting
  • Experience with Infrastructure as Code tools such as Terraform
  • Experience working with GitOps
  • Strong Golang programming skills for automation tools, scripts, or services
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, CA
110 Employees
Year Founded: 2014

What We Do

ArangoDB is the most scalable open-source graph database, with more than 12,000 stargazers on GitHub. Building on the concept of ‘graph and beyond’, ArangoDB combines the analytical power of graphs with JSON documents, a key-value store, and a full-text search engine, enabling developers to access and combine all of these data models with a single, elegant, declarative query language. It serves as the scalable backbone for graph analytics and complex data architectures across many industries. Founded in 2015, ArangoDB Inc. is a privately-held company backed by Bow Capital, Iris Capital, New Forge, and Target Partners. It is headquartered in San Francisco and Cologne, Germany, with offices and employees around the world. Learn more at www.arangodb.com.

Similar Jobs

CSC Logo CSC

Senior Software Engineer

Fintech • Legal Tech • Software • Financial Services • Cybersecurity • Data Privacy
Remote or Hybrid
2 Locations
8500 Employees

CSC Logo CSC

Accountant

Fintech • Legal Tech • Software • Financial Services • Cybersecurity • Data Privacy
Remote or Hybrid
2 Locations
8500 Employees

Tufin Logo Tufin

Professional Services Engineer

Security • Cybersecurity
Remote or Hybrid
India
500 Employees

Ericsson Logo Ericsson

Integration Engineer

Cloud • Information Technology • Internet of Things • Machine Learning • Software • Cybersecurity • Infrastructure as a Service (IaaS)
In-Office
6 Locations
88000 Employees

Similar Companies Hiring

Northslope Thumbnail
Artificial Intelligence • Information Technology • Software • Analytics • Consulting • Generative AI
London, GB
100 Employees
Scotch Thumbnail
Artificial Intelligence • eCommerce • Fintech • Payments • Retail • Software • Analytics
US
35 Employees
Milestone Systems Thumbnail
Artificial Intelligence • Security • Software • Analytics • Big Data Analytics
Lake Oswego, OR
1500 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account