Site Reliability Engineer

Posted Yesterday
Be an Early Applicant
Pune, Mahārāshtra, IND
In-Office
Junior
Information Technology • Internet of Things • Logistics • Software • 3PL: Third Party Logistics
The Role
Manage and improve large-scale 24/7 cloud operations and production reliability. Responsibilities include administering AWS, Linux, Kubernetes, databases, and infrastructure; troubleshooting incidents; developing automation; participating in on-call rotations; meeting uptime SLAs; documenting runbooks; and preparing postmortems. The role requires rotational shifts, night and weekend availability, incident and change management, and collaboration with internal teams and customers.
Summary Generated by Built In
Job Description:
· Intangles Lab is looking for a hands-on Site Reliability Engineer from FinTech background to manage large 24×7 Cloud Operations.
· Looking for a Site Reliability Engineer with 2+ years of experience, having hands-on with the following technologies/skillset:

Must-Required Skills:
·      AWS Cloud (Advanced): Certification is preferred.
·      Networking (Intermediate): Proficiency in networking concepts is necessary.
·      Ubuntu/Linux & OS (Advanced): Strong Linux & Networking basics, Prior working experience is preferred.
·      Database (Basic Knowledge): Familiarity with SQL and NoSQL databases is required, having worked with at least one of them.
· Database Administration (MongoDB & PostgreSQL, Elasticsearch), having hands-on experience of at least one is required
·      Containerization Tools: Docker
·      Kubernetes (Advanced)
·      Knowledge of Amazon EKS is compulsory.
·      Working knowledge of StatefulSets is required.
·      Familiarity with the HELM Chart is necessary.

CI/CD (Advanced):
· Proficiency in at least one CI/CD tool, such as CircleCI, Argo Project, GitHub Actions, or similar, is essential.
· Programming:
a.Basic programming knowledge is required, with the ability to write code.
b.Scripting Language: Python, Shell

Monitoring Stack:
·      Prometheus, Grafana, Alert Mangaer, Istio, Jaeger, Datadog, PagerDuty (or similar). ElasticAPM

Optional Skills:
·      Medium to High Level of Application Development Experience in languages like JavaScript, Python, and Java will be a bonus.
·      Understanding of N-tier Architectures
·      Understanding of REST & gRPC API Frameworks
·      Understanding of Web Servers in NodeJS


Responsibilities:
·      To work in a production environment with technologies like Linux, AWS, Terraform, Kubernetes, MongoDB, Elasticsearch & PostgreSQL Administration.
·      To keep the production environment up & running, i.e. ensuring the reliability of the production environment.
·      To troubleshoot, debug and fix issues in case of failures of the production and QA environment and provide technical solutions.
·      To own the responsibilities of on-call as per the team’s policy.
·      To write and enhance automations as and when needed.
·      To work closely with internal teams and customers to follow the processes and SLAs of uptime.
·      To write, update and enhance documentation, including runbooks/playbooks and prepare postmortem reports for the production incidents.
·      Considering the role is to ensure the platform’s reliability, ready to work in a 24*7 work environment when required. 
 Additional Requirements:
·      One should be aware of change/incident/problem/issue/risk management/escalations.
·      Should be flexible in working in rotational shifts and night hours (Including weekends).
·      Excellent thinking and problem-solving skills.

Skills Required

  • 2+ years of Site Reliability Engineering or relevant experience
  • FinTech industry background
  • Advanced AWS Cloud knowledge
  • AWS certification
  • Intermediate networking proficiency
  • Advanced Ubuntu/Linux and operating system skills
  • Basic knowledge of SQL and NoSQL databases
  • Hands-on database administration experience with MongoDB, PostgreSQL, or Elasticsearch
  • Docker containerization experience
  • Advanced Kubernetes knowledge
  • Amazon EKS experience
  • StatefulSets experience
  • Helm chart familiarity
  • Advanced proficiency with at least one CI/CD tool, such as CircleCI, Argo, or GitHub Actions
  • Basic programming and code-writing ability
  • Python and Shell scripting experience
  • Familiarity with monitoring and observability tools such as Prometheus, Grafana, Alertmanager, Istio, Jaeger, Datadog, PagerDuty, or Elastic APM
  • Application development experience with JavaScript, Python, or Java
  • Understanding of N-tier architectures
  • Understanding of REST and gRPC APIs
  • Understanding of Node.js web servers
  • Availability for rotational shifts, night hours, weekends, and 24/7 operations
  • Knowledge of change, incident, problem, issue, risk management, and escalation processes
  • Excellent problem-solving skills
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Pune, Maharashtra
200 Employees
Year Founded: 2016

What We Do

Intangles is an IOT company based out of Pune with its expertise in Advance Telemetry and Operations Automation. Intangles has developed its own set of proprietary algorithms which allows fleet operators to monitor the performance of the vehicle in real time. The biggest USP is the state of the art proprietary algorithms which allows fleet operators to take informed decisions. Intangles aspires to be the world’s foremost authority in Telemetry and Vehicle Performance monitoring through predictive/prognostic monitoring and benchmarking performance of the vehicle based on its own Data analytics platform “Indium”. Our vision is to help fleet operators monitor, benchmark and do predictive maintenance of assets and identify underperforming assets to increase overall operational profits.

Similar Jobs

Citi Logo Citi

Site Reliability Engineer

Fintech • Financial Services
In-Office
2 Locations
223850 Employees
In-Office
2 Locations
6435 Employees
In-Office
2 Locations
6435 Employees

Hitachi Solutions America Logo Hitachi Solutions America

Site Reliability Engineer

Information Technology • Consulting
In-Office
Pune, Mahārāshtra, IND
768 Employees

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
70 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account