Site Reliability Engineer

Posted Yesterday
Be an Early Applicant
Vilnius, Vilniaus miesto savivaldybė, Vilniaus apskritis, LTU
In-Office
Mid level
Cloud • Information Technology • Consulting • Design • Generative AI
We are aggressively arming companies with GenAI and Public Cloud technologies daily!
The Role
Operate and automate production infrastructure to keep user-facing services reliable and scalable. Respond to on-call incidents, build infrastructure with IaC (Ansible, Terraform), containerize and run workloads on Kubernetes, improve monitoring/alerting, document runbooks, debug distributed systems, and plan infrastructure growth.
Summary Generated by Built In

Our client's Cloud Operations team is expanding its SRE function. Site Reliability Engineers keep all user-facing services and production systems running smoothly. SREs here are a blend of pragmatic operators and software craftspeople who apply sound engineering principles, operational discipline and mature automation to the environment and the codebase. The team specialises in systems — networking, the Linux kernel, and scaling, algorithms and distributed systems.

As an SRE you will

  • Be on an on-call rotation responding to production availability incidents, and support service engineers with customer incidents
  • Use your on-call shift to prevent incidents from ever happening
  • Run infrastructure with Ansible, Puppet, Terraform and Kubernetes
  • Make monitoring and alerting alert on symptoms, not outages
  • Document every action, so findings turn into repeatable actions — and then into automation
  • Improve the deployment process to make it as boring as possible
  • Design, build and maintain core infrastructure that scales to hundreds of thousands of concurrent users
  • Debug production issues across services and levels of the stack
  • Plan the growth of the infrastructure

You may be a fit if you

  • Think cloud-first, regardless of the flavour of public cloud
  • Think security-first
  • Think about systems — edge cases, failure modes, behaviours, specific implementations
  • Know your way around Linux and Windows
  • Know the use of config-management systems like Ansible or Puppet
  • Have strong programming skills — Python, Java, Golang, Node.js
  • Collaborate and communicate asynchronously, and document so nothing is learned twice
  • Have a go-for-it attitude: when you see something broken, you fix it
  • Have experience with Nginx, HAProxy, Docker, Kubernetes, Terraform or similar technologies

Projects you could work on

  • Coding infrastructure automation with Ansible and Terraform
  • Improving Prometheus monitoring or building new metrics
  • Helping release managers deploy and fix new versions of application software
  • Planning and executing the migration from AWS virtual machines to cloud-native, container-based deployments on Kubernetes (EKS)
  • Developing a relationship with a product group and defining their SRE KPIs — the SRE practice here is early in its journey

    Créé par

    Skills Required

    • On-call rotation and production incident response experience
    • Experience with configuration management and IaC: Ansible, Puppet, Terraform
    • Kubernetes experience (running containerized workloads, EKS familiarity)
    • Linux and Windows systems knowledge/administration
    • Strong programming skills in Python, Java, Golang, or Node.js
    • Experience with Docker, Nginx, HAProxy
    • Experience building or improving monitoring (eg. Prometheus) and alerting
    • Cloud experience (public cloud platforms, AWS experience noted)
    • Ability to debug production distributed systems and plan infrastructure scaling
    • Asynchronous collaboration, documentation, and runbook/process creation
    Am I A Good Fit?
    beta
    Get Personalized Job Insights.
    Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

    The Company
    HQ: Chicago, Illinois
    6 Employees
    Year Founded: 2010

    What We Do

    Ontrac Solutions helps organizations adopt emerging technologies to scale smarter. We build GenAI platforms, predictive analytics solutions, and drive cloud adoption. We're also a HubSpot partner, supporting landing page design, website development, CRM integration, workflows, and automation. From infrastructure to marketing ops, we deliver strategy and execution that drives growth.

    Similar Jobs

    In-Office
    Vilnius, Vilniaus miesto savivaldybė, Vilniaus apskritis, LTU
    243 Employees
    5K-7K Annually

    Ignitis Group Logo Ignitis Group

    Site Reliability Engineer

    Utilities • Renewable Energy
    Hybrid
    Vilnius, Vilniaus miesto savivaldybė, Vilniaus apskritis, LTU
    1628 Employees

    Oxylabs Logo Oxylabs

    Site Reliability Engineer

    Big Data • Information Technology
    Hybrid
    Vilnius, Vilniaus miesto savivaldybė, Vilniaus apskritis, LTU
    500 Employees
    4K-7K Annually

    Hostinger Logo Hostinger

    Engineering Manager

    Information Technology • Consulting
    Hybrid
    2 Locations
    1026 Employees
    6K-9K Annually

    Similar Companies Hiring

    Standard Template Labs Thumbnail
    Artificial Intelligence • Information Technology • Software
    New York, NY
    25 Employees
    Golden Pet Brands Thumbnail
    Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
    El Segundo, California
    178 Employees
    LTX Thumbnail
    Robotics • Conversational AI • Generative AI
    Jerusalem, Israel
    200 Employees

    Sign up now Access later

    Create Free Account

    Please log in or sign up to report this job.

    Create Free Account