Senior Site Reliability Engineer (SRE) | Feeld

Posted 5 Days Ago
Hiring Remotely in Brazil
Remote
Senior level
Information Technology
The Role
Own reliability and observability for critical user journeys within a distributed product squad. Build metrics, dashboards, alerts, SLIs, and SLOs; improve monitoring, logging, tracing, and incident response; act as a first responder for P0/P1 incidents; investigate and mitigate production issues; coordinate escalations and drive postmortem-based infrastructure and process improvements.
Summary Generated by Built In

GT was founded in 2019 by a former Apple, Nest, and Google executive. GT’s mission is to connect the world’s best talent with product careers offered by high-growth companies in the UK, USA, Canada, Germany, and the Netherlands.

On behalf of Feeld, GT is looking for a Senior Site Reliability Engineer (SRE) to join a fast-growing consumer mobile product in the online dating space.

 
About the Client

Founded in 2014 as a dating app, Feeld gathered millions of users in one place to create a safer and more inclusive space online for everyone open to experiencing people and relationships in a new way. Their mission is to elevate the human experience of sexuality and relationships and create a world where everyone is more intimately connected to each other and themselves.

About the Project

You’ll join a consumer mobile product with an engineering and product organization of around 50 people distributed across Europe and the US.

The team works in small, autonomous product squads, each responsible for a specific area of the product and critical user journeys. Because the team operates across multiple regions without a full follow-the-sun model, strong observability, monitoring and reliable incident response are essential.

From a technical perspective, the team is focused on building reliable, observable systems that allow engineers to identify issues early, understand their impact and respond quickly when incidents occur.

  • Technology stack: Node.js, TypeScript, AWS, Cloudflare, CloudWatch, Sentry. React Native is used on the mobile side.

  • Team: Cross-functional product squads of approximately 6–8 people, distributed across Europe, the US and LATAM.

About the Role

We are looking for an experienced Site Reliability Engineer with a strong backend engineering background in Node.js and TypeScript.

The ideal profile is someone who started in backend/software engineering and has moved into SRE or reliability-focused work, combining a strong understanding of application code with hands-on experience in observability, monitoring and production incident management.

You will be embedded within a product squad and take ownership of the reliability and observability of critical user journeys. An important part of the role is being able to interpret production signals, identify when something is going wrong and begin mitigating incidents independently while bringing in the wider engineering team when needed.

Responsibilities:
  • Own observability for critical product and user journeys within your squad.

  • Define, build and maintain meaningful metrics, dashboards and alerts.

  • Define and maintain SLIs/SLOs for key services and product-level metrics.

  • Improve monitoring, logging, tracing and alerting across the squad’s systems.

  • Act as the first responder for critical P0/P1 production incidents, including out-of-hours incidents.

  • Investigate production signals, identify potential root causes and begin mitigating issues independently.

  • Coordinate with other engineers when broader support or escalation is required.

  • Participate in incident triage, mitigation and postmortems.

  • Identify recurring reliability issues and drive improvements to infrastructure, tooling and incident-response processes.

  • Work closely with backend and product engineers in a distributed, autonomous squad.

Essential knowledge, skills & experience:
  • Strong previous experience as a Backend / Software Engineer, with senior-level hands-on experience in Node.js and TypeScript.

  • Hands-on experience working in an SRE, Production Engineering or similar reliability-focused role.

  • Strong production experience with AWS.

  • Experience with Cloudflare and CloudWatch.

  • Experience with monitoring and observability across metrics, logging, tracing and alerting.

  • Practical experience responding to production incidents, including triage, mitigation and postmortems.

  • Ability to interpret monitoring signals and independently investigate and begin resolving production issues.

  • Understanding of both the application and infrastructure layers rather than infrastructure-only experience.

  • Strong communication skills and the ability to work autonomously within a distributed engineering team.

  • Comfortable participating in out-of-hours incident response as part of the team’s coverage model.

Nice-to-have
  • Experience with observability tools such as Sentry.

  • Experience defining SLIs and SLOs for product-level metrics.

  • Experience with React Native or exposure to mobile application environments.

  • Previous experience with consumer mobile products or high-traffic B2C systems.

Interview Steps
  1. GT interview with Recruiter

  2. Technical interview

  3. Final interview

  4. Reference Check

We go beyond usual perks… By working with us, you will get:
  • Health insurance.

  • Wellbeing budget.

  • Sport coverage.

  • Learning budget.

  • 18 business days of paid vacation days per year

  • Paid sick leaves.

  • All public holidays are paid days off.

GT working model:

You will work directly with a client through our Extended Team model. We try to do things differently and put our efforts into integrating you as deeply as possible into the client’s team. You work with the same tools and technologies as they do and are managed directly by the client without any intermediary in between. We help you build relationships and create an environment where you genuinely feel like a member of the client’s team. We also encourage trips to a client and join teambuilding and after-work activities. Our Extended Team model is focused on long-term projects that last over several years.

Skills Required

  • Senior-level backend or software engineering experience with Node.js and TypeScript
  • Hands-on experience in SRE, Production Engineering, or a similar reliability-focused role
  • Strong production experience with AWS
  • Experience with Cloudflare and CloudWatch
  • Experience with metrics, logging, tracing, monitoring, and alerting
  • Practical production incident response experience, including triage, mitigation, and postmortems
  • Ability to investigate and begin resolving production issues independently
  • Understanding of both application and infrastructure layers
  • Strong communication skills and ability to work autonomously in a distributed engineering team
  • Availability for out-of-hours incident response
  • Experience with observability tools such as Sentry
  • Experience defining SLIs and SLOs for product-level metrics
  • Experience with React Native or mobile application environments
  • Experience with consumer mobile products or high-traffic B2C systems
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
686 Employees
Year Founded: 2019

What We Do

GT was founded in 2019 by a former executive from Apple, Nest, and Google. GT’s mission is to build teams and products that address some of the more challenging requirements from fast-growth clients in Europe and North America.

Similar Jobs

Luxury Presence Logo Luxury Presence

Senior Devops Engineer

Marketing Tech • Real Estate • Software • PropTech • SEO
Easy Apply
Remote or Hybrid
12 Locations
500 Employees

Milestone Systems Logo Milestone Systems

Channel Business Manager - Brazil

Artificial Intelligence • Security • Software • Analytics • Big Data Analytics
Remote or Hybrid
2 Locations
1500 Employees

CrowdStrike Logo CrowdStrike

Regional Sales Manager

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
Brazil
11000 Employees

CrowdStrike Logo CrowdStrike

Technical Support

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
Brazil
11000 Employees

Similar Companies Hiring

Standard Template Labs Thumbnail
Artificial Intelligence • Information Technology • Software
New York, NY
25 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account