SRE — Incidents & Monitoring

Reposted 6 Hours Ago
Be an Early Applicant
Hiring Remotely in São Paulo, BRA
In-Office or Remote
Senior level
Fintech • Software • Financial Services
The Role
Lead response to highest-severity production incidents, diagnose and resolve outages, run post-mortems, design and maintain observability (metrics, logs, tracing, alerts), reduce MTTR, expand monitoring coverage, participate in on-call rotation and improve runbooks.
Summary Generated by Built In

Why work with us? 

 

We are a fast-growing company that is revolutionizing the world of SaaS platform and data in Latin America! 

CIAL Dun & Bradstreet is the leading provider of business decisioning solutions and commercial data across Latin America and the Caribbean. Our solutions are designed to transform how businesses manage risk and make critical decisions about the companies they rely on. 

It’s our people, not technology, that makes what we do possible. It’s our people, not data, that turns information into insights. And it’s our people, not algorithms, that help our clients make better informed decisions. We are innovative, agile, and inspired by SaaS solutions and data – and we are looking for people that share those values to join our mission to build stronger businesses! 


About Us 

Dunsguide, by CIAL Dun & Bradstreet, is transforming how Latin American businesses discover and evaluate potential partners. As the region's leading platform for customer prospecting and supplier discovery, we help companies make smarter, data-driven decisions about their business relationships. We're growing rapidly and expanding our suite of solutions that combine deep business intelligence with modern, intuitive tools. 

Sobre a posição

Atuar na linha de frente de confiabilidade da CIAL: resposta a incidentes de produção de classificação mais crítica (highest severity) e monitoramento contínuo de aplicações, serviços e bancos de dados. É uma posição sênior — a pessoa precisa tomar decisão sob pressão em incidente e evoluir a observabilidade da plataforma.

Responsabilidades

  • Responder a incidentes de produção críticos, conduzindo o diagnóstico até a resolução e o post-mortem.
  • Projetar e manter a stack de observabilidade (métricas, logs, tracing, alertas) das aplicações e serviços.
  • Reduzir MTTR e aumentar a cobertura de monitoramento e a disponibilidade dos serviços.
  • Participar de escala de on-call e melhorar continuamente os runbooks.

Requisitos obrigatórios

  • Experiência sênior em SRE / DevOps / Engenharia de Produção.
  • Domínio de ferramentas de observabilidade (ex.: Datadog, Grafana, Prometheus ou equivalentes).
  • Sólida experiência com ambientes cloud e infraestrutura em produção.
  • Prática consolidada em gestão de incidentes e resposta a on-call.
  • Scripting para automação (Python, Bash ou similar).
  • Inglês avançado

Diferenciais

  • Certificações cloud (AWS/GCP/Azure).
  • Experiência com Kubernetes e Infrastructure as Code.
  • Vivência em ambientes de dados / pipelines de crédito.
    CIAL provides equal employment opportunities to all employees and applicants for employment and prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state, or local laws.  

Skills Required

  • Senior experience in SRE / DevOps / Production Engineering
  • Mastery of observability tools (Datadog, Grafana, Prometheus or equivalents)
  • Solid experience with cloud environments and production infrastructure
  • Established practice in incident management and on-call response
  • Scripting for automation (Python, Bash or similar)
  • Cloud certifications (AWS/GCP/Azure)
  • Experience with Kubernetes
  • Infrastructure as Code
  • Experience in data environments / credit pipelines
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Teaneck, NJ
418 Employees
Year Founded: 2016

What We Do

Data, platforms and technology to transform processes and improve your B2B business decision globally.

Similar Jobs

CrowdStrike Logo CrowdStrike

Technical Support

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
Brazil
11000 Employees

JumpCloud Logo JumpCloud

Partner Sales - Brazil

Cloud • Information Technology • Security • Software
Easy Apply
In-Office or Remote
São Paulo, BRA
800 Employees

Tapestry - Coach and Kate Spade Logo Tapestry - Coach and Kate Spade

Sr. Sales Associate I

eCommerce • Fashion • Retail • Sales • Wearables • Design
Remote or Hybrid
14 Locations
16000 Employees
15-20 Hourly

GitLab Logo GitLab

Area Vice President - LATAM

Cloud • Security • Software • Cybersecurity • Automation
Easy Apply
Remote
Brazil
2500 Employees

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account