Senior SRE Engineer

Posted 6 Days Ago
Be an Early Applicant
New York, NY, USA
Hybrid
Senior level
AdTech • Big Data • Internet of Things • Marketing Tech • Mobile • Software • Analytics
Fans to Brands
The Role
Senior Site Reliability Engineer responsible for improving platform availability, scalability, resilience, and observability. The role manages AWS and EKS infrastructure with Terraform, builds CI/CD pipelines using GitHub Actions, expands GitOps through ArgoCD and Helm, strengthens disaster recovery, supports incident response, and develops monitoring, alerting, dashboards, and SLOs. The engineer partners with product teams, leads infrastructure initiatives, troubleshoots Kubernetes, and supports high-availability distributed systems.
Summary Generated by Built In
Senior SRE Reliability Engineer

Location: New York, NY (Hybrid) / Remote
Department: Engineering

The Role

Flowcode is seeking a Senior Site Reliability Engineer (SRE) to work on reliability and infrastructure efforts across our platforms. This role will help grow and drive our infrastructure strategy, operational rigor and observability while building and supporting the systems and tooling required to support Flowcode’s continued growth.

As an individual contributor within our engineering organization, you will develop and operate scalable cloud infrastructure, establish best practices around deployment and reliability, and partner closely with engineering teams to ensure systems are scalable, resilient and observable. 

What You’ll DoReliability & Infrastructure
  • Improve system availability, scalability, and resilience across Flowcode's platforms
  • Own key pieces of our EKS-based infrastructure end-to-end
  • Contribute to incident response and postmortems, turning findings into durable fixes
  • Support engineering teams with infrastructure questions, escalations, and day-to-day unblocking
Cloud & Platform Engineering
  • Manage and scale our core AWS footprint (EKS, VPC, RDS) through Infrastructure as Code (Terraform)
  • Enhance disaster recovery and failover mechanisms to protect mission-critical workloads
  • Collaborate with product engineering to streamline and optimize internal developer experience
CI/CD & Deployment Automation
  • Design and scale deployment pipelines using GitHub Actions
  • Expand GitOps practices and tooling through ArgoCD
  • Facilitate secure delivery with automated validation and progressive rollout strategies
Observability & Monitoring
  • Oversee and optimize the organization's monitoring, logging, and alerting infrastructure
  • Develop high-signal metrics, tracing, and visualization dashboards while minimizing operational noise
  • Establish and monitor Service Level Objectives for managed platform components
QualificationsRequired
  • 4+ years of professional experience across SRE, DevOps, or Platform Engineering domains
  • Technical proficiency in Kubernetes, including cluster troubleshooting and managing controllers or CRDs
  • Advanced Terraform or OpenTofu expertise, encompassing module architecture and production state management
  • Hands-on operational experience with GitOps workflows via ArgoCD and Helm-based deployments
  • Ability to author production-grade code in Go or Python alongside robust shell scripting
  • Mastery of core AWS services, specifically EKS, Networking/VPC, RDS, and IAM
  • Experience maintaining and scaling CI/CD automation using GitHub Actions within collaborative environments
  • Proven track record of leading infrastructure initiatives from initial design through to long-term operation
  • Background in supporting large-scale distributed systems within high-availability production environments
  • Adept at navigating interrupt-driven workflows, balancing strategic project delivery with day-to-day operational support
Preferred
  • Exposure to Crossplane or alternative Kubernetes-native solutions for infrastructure provisioning
  • Deep observability experience utilizing Datadog or Prometheus to engineer SLOs, high-signal dashboards, and intelligent alerting
  • Practical knowledge of modern secrets management frameworks and implementation
  • Experience optimizing cluster efficiency through autoscaling technologies such as Karpenter or Cluster Autoscaler

Flowcode is not for everyone. We hire with a pinhole lens — only those with the rare combination of intellectual horsepower, execution velocity, and uncompromising drive will thrive here. If you are seeking to operate at the highest levels of performance and impact, we want to meet you.

How to Apply

We are an equal opportunity employer and value diversity at our company. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.

A successful candidate’s starting pay will be determined based on the role, job-related skills, experience, qualifications, work location, and market conditions. 

Skills Required

  • 4+ years of professional experience in SRE, DevOps, or Platform Engineering
  • Technical proficiency in Kubernetes, including cluster troubleshooting and managing controllers or CRDs
  • Advanced Terraform or OpenTofu expertise, including module architecture and production state management
  • Hands-on experience with GitOps workflows using ArgoCD and Helm-based deployments
  • Ability to write production-grade code in Go or Python and robust shell scripts
  • Mastery of AWS services including EKS, VPC/networking, RDS, and IAM
  • Experience maintaining and scaling CI/CD automation using GitHub Actions
  • Proven experience leading infrastructure initiatives from design through long-term operation
  • Experience supporting large-scale distributed systems in high-availability production environments
  • Ability to manage interrupt-driven operational workflows while delivering strategic projects
  • Exposure to Crossplane or alternative Kubernetes-native infrastructure provisioning solutions
  • Deep observability experience with Datadog or Prometheus, including SLOs, dashboards, and alerting
  • Practical knowledge of modern secrets management frameworks and implementation
  • Experience with cluster autoscaling technologies such as Karpenter or Cluster Autoscaler
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: New York, NY
110 Employees
Year Founded: 2019

What We Do

Flowcode instantly connects your brand to high-value customers, capturing who they are and how they engage. Each premium-designed Flowcode QR code acts as your brand’s front door, delivering a seamless, on-brand experience and unlocking valuable data. Acquire your best audiences at the lowest cost and turn every connection into measurable results.

Why Work With Us

Founded by the former AOL CEO, Tim Armstrong. We are a team of large company executives, startup founders, engineers, scientists, artists, designers, and creators, all data obsessed. We are focused on building a powerfully diverse workforce, not just because it is the right thing to do but because it expands the power of our team.

Gallery

Gallery

Similar Jobs

PwC Logo PwC

Site Reliability Engineer

Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Hybrid
9 Locations
370000 Employees
124K-280K Annually

CrowdStrike Logo CrowdStrike

Senior Engineer

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
USA
11000 Employees
140K-215K Annually

CrowdStrike Logo CrowdStrike

Senior Site Reliability Engineer

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Hybrid
4 Locations
11000 Employees
140K-215K Annually
Remote or Hybrid
United States
1557 Employees
155K-172K Annually

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account