Senior Software Engineer, Infrastructure & Systems

Posted 15 Hours Ago
Be an Early Applicant
New York City, NY, USA
Hybrid
200K-300K Annually
Senior level
Artificial Intelligence • Big Data • Cloud • Software • Analytics • Infrastructure as a Service (IaaS) • Big Data Analytics
Astronomer empowers data teams to bring mission-critical software, analytics, and AI to life.
The Role
Design, build, and own control-plane systems and APIs that provision, manage, and secure infrastructure for Airflow at scale. Lead architecture, observability, and security efforts, participate in on-call incident response, and mentor engineers while driving operational rigor and lifecycle management across multi-tenant and air-gapped environments.
Summary Generated by Built In

Astronomer empowers data teams to bring mission-critical software, analytics, and AI to life and is the company behind Astro, the industry-leading unified DataOps platform powered by Apache Airflow®. Astro accelerates building reliable data products that unlock insights, unleash AI value, and powers data-driven applications. Trusted by more than 800 of the world's leading enterprises, Astronomer lets businesses do more with their data. To learn more, visit www.astronomer.io.

About This Role

Senior and Staff Software Engineers on our Infrastructure team own the architecture of the systems that stand up, configure, and manage the infrastructure the underlying Airflows run on. This includes the security of those systems and the reliability and observability capabilities that keep Astro running at scale, across both Astronomer's multi-tenant cloud platform and Astro Private Cloud, our self-hosted distribution for regulated organizations running Airflow in their own private clouds or fully isolated data centers.

You'll design and build the control-plane services, the APIs that drive infrastructure lifecycle management, and the control/data-plane interaction patterns that make it possible to bring up and manage that infrastructure — including how those interactions hold up across network topologies from standard private cloud to zero-egress, air-gapped environments.

You'll bring deep, hands-on expertise in distributed systems, API design, and infrastructure lifecycle management, own the technical direction, design, and implementation of key areas of the platform, and mentor engineers earlier in their careers. This posting covers both Senior Software Engineer and Staff Software Engineer, with the level you're hired at depending on experience and demonstrated scope.

What You'll Do

  • Design and own the architecture of the systems that stand up, scale, and manage the infrastructure Airflow runs on, leading these efforts end-to-end (design, implementation, tests, documentation, and rollout).

  • Design and evolve the APIs (REST/gRPC) that drive infrastructure lifecycle management, powering customer-facing UI and programmatic integrations, including data modeling, versioning, and backward compatibility.

  • Own the networking and security posture of these systems, including how they authenticate, authorize, and communicate securely, and drive CVE remediation, image hardening, and Pod Security Standards compliance.

  • Architect observability and traceability across the stack: metrics, logs, and traces that let a request be followed end-to-end across service and network boundaries.

  • Participate in on-call rotation, contributing to incident diagnosis, resolution, and post-mortems for these systems and ensuring remediation items are tracked into sprints.

  • Mentor engineers and raise the bar on design and code review, testing, and operational rigor.

What You'll Bring

  • 5+ years of experience in infrastructure, platform, or systems engineering, with a track record of owning production systems at scale (8+ years typical for Staff).

  • Deep, hands-on expertise with Kubernetes, including building custom Operators/CRDs and managing Helm-based deployments for infrastructure lifecycle workflows.

  • Experience designing APIs (REST or gRPC) that serve as the core engine behind a customer-facing UI or programmatic integrations, including versioning and backward compatibility.

  • Strong systems fundamentals: networking, distributed systems, reliability engineering, and security best practices.

  • Proficiency in Go and/or TypeScript, or a demonstrated ability to pick up new languages quickly.

  • Practical experience designing and/or building observability and distributed tracing for multi-layer systems

  • Demonstrated ability to lead technical design across teams and communicate trade-offs to both engineers and stakeholders.

  • Experience mentoring engineers and improving team-wide practices.

Nice to Have

  • Experience designing or operating multi-tenant control planes.

  • Experience shipping software into air-gapped or highly regulated environments (finance, healthcare, government).

  • Familiarity with Apache Airflow internals or other data-orchestration platforms.

  • Incident command / on-call leadership experience.

The estimated total compensation for this role ranges from $200,000 - $300,000 based on leveling and geography, along with an equity component and a comprehensive benefits package. This range is merely an estimate; actual compensation may deviate from this range based on skills, experience, and qualifications.

#LI-Fulltime #LI-Hybrid

At Astronomer, we value diversity. We are an equal opportunity employer: we do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.

Skills Required

  • 5+ years of experience in infrastructure, platform, or systems engineering
  • Track record of owning production systems at scale (8+ years typical for Staff)
  • Deep, hands-on expertise with Kubernetes, including building custom Operators/CRDs
  • Managing Helm-based deployments for infrastructure lifecycle workflows
  • Experience designing APIs (REST or gRPC), including versioning and backward compatibility
  • Strong systems fundamentals: networking, distributed systems, reliability engineering, and security best practices
  • Proficiency in Go and/or TypeScript, or demonstrated ability to pick up new languages quickly
  • Practical experience designing and/or building observability and distributed tracing for multi-layer systems
  • Demonstrated ability to lead technical design across teams and communicate trade-offs to engineers and stakeholders
  • Experience mentoring engineers and improving team-wide practices
  • Experience designing or operating multi-tenant control planes
  • Experience shipping software into air-gapped or highly regulated environments (finance, healthcare, government)
  • Familiarity with Apache Airflow internals or other data-orchestration platforms
  • Incident command / on-call leadership experience
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: New York, New York
344 Employees
Year Founded: 2018

What We Do

Astronomer empowers data teams to bring mission-critical software, analytics, and AI to life and is the company behind Astro, the industry-leading unified DataOps platform powered by Apache Airflow®. Astro accelerates building reliable data products that unlock insights, unleash AI value, and powers data-driven applications. Trusted by more than 800 of the world's leading enterprises, Astronomer lets businesses do more with their data. To learn more, visit www.astronomer.io. Apache® and Apache Airflow® are either registered trademarks or trademarks of the Apache Software Foundation in the United States and/or other countries. No endorsement by the Apache Software Foundation is implied by the use of these marks. All other trademarks are the property of their respective owners.

Similar Jobs

UL Solutions Logo UL Solutions

Engineer Project Associate - Electric Vehicle Charging

Automotive • Professional Services • Software • Consulting • Energy • Chemical • Renewable Energy
Hybrid
2 Locations
15000 Employees
76K-112K Annually
In-Office or Remote
New York, NY, USA
5800 Employees

Domino Data Lab Logo Domino Data Lab

Senior Quality Engineer

Artificial Intelligence • Machine Learning
Remote or Hybrid
US
200 Employees
145K-175K Annually

Datadog Logo Datadog

Product Marketing Manager

Artificial Intelligence • Cloud • Security • Software • Cybersecurity
Easy Apply
Hybrid
New York, NY, USA
6500 Employees
123K-164K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account