Data Platform Infrastructure Manager

Posted Yesterday
Hiring Remotely in United States
Remote or Hybrid
125K-210K Annually
Senior level
Artificial Intelligence • Cloud • Sales • Security • Software • Cybersecurity • Data Privacy
The Role
Leads the data infrastructure platform team responsible for production services supporting batch, streaming, analytics, and machine learning workloads. Oversees people leadership, platform strategy, reliability, security, scalability, cost efficiency, incident response, roadmap execution, and cross-functional collaboration. Builds self-service developer experiences, standardized CI/CD, observability, infrastructure automation, and service ownership practices across Airflow, Flink, Spark, Kafka, Snowflake, Iceberg, AWS, and ML platforms.
Summary Generated by Built In

About SailPoint

SailPoint is the leader in identity security for the cloud enterprise. Our identity security solutions secure and enable thousands of companies worldwide, giving our customers unmatched visibility into the entirety of their digital workforce and ensuring workers have the right access to do their jobs—no more, no less.

Built on a foundation of AI and machine learning, SailPoint Identity Security delivers the right level of access to the right identities and resources at the right time—matching the scale, velocity, and changing needs of today’s cloud-oriented, modern enterprise.

About the role

As the Data Infrastructure Platform Manager, you will lead the team responsible for the runtime and platform layer beneath SailPoint’s batch, streaming, analytics, and machine learning workloads. Your team owns the availability, scalability, security, performance, lifecycle, and cost efficiency of shared services such as Airflow, Flink, Spark on AWS EMR, Kafka, Snowflake, and Iceberg. Data and product engineering teams own the workload-specific pipelines and processing logic that run on those services.

Your day will span people leadership, production operations, technical strategy, and cross-functional execution. You will review service health and incidents, set priorities across operational and roadmap work, coach and unblock engineers, make architectural and investment tradeoffs, and partner with data engineering, Developer Platform, SRE, Observability, Security, and Infrastructure teams. You will make the platform easier to consume through paved roads, CI/CD, configuration as code, observability, automation, and self-service—all while protecting reliability, quality, and cost efficiency in a fast-moving environment.

About the team

The Data Infrastructure Platform team designs, builds, and operates the production-grade data processing infrastructure that powers SailPoint Identity Security. We provide reliable, scalable, and secure data platforms as services so data and product engineers can focus on DAGs, streaming jobs, pipelines, models, and business logic rather than provisioning and operating the underlying infrastructure. The team values engineering and operations excellence, service ownership, practical automation, constructive debate, continuous learning, and a “strong opinions, loosely held” mindset.

Roadmap for success

Success in this role will be measured through the following outcomes:

30 days

  • Build working relationships with team members and key partners across data and product engineering, Data Operations, Developer Platform, SRE, Observability, Security, and Infrastructure.
  • Document and align stakeholders on the team charter, service catalog, ownership boundaries, escalation paths, and the distinction between platform ownership and workload-specific pipeline ownership.
  • Complete an initial assessment of the team’s people, platforms, roadmap, on-call load, incidents, operational risks, capacity, and cloud costs; share a prioritized set of immediate risks and opportunities.
  • Establish a regular operating cadence for team priorities, service health, incidents, roadmap delivery, and cross-team dependencies.

90 days

  • Publish an outcome-oriented 12-month platform roadmap, developed with technical leads and partner teams, that balances reliability, scalability, security, developer experience, and cost.
  • Define the operating model for the team’s highest-criticality services, including named ownership, service-level objectives, observability expectations, on-call practices, capacity planning, lifecycle management, and incident follow-up.
  • Select and begin delivery of the first high-value paved-road or self-service improvement for deploying and operating Airflow DAGs, Flink or Spark jobs, or another priority data workload.
  • Set clear performance expectations and development goals for each team member, identify capability or staffing gaps, and establish a hiring and development plan where needed.
  • Baseline the platform’s key reliability, delivery, toil, utilization, and cost measures so subsequent improvements can be demonstrated with data.

6 months

  • Have service-level objectives, actionable dashboards, alerts, and recurring service reviews in place for the highest-criticality data platform services, with measurable progress against the 90-day reliability baseline.
  • Deliver at least one production self-service or standardized delivery capability that reduces the effort and lead time required for developers to deploy data workloads safely across supported environments.
  • Implement a cost and capacity management program with service-level visibility, accountable owners, prioritized optimization work, and documented efficiency gains.
  • Strengthen incident response, change management, disaster recovery, vulnerability remediation, and operational runbooks; demonstrate reduced recurring toil or faster recovery for priority failure modes.
  • Establish a clear platform approach for supporting machine learning workloads, including appropriate use of AWS SageMaker or equivalent capabilities, model lifecycle needs, and operational guardrails.

1 year

  • Operate the data infrastructure platform as a mature internal product with a clear service catalog, documented support model, paved roads, self-service capabilities, standardized CI/CD, and transparent reliability and cost reporting.
  • Deliver the highest-priority roadmap outcomes and demonstrate measurable year-over-year improvement in platform availability, incident recovery, deployment lead time, developer effort, operational toil, and cost efficiency.
  • Build a healthy, high-performing team with clear ownership, strong technical leadership, meaningful career growth, effective succession coverage, and the capability to execute both roadmap and operational work predictably.
  • Establish a durable multi-year strategy for Airflow, Flink, Spark and AWS EMR, Kafka, Snowflake, Iceberg, and ML infrastructure that anticipates growth, regional expansion, security requirements, and evolving developer needs.
  • Be recognized by partner teams as a responsive, reliable platform organization that enables them to deliver data products and business value faster without assuming the burden of operating shared infrastructure.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent professional experience.
  • Proven experience managing and leading a cloud infrastructure, platform engineering, SRE, or data infrastructure team responsible for business-critical production services.
  • Deep understanding of cloud infrastructure, distributed systems, and production operations, including:
    • 5+ years of production experience with AWS.
    • 3+ years of experience with Kubernetes and containerized workloads.
    • Ability to read, review, and troubleshoot software written in Python, Go, Java, or a comparable language.
    • Experience with infrastructure as code, preferably Terraform, and modern CI/CD or GitOps practices.
    • Strong systems, networking, security, and distributed-systems troubleshooting fundamentals.
  • Substantial hands-on experience operating production data platforms at scale, with depth in several of the following: Apache Airflow, Apache Kafka, Apache Flink, Apache Spark, AWS EMR, Snowflake, and Apache Iceberg. Experience operating these technologies as shared services—not only consuming them—is strongly preferred.
  • Experience with machine learning systems and their operational lifecycle, including practical experience training, evaluating, or deploying ML models and familiarity with AWS SageMaker or an equivalent ML platform. Databricks experience is a plus.
  • Working knowledge of modern LLM capabilities and AI-assisted engineering practices, with sound judgment about responsible use, validation, and production guardrails.
  • Demonstrated success building internal platforms as products, including self-service developer experiences, standardized delivery workflows, CI/CD, observability, and clear service ownership.
  • Strong production-operations discipline, including metrics and observability, SLOs, incident response, capacity planning, disaster recovery, and continuous reliability improvement.
  • Experience managing cloud infrastructure cost, capacity, and performance, with a track record of making measurable efficiency improvements.
  • Excellent leadership, communication, negotiation, and cross-functional collaboration skills, including the ability to align teams with competing goals and clarify ambiguous ownership.
  • Ability to thrive in a fast-paced environment with shifting priorities and incomplete information while preserving engineering rigor, production uptime, and quality.
  • Experience with Agile or similar iterative planning and delivery practices, applied pragmatically to a team that balances roadmap work with operational demand.
  • Passion for engineering excellence, reliability, continuous learning, and developing people.

The Tech Stack

  • Cloud and infrastructure: AWS, Amazon EKS, Kubernetes, Terraform, and configuration as code.
  • Data orchestration and processing: Apache Airflow, Apache Flink, Apache Spark, and AWS EMR.
  • Streaming, storage, and analytics: Apache Kafka, Snowflake, Apache Iceberg, and AWS data services.
  • Delivery and operations: CI/CD, GitOps, metrics, logs, traces, alerting, SLOs, and incident management.
  • Software: Python, Go, Java, or comparable languages.
  • AI and machine learning: AWS SageMaker or equivalent ML platforms, modern LLM capabilities, and AI-assisted engineering tools. Databricks experience is a plus.

SailPoint is an equal opportunity employer, and we welcome everyone to our team. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or veteran status.



Benefits and Compensation listed vary based on the location of your employment and the nature of your employment with SailPoint.


As a part of the total compensation package, this role may be eligible for the SailPoint Corporate Bonus Plan or a role-specific commission, along with potential eligibility for equity participation. SailPoint maintains broad salary ranges for its roles to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect SailPoint’s differing products, industries, and lines of business. Candidates are typically placed into the range based on the preceding factors as well as internal peer equity. We estimate the base salary, for US-based employees, will be in this range from (min-max, USD):

$124,700 - $210,158.00

Base salaries for employees based in other locations are competitive for the employee’s home location.

Benefits Overview

1. Health and wellness coverage: Medical, dental, and vision insurance

2. Disability coverage: Short-term and long-term disability

3. Life protection: Life insurance and Accidental Death & Dismemberment (AD&D)

4. Additional life coverage options: Supplemental life insurance for employees, spouses, and children

5. Flexible spending accounts for health care, and dependent care; limited purpose flexible spending account

6. Financial security: 401(k) Savings and Investment Plan with company matching

7. Time off benefits: Flexible vacation policy

8. Holidays: 8 paid holidays annually

9. Sick leave

10. Parental support: Paid parental leave

11. Employee Assistance Program (EAP) and Care Counselors

12. Voluntary benefits: Legal Assistance, Critical Illness, Accident, Hospital Indemnity and Pet Insurance options

13. Health Savings Account (HSA) with employer contribution

SailPoint is an equal opportunity employer and we welcome all qualified candidates to apply to join our team.  All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, protected veteran status, or any other category protected by applicable law.  

Alternative methods of applying for employment are available to individuals unable to submit an application through this site because of a disability. Contact [email protected] or mail to 11120 Four Points Dr, Suite 100, Austin, TX 78726, to discuss reasonable accommodations.  NOTE: Any unsolicited resumes sent by candidates or agencies to this email will not be considered for current openings at SailPoint.

Skills Required

  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent professional experience
  • Experience managing and leading a cloud infrastructure, platform engineering, SRE, or data infrastructure team responsible for business-critical production services
  • 5+ years of production experience with AWS
  • 3+ years of experience with Kubernetes and containerized workloads
  • Ability to read, review, and troubleshoot software written in Python, Go, Java, or a comparable language
  • Experience with infrastructure as code, preferably Terraform, and modern CI/CD or GitOps practices
  • Strong systems, networking, security, and distributed-systems troubleshooting fundamentals
  • Hands-on experience operating production data platforms at scale
  • Experience operating shared data services including Apache Airflow, Apache Kafka, Apache Flink, Apache Spark, AWS EMR, Snowflake, or Apache Iceberg
  • Experience with machine learning systems and their operational lifecycle, including training, evaluating, or deploying ML models
  • Familiarity with AWS SageMaker or an equivalent ML platform
  • Working knowledge of modern LLM capabilities and AI-assisted engineering practices
  • Experience building internal platforms as products, including self-service developer experiences, standardized delivery workflows, CI/CD, observability, and clear service ownership
  • Production-operations experience with metrics, observability, SLOs, incident response, capacity planning, disaster recovery, and reliability improvement
  • Experience managing cloud infrastructure cost, capacity, and performance with measurable efficiency improvements
  • Excellent leadership, communication, negotiation, and cross-functional collaboration skills
  • Experience working in fast-paced environments with shifting priorities and incomplete information
  • Experience with Agile or similar iterative planning and delivery practices
  • Passion for engineering excellence, reliability, continuous learning, and developing people
  • Databricks experience

SailPoint Compensation & Benefits Highlights

  • Healthcare Strength Coverage spans medical, dental, vision, and mental health, with FSA/HSA options and even pet insurance. The breadth and inclusion of wellness resources indicate a robust core health package.
  • Leave & Time Off Breadth Time off is generous, including flexible or unlimited vacation, paid holidays, sick leave, and volunteer time off. Generous parental leave and family medical leave further expand the time-away options.
  • Retirement Support A 401(k) plan with company matching is offered, complemented by income‑protection benefits such as disability and life insurance. These elements strengthen long‑term financial security for employees.

SailPoint Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Austin, TX
2,461 Employees
Year Founded: 2005

What We Do

At SailPoint, we believe enterprise security must start with identity at the foundation. Today’s enterprise runs on a diverse workforce of not just human but also digital identities—and securing them all is critical. Through the lens of identity, SailPoint empowers organizations to seamlessly manage and secure access to applications and data at speed and scale. Our unified, intelligent, and extensible platform delivers identity-first security, helping enterprises defend against dynamic threats while driving productivity and transformation. Trusted by many of the world’s most complex organizations, SailPoint secures the modern enterprise.

Why Work With Us

Together, we’re redefining identity’s place in the security ecosystem. We love taking on new challenges that seem daunting to others. We hold ourselves to the highest standards and deliver upon our promises to our customers. We bring out the best in each other, and we’re having a lot of fun doing it.

Gallery

Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery

SailPoint Teams

Team
International Culture
Team
Engineering
Team
Professional Services
Team
Sales
About our Teams

SailPoint Offices

Hybrid Workspace

Employees engage in a combination of remote and on-site work.

Typical time on-site: Flexible
HQAustin, TX
Amsterdam, NL
Coyoacán, Ciudad de México
London, GB
Pune, Maharashtra
Toronto, Ontario
Learn more

Similar Jobs

SailPoint Logo SailPoint

Field Marketing Manager

Artificial Intelligence • Cloud • Sales • Security • Software • Cybersecurity • Data Privacy
Remote or Hybrid
9 Locations
2461 Employees
120K-202K Annually

SailPoint Logo SailPoint

Account Executive

Artificial Intelligence • Cloud • Sales • Security • Software • Cybersecurity • Data Privacy
Remote or Hybrid
United States
2461 Employees
109K-165K Annually

SailPoint Logo SailPoint

Account Executive

Artificial Intelligence • Cloud • Sales • Security • Software • Cybersecurity • Data Privacy
Remote or Hybrid
2 Locations
2461 Employees
109K-165K Annually

SailPoint Logo SailPoint

Consultant

Artificial Intelligence • Cloud • Sales • Security • Software • Cybersecurity • Data Privacy
Remote or Hybrid
United States
2461 Employees
95K-160K Annually

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account