Principal Dev-Ops Architect

Posted 6 Days Ago
Be an Early Applicant
Hiring Remotely in Toronto, CAN
In-Office or Remote
Expert/Leader
Other
The Role
Principal-level individual contributor architect responsible for designing and operating a global AWS SaaS platform. Owns Terraform infrastructure, CI/CD, Kubernetes, observability, SLOs, reliability, release management, security controls, and AI/ML platform operations. Establishes engineering standards and reference implementations, governs LLMOps, and ensures HIPAA and SOC 2 Type 2 readiness through secrets management, supply-chain security, audit controls, and FinOps.
Summary Generated by Built In

This is a remote position.

Senior technical authority for a cloud platform running global, multi-tenant SaaS services. This is a hands-on individual-contributor architect role — not people management. The architect designs the platform, sets standards and reference implementations other teams build on, and still writes Terraform, builds CI/CD pipelines, and stands up the AI/ML platform personally. Defines how reliability is measured against SLOs, how releases ship, and how the AI/ML platform is built and governed while keeping the environment HIPAA-compliant and SOC 2 Type 2 audit-ready. Influences products, software, and QA through architecture and example.

Core Responsibilities

     Own platform architecture and technical roadmap for infrastructure, deployment, observability, and the AI/ML platform.

     Set engineering standards, patterns, and golden paths for Infrastructure as Code (IaC), CI/CD, and AI tooling; drive adoption through reference implementations and architecture reviews.

     Manage all cloud infrastructure as code in Terraform — reusable modules, remote state, peer-reviewed PRs, drift detection, and automated plan/apply in CI/CD.

     Enforce policy-as-code (OPA, Sentinel, or equivalent) so infrastructure changes meet security and cost guardrails before merge.

     Design and deploy AWS infrastructure across dev, UAT, staging, and production for performance, availability, recoverability, and security (CIS Critical Security Controls).

     Build and operate CI/CD pipelines for large-scale applications on AWS; own release management, rollback, blue/green, canary, and release gates.

     Package and run containerized workloads on Docker and Kubernetes (EKS).

     Lead the SLI/SLO/SLA program and modern observability using OpenTelemetry; drive down MTTD and MTTR; lead blameless post-incident reviews and participate in on-call.

     Provision and operate the AI/ML platform — Anthropic Claude via AWS Bedrock and internal MCP services — all managed as IaC.

     Build LLMOps practices: prompt versioning, evaluation pipelines, token cost attribution, guardrails, and audit logging of agent actions; enforce the PHI data boundary to BAA-covered providers only.

     Operate and evidence the platform controls required for SOC 2 Type 2 and HIPAA; own secrets management, supply-chain security (SBOM, image and dependency scanning), and FinOps.

Required Qualifications

     Bachelor's degree in Software Engineering or equivalent combination of technical education and work experience.

     10+ years in SRE / DevOps / Platform Engineering delivering CI/CD, REST API deployment, containerization, IaaS/PaaS, data pipelines, and application observability — including time at a senior IC or architect level (Staff, Principal, or Architect).

     Proven technical authority across teams: sets architecture and standards and influences delivery through expertise and example rather than direct management.

     Demonstrated experience driving adoption of a new practice or platform (IaC, CI/CD overhaul, or an AI/ML platform) across multiple teams.

     Hands-on Terraform, including reusable modules other teams consume via self-service, remote state, and change management in a CI/CD pipeline.

     Building and operating CI/CD pipelines for large-scale applications on AWS (GitHub Actions, Jenkins, GitLab, or AWS-native).

     Running containerized workloads on Docker and Kubernetes.

     Monitoring and troubleshooting using cloud-native tooling and OpenTelemetry.

     Linux system administration, Unix scripting, and automation.

     Experience working in a HIPAA / HITECH / HITRUST / PHI / PII or PCI DSS environment.



Skills Required

  • Bachelor's degree in Software Engineering or equivalent technical education and work experience
  • 10+ years of experience in SRE, DevOps, or Platform Engineering
  • Experience delivering CI/CD, REST API deployment, containerization, IaaS/PaaS, data pipelines, and application observability
  • Experience at a senior individual-contributor or architect level, such as Staff, Principal, or Architect
  • Ability to set architecture and standards and influence delivery without direct management
  • Experience driving adoption of IaC, CI/CD transformation, or an AI/ML platform across multiple teams
  • Hands-on Terraform experience with reusable modules, remote state, self-service, and CI/CD change management
  • Experience building and operating CI/CD pipelines for large-scale AWS applications
  • Experience with GitHub Actions, Jenkins, GitLab, or AWS-native CI/CD
  • Experience running containerized workloads on Docker and Kubernetes
  • Experience monitoring and troubleshooting with cloud-native tooling and OpenTelemetry
  • Linux system administration, Unix scripting, and automation experience
  • Experience in HIPAA, HITECH, HITRUST, PHI, PII, or PCI DSS environments
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Toronto
115 Employees
Year Founded: 2003

What We Do

CDIT, headquartered in Slidell, LA, has provided technical services for both commercial and Federal customers for the past 18 years. We deliver high-value services with our Agile integrated approach, consisting of Lean-Agile frameworks, process maturity, best practices combined with information security and quality management standards. This integrated approach is paired with the principles of accountability, collaboration, and delivery established our core CDIT execution model. This model allows us to successfully deliver and perform on small to large-scale programs remotely and on-site. CMMI III DEV | ISO 9001:2015 | ISO 27001:2015

Similar Jobs

Square Logo Square

Software Engineer

eCommerce • Fintech • Hardware • Payments • Software • Financial Services
Remote or Hybrid
8 Locations
12000 Employees
185K-327K Annually

Block Logo Block

Software Engineer

Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
In-Office or Remote
8 Locations
12000 Employees
185K-327K Annually

SailPoint Logo SailPoint

Sales Representative

Artificial Intelligence • Cloud • Sales • Security • Software • Cybersecurity • Data Privacy
Remote or Hybrid
Canada
2461 Employees
84K-120K Annually

Square Logo Square

Senior Manager, Business Development

eCommerce • Fintech • Hardware • Payments • Software • Financial Services
Remote or Hybrid
8 Locations
12000 Employees
149K-248K Annually

Similar Companies Hiring

T-Mobile Thumbnail
Other • Utilities
Bellevue, WA
89016 Employees
Rosendin Thumbnail
Other • Manufacturing
San Jose, CA
6219 Employees
OmniCable Thumbnail
Other
Houston, Texas
815 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account