Platform Engineer II

Posted 2 Days Ago
Hiring Remotely in United States
Remote
130K-140K Annually
Senior level
Artificial Intelligence • Enterprise Web • Software • PropTech
The Role
Design, build, and operate core platform infrastructure on AWS using Terraform and Kubernetes. Own end-to-end projects, CI/CD pipelines, observability, security best practices, and production reliability while contributing to architecture decisions, documentation, runbooks, and mentoring. Participate in on-call rotation and automate recurring work.
Summary Generated by Built In
About Fexa

Fexa builds facilities-management (CMMS) software for companies with large retail footprints. These companies run their maintenance operations on us. When our platform slows down, our customers feel it first, so we treat reliability and customer impact as the top bar for everything we ship.

Platform Engineering owns AWS, CI/CD, and the platform our product runs on. We are a small, fast team building the function from the ground up: the standards, the automation, and the infrastructure. We run on Terraform, AWS (EKS, Fargate, Aurora Postgres), GitHub Actions and Jenkins, Jira, Confluence, and Slack, and we are AI-first in how we work.

The role

You own components and projects end-to-end, and you are ready to take on larger, more ambiguous problems. Given a goal, you design and deliver the "how" yourself with light oversight. On bigger, less-defined work, you shape the approach and contribute to the design. You will work within the team's technical direction and constraints, bringing a plan forward and aligning with the Director and team members before building.

What you'll do
  • Infrastructure as Code (IaC): Author reusable Terraform modules that other engineers build on, and make changes confidently across our AWS infrastructure within our established conventions.

  • CI/CD Pipelines: Design, optimize, and standardize continuous integration and deployment pipelines to accelerate time-to-market.

  • Ownership: Own components and projects end-to-end, while taking on larger initiatives such as database engine migrations, cronjobs to EventBridge, or data warehouse infrastructure, shaping the approach with direction and support.

  • Architecture: Contribute to Architecture Decisions and help refine our Terraform and observability standards.

  • Observability: Instrument systems, build dashboards and alerts, and begin defining the observability strategy for the systems you work on.

  • Security & Compliance: Apply security by default, including least-privilege IAM, proper secrets handling, and key and certificate rotation.

  • Operations and Reliability: Carry a share of the team's day-to-day operational load alongside project work, including production support, escalations, maintenance, and keeping existing systems healthy. This role is not projects alone.

  • Collaboration: Write the runbooks that let others operate what you build, create documentation for both the team and the broader engineering organization, and help onboard and mentor newer engineers.

On-call

You will take part in the team's on-call rotation as an escalation point of contact. You will step in when first-line remediation and runbooks haven't resolved an incident, and you will drive the issue to resolution.

Education & Qualifications
  • Bachelor’s degree in Computer Science, Engineering, or a related field with 5+ years of relevant experience, or an equivalent combination of education and 8+ years of professional experience in platform engineering, infrastructure, or DevOps.

Required skills
  • 5+ years in platform engineering, infrastructure, or DevOps roles, with a proven track record of building internal tooling and Production support.

  • Strong hands-on experience across AWS infrastructure.

  • Advanced Terraform: composing and versioning reusable modules; remote state with locking and workspaces/backends; meta-arguments and dynamic blocks (for_each, count, dynamic); safe state operations and refactoring (import, moved blocks, targeted state changes); drift detection; and policy/testing in CI (e.g., Terratest, tflint, Sentinel or OPA).

  • Building and debugging CI/CD pipelines independently (GitHub Actions and/or Jenkins).

  • Container orchestration on EKS/Kubernetes and AWS Fargate.

  • GitOps workflows (e.g., ArgoCD).

  • Observability in practice: instrumenting services and building meaningful dashboards and alerts across metrics, logs, and APM.

  • Security fundamentals applied by default: least-privilege IAM, secrets handling, key and certificate rotation.

  • Strong proficiency with python and shell scripting to automate recurring work and build internal tooling.

  • Clear written and verbal communication — able to plan a piece of work, document the reasoning, and align with the team before building.

  • AWS certification, such as Solutions Architect Associate/Professional or SysOps/DevOps Engineer.

What we're looking for
  • Owns components and small projects end-to-end and makes sound trade-off calls at that scope.

  • Ready to take ambiguous, larger problems from idea to delivery with direction. You will shape the approach and be accountable for the result.

  • Contributes to Architecture Decisions and design discussions, and communicates and plans clearly.

  • Proactively automates recurring work rather than absorbing it.

  • A reliable collaborator who reviews peers' pull requests and helps onboard and mentor newer engineers.

  • Flexibility to support scheduled off-hours deployments and critical maintenance windows.

Working with AI

At Fexa, AI isn't just a tool we're testing; it is foundational to how we build and scale. We expect engineers to weave AI deeply into their daily workflows, from architecting infrastructure and automating complex pipelines to rapid troubleshooting and documentation. You aren't just using LLMs; you are authoring custom skills, developing multi-agent workflows, and proactively sharing your prompt engineering breakthroughs to level up the entire team. We value the visible uplift in speed and quality that comes from an AI-augmented engineering culture.

How we work

We run a Kanban board with WIP limits and expect steady flow, not stale tickets. Our Definition of Done includes documentation and written Architecture Decisions, as we record our design decisions. We review each other's pull requests against a shared standard, take structured platform requests from other teams, and keep an eye on cloud cost as we build.

Nice to have
  • Ruby or Postgres performance tuning experience.

  • FinOps or cloud-cost optimization: right-sizing, cost tagging, and reporting.

  • OpenSearch/Elasticsearch operation and migration.

  • Data warehouse / analytics infrastructure on AWS (Redshift, Glue, Step Functions).

  • Disaster recovery and backup/restore testing; exposure to RTO/RPO planning.

  • Experience migrating between CI/CD systems or source-control platforms, such as Jenkins to GitHub Actions or GitLab to GitHub.

  • Networking depth: VPC, DNS, load balancers, and WAF.

  • Leading or co-leading incident response and postmortems.

  • Experience on a small platform/DevOps team moving fast on multiple initiatives concurrently.

  • Building or operating AI/LLM backends on AWS, including Bedrock, agentic/tool-calling services, MCP servers, and vector stores.

  • Hands-on experience with Azure infrastructure.

Skills Required

  • Bachelor's degree in Computer Science, Engineering, or related field with 5+ years experience, or equivalent (8+ years)
  • 5+ years in platform engineering, infrastructure, or DevOps with production support and internal tooling experience
  • Strong hands-on experience with AWS infrastructure
  • Advanced Terraform: reusable modules, remote state/backends, meta-arguments, safe state operations, drift detection, policy/testing (Terratest, tflint, Sentinel or OPA)
  • Designing and debugging CI/CD pipelines (GitHub Actions and/or Jenkins)
  • Container orchestration on EKS/Kubernetes and AWS Fargate
  • GitOps workflows (e.g., ArgoCD)
  • Practical observability: instrumenting services and building dashboards/alerts across metrics, logs, and APM
  • Security fundamentals applied by default: least-privilege IAM, secrets handling, key and certificate rotation
  • Strong proficiency with Python and shell scripting for automation and internal tooling
  • Clear written and verbal communication; able to document designs, write runbooks, and align with team
  • Willingness to participate in on-call rotation and support off-hours deployments/maintenance
  • AWS certification (e.g., Solutions Architect Associate/Professional or SysOps/DevOps Engineer)
  • Ruby or Postgres performance tuning experience
  • FinOps / cloud-cost optimization experience
  • OpenSearch/Elasticsearch operation and migration experience
  • Data warehouse / analytics infrastructure on AWS (Redshift, Glue, Step Functions)
  • Disaster recovery, backup/restore testing, RTO/RPO planning exposure
  • Experience migrating CI/CD or source-control systems (e.g., Jenkins to GitHub Actions)
  • Networking depth: VPC, DNS, load balancers, WAF
  • Leading or co-leading incident response and postmortems
  • Experience building or operating AI/LLM backends on AWS (Bedrock, agentic services, vector stores)
  • Hands-on experience with Azure infrastructure
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
92 Employees
Year Founded: 2011

What We Do

Fexa is an AI-native facilities management SaaS platform built for enterprise, multi-site operations. It provides a highly configurable solution to help facilities and operations teams improve efficiency, reduce costs, and maintain oversight at scale. Serving diverse markets like retail, healthcare, and restaurants, Fexa’s portfolio includes a CMMS, refrigerant management software, and a vendor sourcing network to streamline maintenance and compliance.

Similar Jobs

Samsara Logo Samsara

Senior Software Engineer

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
6 Locations
4000 Employees
131K-198K Annually

Affirm Logo Affirm

Software Engineer

Big Data • Fintech • Mobile • Payments • Financial Services
Easy Apply
Remote
United States
2200 Employees
146K-225K Annually

MeridianLink Logo MeridianLink

Artificial Intelligence Engineer

Software • Financial Services
Remote
US
522 Employees
104K-178K Annually

Elastic Logo Elastic

Software Engineer

Cloud • Security • Software • Generative AI
Remote
United States
3222 Employees
111K-175K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account