Platform Engineer

Posted Yesterday
Be an Early Applicant
Hiring Remotely in Office, Machaze, Manica, MOZ
Remote
Mid level
Cloud • Software
The Role
Build and operate the company’s platform infrastructure, including Terraform-managed cloud resources, Kubernetes clusters, Helm workloads, GitHub Actions pipelines, observability, cost controls, security guardrails, and AI-agent access controls. Support on-call operations, diagnose production incidents, improve reliability, and create internal tooling for engineering teams. The role also involves managing cloud identity, secrets, networking, scaling, capacity, and infrastructure change reviews.
Summary Generated by Built In

About Perry Weather

Perry Weather is the weather safety standard for organizations that operate outdoors. We combine on-site weather hardware with a software platform that automates alerts, siren activations, and safety decisions for more than 3,000 organizations — from the PGA of America, the NFL, and MLB to Turner Construction, thousands of school districts, cities, and golf courses. We've grown 80%+ annually for five consecutive years, and are continuing to grow.

 

We're headquartered in Dallas at The Centrum, in the Oak Lawn neighborhood, where our teams work together every day to help organizations make faster and safer decisions when weather puts people, assets, and operations at risk.

About the Role

We're looking for a Platform Engineer to help build and run the infrastructure the rest of engineering depends on. You'll partner closely with our Lead Platform Engineer, designing and delivering the work together, and own real surface area from the start: infrastructure as code, our Kubernetes environments, delivery pipelines, observability, and the guardrails that keep all of it safe to change.

Engineering is your customer. Success in this role looks like fast service setup, uneventful deploys, and engineers who can diagnose a failing workload without asking for help.

Engineering here also works with AI coding agents daily, which adds to the platform's job: checks that catch generated infrastructure errors before they apply, scoped credentials and audit trails for automated actors, and cost reporting that accounts for agent usage. You'll use agents in your own work and build these controls for everyone else.

What You'll Own

  • Infrastructure as code. Write and maintain the Terraform that defines our cloud footprint, review changes for blast radius, and keep our modules usable by engineers outside the platform team.

  • Kubernetes and workloads. Run our clusters and the Helm-deployed services on them, including resource tuning, scaling behavior, and the failure modes that only show up under load.

  • Delivery pipelines. Build and maintain CI/CD in GitHub Actions so that builds, tests, and deploys are fast and dependable enough that engineers rely on them.

  • Observability and reliability. Improve the instrumentation, dashboards, and alerting that tell us how the platform is behaving, and make alerts specific enough that people act on them.

  • Cost and capacity. Track where our cloud spend goes, find the waste, and add controls that catch expensive changes before they ship.

  • Guardrails. Secrets management, least-privilege access, dependency and image scanning, and policy checks that catch mistakes during review.

  • AI and automated actors. Coding agents and automation are regular consumers of our platform. You'll give them scoped identities and audit trails, add the checks that catch generated infrastructure mistakes before they apply, and track what their usage costs alongside the rest of our spend.

  • On-call and incident support. Share the on-call rotation, respond when the platform degrades, and make the change that prevents a repeat.

Requirements

  • 4+ years in platform, infrastructure, DevOps, or backend engineering, with real ownership of production infrastructure

  • Terraform you've written and maintained. You've reviewed a plan and caught something you didn't want applied, and you know how state drift happens

  • Kubernetes in production. You can take a failing workload, trace it through configuration, networking, and resource limits, and explain what went wrong to the engineer who owns the service

  • CI/CD you've built and debugged, ideally GitHub Actions. You've fixed a pipeline engineers had stopped trusting, and you can separate a failing test from failing infrastructure

  • Working depth in a major cloud, including managed databases, networking, identity, and secrets. We're primarily on Azure, and AWS or GCP experience transfers

  • Scripting in Python, Go, or Bash good enough to automate a manual process end to end and leave it maintainable for someone else

  • Observability work you've done yourself: instrumentation you added, a dashboard other people used, and an alert you tuned because it fired too often

  • On-call experience for production systems, including at least one incident you drove to resolution and wrote up afterward

  • Product instincts about internal tooling. You've built something for other engineers, watched them use it, and changed it based on what you saw

  • Fluency with AI-native engineering tools and agentic workflows (e.g., terminal-native coding agents, LLM-assisted code refactoring and generation) to multiply technical output and speed up development cycles

  • Least-privilege credential design for automated systems: CI service accounts, scoped tokens, short-lived credentials, audit trails. Agents are the newest consumers of that work, and the principles carry over

  • A view on how automated changes should be reviewed. More code and configuration now arrive generated, which puts weight on the checks that run before a merge or an apply, and you should have opinions about what those checks need to cover

A strong plus is experience with policy-as-code, cloud cost management or FinOps practices, API gateway operations, DNS and CDN configuration, Helm chart authoring, managed time-series or high-volume databases, load and performance testing, running AI or ML workloads on shared infrastructure, and exposure to fleets of connected hardware.

Benefits

  • You'll want to come into the office. Our Oak Lawn office isn't just a place to sit — it's where ideas move fast and culture stays strong. The whole team is here Monday through Friday, which means real collaboration, no chasing people down over Slack, and a genuinely fun place to spend your work days.

  • Your wellbeing is covered. Competitive health insurance, 401(k) with employer matching, and a full suite of voluntary benefits, because you shouldn't have to think twice about the basics.

  • Good people, good times. Monthly All-Hands, Office Olympics, happy hours, and more. We take the work seriously and the culture seriously too.

  • You're getting in early, and that matters. We're growing fast, but the biggest opportunities are still ahead. The people joining now will help shape what Perry Weather becomes.

Skills Required

  • 4+ years of experience in platform, infrastructure, DevOps, or backend engineering with production infrastructure ownership
  • Experience writing and maintaining Terraform and reviewing infrastructure plans and state drift
  • Production Kubernetes experience, including troubleshooting configuration, networking, resource limits, and workload failures
  • Experience building and debugging CI/CD pipelines, ideally with GitHub Actions
  • Working knowledge of a major cloud platform, including managed databases, networking, identity, and secrets; Azure experience is primary
  • Scripting ability in Python, Go, or Bash
  • Hands-on observability experience with instrumentation, dashboards, and alert tuning
  • Production on-call experience, including resolving and documenting an incident
  • Experience building and improving internal tooling for other engineers
  • Fluency with AI-native engineering tools and agentic workflows
  • Experience designing least-privilege credentials, scoped tokens, short-lived credentials, and audit trails for automated systems
  • Understanding of review and validation checks for generated code and infrastructure changes
  • Experience with policy-as-code, cloud cost management, or FinOps practices
  • Experience with API gateway operations, DNS, CDN configuration, Helm chart authoring, managed time-series or high-volume databases, load testing, AI/ML workloads, or connected hardware fleets
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Dallas, TX
32 Employees
Year Founded: 2013

What We Do

The modern weather safety platform for athletics, parks & rec, golf, education, and others facing disruptive weather. Powerful, intuitive cloud-based software and connected hardware that protects lives, enforces policy, and minimizes the impact of weather.

Similar Jobs

Kone Logo Kone

Platform Engineer

Logistics • Transportation • Design • Automation • Manufacturing
Remote
Office, Machaze, Manica, MOZ
31273 Employees

Novartis Logo Novartis

Platform Engineer

Biotech • Pharmaceutical
Remote
Office, Machaze, Manica, MOZ
110000 Employees

FTMO Logo FTMO

Platform Engineer

Edtech • Fintech • Financial Services
Remote
Office, Machaze, Manica, MOZ
In-Office or Remote
3 Locations
2267 Employees

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software • Productivity
US
15 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account