Senior Engineer, SRE

Posted Yesterday
28 Locations
In-Office or Remote
Senior level
Artificial Intelligence • Information Technology • Software
The Role
Design, build, and operate Zencoder's cloud and Kubernetes infrastructure to improve reliability, security, scalability, cost efficiency and developer tooling. Automate operations, enhance CI/CD and GitOps, operate data systems (OpenSearch/Postgres), improve observability and incident response, and prepare the platform for growth and AI workloads.
Summary Generated by Built In
About Zencoder

At Zencoder.ai, we build and orchestrate AI agents that ship real work - code, research, operations, and more. What started as developer tooling is becoming a platform where people and agents collaborate across knowledge work tasks.

About the role

We’re looking for an Engineer to help build and operate the infrastructure behind Zencoder’s AI-powered products.

You’ll work across our production platform, improving its reliability, security, scalability and cost efficiency. This includes our Kubernetes foundations, cloud infrastructure, networking, data systems and the internal tooling that enables engineers to deploy and operate services confidently.

This is a hands-on engineering role rather than a traditional operations position. You’ll write code, automate infrastructure, investigate production issues and design systems that reduce operational complexity as the company grows.

The exact problems will evolve quickly. You should be comfortable taking ownership of unfamiliar systems, identifying the highest-leverage improvements and moving between immediate production needs and longer-term platform investments.

Example projects include
  • Owning our Kubernetes foundations: building and operating production GKE clusters with reliable networking, ingress, service-to-service communication, workload isolation, autoscaling and deployment patterns.
  • Improving cloud security and networking: evolving our GCP architecture across VPCs, IAM, workload identity, secrets, firewalls, WAF, CDN and other security controls.
  • Building dependable search infrastructure: improving the deployment, scaling, performance and operational reliability of OpenSearch and other data-intensive systems.
  • Reducing infrastructure cost: developing better cost attribution, capacity planning and optimisation across compute, storage, networking, observability and managed cloud services.
  • Making deployments safer: improving CI/CD, GitOps, progressive delivery, automated rollback and the tooling engineers use to deploy and operate their services.
  • Strengthening production reliability: improving observability, alerting, incident response, disaster recovery and the resilience of critical customer-facing systems.
  • Automating operational work: replacing manual procedures with software, infrastructure-as-code and reusable platform capabilities.
  • Preparing the platform for growth: identifying architectural bottlenecks and evolving our infrastructure to support increasing usage, larger customers and new AI workloads.
You may be a fit if
  • You have strong software-engineering skills and regularly write production code.
  • You have experience building and operating infrastructure on GCP, AWS or another major cloud platform.
  • You have hands-on experience with Kubernetes in production.
  • You understand cloud networking and security concepts such as VPCs, IAM, load balancing, firewalls, WAFs, CDNs, DNS and service identity.
  • You have experience with infrastructure-as-code and automated deployment systems.
  • You are comfortable debugging problems across application, infrastructure, networking and data-system boundaries.
  • You have operated distributed systems such as OpenSearch, Elasticsearch, PostgreSQL or similar technologies at scale.
  • Experience deploying or operating large language models with serving frameworks such as vLLM or SGLang is a plus, but not required.
  • You care about reliability, security, developer experience and cost - not just whether infrastructure is technically running.
  • You look for ways to remove operational toil rather than accepting repetitive manual work.
  • You take ownership of important problems and are comfortable working across traditional team boundaries.

Experience with every technology we use is not required. We value strong engineering fundamentals, good judgement and the ability to learn unfamiliar systems quickly.


Why Join Zencoder?
  • Shape the Future of Software Creation: We’re not just improving how developers write code — we’re redefining how ideas turn into reality. By closing the gap between concept and execution, we’re creating tools that will influence every industry that relies on software.
  • Massive Impact, Real Ownership: At Zencoder, you’ll have full visibility into how your work moves the product and the company forward. You’ll ship features that matter, see the immediate impact of your decisions, and get feedback directly from users — fast.
  • ICs Are the Core: Individual Contributors are the highest-status role at Zencoder. Our culture celebrates those who lead by doing — who create momentum, inspire others, and turn ideas into shipped products.
  • High-Caliber Team & Founder: Work alongside exceptional AI and software engineers, and learn directly from Andrew Filev, founder of a unicorn startup, who brings deep expertise in scaling world-class technology companies.
  • Global & Flexible: We hire talent, not coordinates. Work from wherever you’re happiest and most productive — as long as you bring the energy, focus, and results.
  • Aligned Incentives: Our equity plan ensures that when we succeed, you succeed. Your impact compounds as the company grows.

Zencoder is an equal-opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.

Skills Required

  • Strong software engineering skills and regularly write production code.
  • Experience building and operating infrastructure on GCP, AWS, or another major cloud platform.
  • Hands-on production experience with Kubernetes (GKE preferred).
  • Understanding of cloud networking and security concepts (VPCs, IAM, load balancing, firewalls, WAFs, CDN, DNS, service identity).
  • Experience with infrastructure-as-code and automated deployment systems (CI/CD, GitOps).
  • Ability to debug across application, infrastructure, networking, and data-system boundaries.
  • Experience operating distributed systems such as OpenSearch, Elasticsearch, PostgreSQL at scale.
  • Experience deploying or operating large language models with vLLM or similar frameworks.
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
Campbell, California
25 Employees
Year Founded: 2023

What We Do

At Zencoder, we're transforming the landscape of software development by empowering developers with AI coding agents embedded into their workflow that help create high-quality software and accelerate product delivery. Zencoder brings the zen back in coding

Similar Jobs

Pragmatike Logo Pragmatike

Senior Site Reliability Engineer

Information Technology • Software
In-Office or Remote
10 Locations
11 Employees
Remote
26 Locations
1004 Employees
53K-120K Annually
Remote
26 Locations
1004 Employees
54K-150K Annually

Miro Logo Miro

Site Reliability Engineer

Cloud • Information Technology • Internet of Things • Productivity • Software
Remote or Hybrid
28 Locations
2500 Employees

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account