AI Native, Tech Ops Engineer

Posted 2 Days Ago
Hiring Remotely in United States
Remote
Senior level
Consumer Web • Digital Media • eCommerce • News + Entertainment • Analytics
We want to prevent one-star experiences for consumers and brands. Our business model is set up to support this mission.
The Role
Build and maintain cloud infrastructure and observability for scalable systems. Monitor, troubleshoot, patch, and automate operations using IaC, CI/CD, Kubernetes, and AI-driven automation. Manage backups, disaster recovery, log pipelines, databases, and message streaming. Collaborate with engineering, security, and business teams to improve reliability, incident response, and infrastructure lifecycle (including Terraform state cleanup and decommissioning).
Summary Generated by Built In

The Tech Ops Engineer serves as a force multiplier for operational excellence, combining infrastructure expertise with AI-powered automation to create self-healing, highly observable, and scalable technology systems. This role harnesses machine intelligence, predictive monitoring, and automated workflows to anticipate issues before they affect users, accelerate incident response, and continuously improve platform performance. Through close partnership with engineering, security, and business teams, the Tech Ops Engineer helps build an AI-native operational environment that maximizes reliability, efficiency, and business agility.

Responsibilities

System Monitoring and Maintenance:

  • Monitor and maintain the organization’s infrastructure, including servers, networks, storage systems, and applications.
  • Perform routine system checks and preventive maintenance to ensure optimal performance and uptime.
  • Respond to system alerts and incidents, diagnosing and resolving issues promptly to minimize downtime.

Troubleshooting and Support:

  • Provide technical support to resolve infrastructure-related issues, working closely with other technical teams.
  • Troubleshoot and resolve hardware, software, and network issues, escalating to higher-level support when necessary.
  • Maintain detailed documentation of issues, solutions, and processes to improve the team’s knowledge base.

System Upgrades and Patching:

  • Plan and execute system upgrades, patches, and configuration changes, ensuring minimal disruption to business operations.
  • Test and validate updates in development environments before deploying them to production.
  • Ensure that all systems comply with security standards and best practices.

Automation and Optimization:

  • Identify opportunities to automate routine tasks and processes, improving operational efficiency and reducing manual workload.
  • Implement scripts, automation tools, and AI skills to streamline system management and monitoring.
  • Continuously evaluate and optimize infrastructure performance, capacity, and resource utilization.

Disaster Recovery and Backup:

  • Support the development and execution of disaster recovery plans to ensure business continuity in case of system failures.
  • Manage backup and restore processes for critical systems and data, ensuring data integrity and availability.
  • Participate in regular disaster recovery testing and drills.

Infrastructure Lifecycle Management:

  • Plan and execute decommissioning of legacy infrastructure, including EC2 instances, VPCs, and load balancers, coordinating Terraform state cleanup and DNS cutover.

Collaboration and Communication:

  • Work closely with development, network, and security teams to ensure alignment and effective communication on infrastructure projects.
  • Provide input on infrastructure design and architecture to support new projects and initiatives.
  • Communicate effectively with non-technical stakeholders, providing updates on system status and issues.

Requirements

Minimum Qualifications & Credentials

  • Bachelor’s degree in Computer Science, Information Technology, or a related field, or equivalent work experience.
  • 5+  years of experience in system administration, or a similar role.
  • 5+ years of professional experience in Linux administration, managing AWS resources, developing CI/CD and server orchestration pipelines, scripting and monitoring.

Hard/Technical Skills

  • You are an expert in:
    • Cloud-based production systems at scale
    • Amazon Web Services (EC2, VPC, EFS, S3, EKS etc.)
  • Production experience running workloads in Kubernetes (EKS), including ArgoCD GitOps deployments
  • Infrastructure as Code tools, primarily Terraform
  • You have experience with:
    • Working in a Python and JavaScript-centric codebase and are familiar with their related best-practices
    • Creating CI/CD pipelines with Jenkins, Concourse or other CI/CD implementation
  • Monitoring tools, like Datadog or Prometheus
  • Scripting for server side automation, auditing, and monitoring
  • Experience maintaining logging, monitoring, and alerting capabilities using OpenSearch, Vector log pipelines, Prometheus, and Kafka.
  • Configuring and managing data sources like PostgreSQL/Aurora RDS, OpenSearch, Redis, and message streaming platforms like Kafka
  • Design and maintain log ingestion pipelines (e.g., Vector → OpenSearch, Vector → Kafka), including index retention, document shape optimization, and failure recovery.
  • Triage and remediate security vulnerabilities (CVEs) across infrastructure components, including container base images, OS packages, and third-party services.
  • The ideal candidate would also have:
    • A working knowledge of modern software practices and technologies such as Agile methodologies
    • Promoting and establishing development standard methodologies for AWS infrastructure-as-code
    • Experience with AWS Well-Architected principles
    • Experience with High Availability implementations
    • Experience around Security and Compliance
    • Exceptional analytical and problem-solving skills

Soft Skills

  • Intellectual curiosity, a willingness to learn new skills and the ability to contribute new ideas
  • Detail-oriented with a focus on maintaining high standards of operational reliability.
  • Adaptable and flexible, able to manage multiple tasks and prioritize effectively in a fast-paced environment.
  • Strong communicator with the ability to collaborate across teams and provide clear, concise technical support.
  • Proactive and self-motivated, with a passion for continuous learning and improvement in technology operations.
  • Learns quickly and using whatever resources  to solve new problems.
  • Obsessed with ensuring an exceptional customer experience- for both internal and external customers. 
  • Stands up for decisions, takes responsibility for results, and shares both good and bad outcomes transparently.
  • Demonstrates a relentless focus on results with a commitment to deliver; 
  • Takes decisive action, and confidently changes course if unsuccessful.
  • Displays a growth mindset to continually improve; encourages everyone around them to be tenacious and never settle.
  • Constantly seeks feedback to improve; Focuses on solving issues through teamwork, and collaboration
  • Acts with urgency; delivers top results in hours and days instead of weeks and months.
  • Relentless in their pursuit of success and possessing the willpower to embrace challenges as opportunities.

BenefitsWhy You’ll Love Working Here

At ConsumerAffairs, your voice matters. We foster a collaborative environment where you’re encouraged to take initiative, experiment boldly, and grow professionally. We're committed to work-life harmony, career development, and celebrating wins together.

  • Health Care Plan (Medical, Dental & Vision)
  • Retirement Plan (401k)
  • Life Insurance (Basic, Voluntary & AD&D)
  • Paid Time Off (Vacation, Sick & Public Holidays)
  • Family Leave (Maternity, Paternity)
  • Short Term & Long Term Disability
  • Training & Development

Skills Required

  • Bachelor's degree in Computer Science, Information Technology, or equivalent experience
  • 5+ years experience in system administration or similar role
  • 5+ years professional Linux administration experience
  • 5+ years managing AWS resources (EC2, VPC, EFS, S3, EKS etc.)
  • Expertise with cloud-based production systems at scale
  • Production experience running workloads in Kubernetes (EKS), including ArgoCD GitOps deployments
  • Infrastructure as Code experience, primarily Terraform
  • Experience working in Python and JavaScript-centric codebases
  • Creating CI/CD pipelines with Jenkins, Concourse, or similar tools
  • Experience with monitoring tools such as Datadog or Prometheus
  • Scripting for server-side automation, auditing, and monitoring
  • Maintaining logging, monitoring, and alerting using OpenSearch, Vector, Prometheus, and Kafka
  • Configuring and managing PostgreSQL/Aurora RDS, OpenSearch, Redis, and Kafka
  • Design and maintain log ingestion pipelines (Vector -> OpenSearch/Vector -> Kafka), index retention and failure recovery
  • Triage and remediate security vulnerabilities (CVEs) across infrastructure components and container base images
  • Experience with Agile methodologies and AWS Well-Architected principles
  • Experience with high availability implementations, security, and compliance
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Tulsa, OK
150 Employees
Year Founded: 1998

What We Do

ConsumerAffairs is a rapidly growing online marketplace where each month millions of consumers research purchases, connect with brands, transact, write reviews and stay up to date on important consumer news. Brands utilize our software-as-a-service platform to connect with customers, collect reviews and generate sales. ConsumerAffairs has a creative, driven and fast-paced entrepreneurial environment. We are looking for teammates that want to win, are self-motivated, high performing and who yearn to build something big.

Why Work With Us

ConsumerAffairs is built on culture, innovation and hard work. We cultivate a fun, creative and dynamic entrepreneurial environment. Our team is energetic, passionate and - dare we say - brilliant.

Gallery

Gallery

Similar Jobs

MetLife Logo MetLife

LTC Care Coordinator - 12/9/24

Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Remote or Hybrid
United States
43000 Employees

MetLife Logo MetLife

Customer Care Advocate Cary 11.25.24

Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Remote or Hybrid
United States
43000 Employees

MetLife Logo MetLife

Call Center Supervisor

Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Remote or Hybrid
United States
43000 Employees
67K-67K Annually

Dynatrace Logo Dynatrace

Solutions Engineer

Artificial Intelligence • Big Data • Cloud • Information Technology • Software • Big Data Analytics • Automation
Remote or Hybrid
Phoenix, AZ, USA
5600 Employees
130K-160K Annually

Similar Companies Hiring

Scotch Thumbnail
Artificial Intelligence • eCommerce • Fintech • Payments • Retail • Software • Analytics
US
35 Employees
Milestone Systems Thumbnail
Artificial Intelligence • Security • Software • Analytics • Big Data Analytics
Lake Oswego, OR
1500 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account