DevOps Engineer Internship

Posted 17 Hours Ago
Hiring Remotely in United States
Remote
Internship
Artificial Intelligence • Cloud • Information Technology • Consulting
The Role
Build and maintain automated infrastructure and CI/CD pipelines using Terraform, Docker, Traefik, and Makefile. Implement observability (Prometheus, Grafana, Loki, Tempo, OpenTelemetry), backups (pgBackRest/Postgres15), and cloud/Linux administration. Automate deployment and monitoring for AI agent services and support monorepo workflow with SRE and DevSecOps teams.
Summary Generated by Built In
Overview: Infrastructure Automation and CI/CD

The DevOps Engineer is responsible for automating, streamlining, and maintaining the infrastructure and deployment pipelines for our entire integrated platform. You'll ensure rapid, reliable, and consistent delivery of our e-commerce storefront, internal supply chain tools (MES, WMS, OMS), and cutting-edge AI agent services, primarily utilizing Infrastructure-as-Code (IaC) and robust CI/CD practices.

Internship Details

Duration: 3 months
Start Date: Immediate
Location: Remote
Stipend: None initially. Based on your first-quarter performance, you may be offered a paid full-time opportunity, or even be absorbed directly by the client as an FTE.

Key Responsibilities & Core Projects

You will build and maintain the fully automated platform that underpins our entire tech stack.

  • Infrastructure-as-Code (IaC): Design, implement, and manage infrastructure provisioning across all environments using Terraform for our Oracle Cloud Free VMs (or equivalent cloud resources). Ensure infrastructure is auditable, repeatable, and secure.

  • CI/CD Pipeline Management: Set up and maintain the Continuous Integration and Continuous Deployment (CI/CD) pipelines, primarily driven by Makefile and automated testing, for the Node.js/NestJS modular monolith and Next.js frontend applications.

  • Containerization & Orchestration: Manage application containerization using Docker. Define deployment strategies, service discovery, and traffic routing using Traefik for our containerized services.

  • Observability Implementation: Implement, manage, and optimize the comprehensive logging, monitoring, and alerting system using our selected stack: Prometheus, Grafana, Loki, Tempo, and OpenTelemetry. Ensure end-to-end tracing is functional across the complex business flow (MES → WMS → OMS).

  • Resilience & Backups: Collaborate with the SRE team to implement high-availability features and maintain automated backup solutions, including pgBackRest for our PostgreSQL 15 database.

  • Workflow: Maintain the Monorepo structure for streamlined code management and deployment separation across applications (web / admin / API) and domain packages.

Required Technologies & Tools

Candidates must possess mandatory expertise in our core infrastructure and automation stack:

  • Infrastructure-as-Code: Expert proficiency in Terraform.

  • Containerization: Expert proficiency in Docker and deployment strategies (e.g., Traefik, orchestration concepts).

  • CI/CD: Hands-on experience building and maintaining complex pipelines (Makefile, Jenkins/GitHub Actions/GitLab CI concepts).

  • Observability: Strong implementation experience with Prometheus, Grafana, Loki, and OpenTelemetry.

  • Cloud & Linux: Experience with Linux administration and managing cloud resources (Oracle Cloud or equivalent).

AI Agent Focus

You will ensure the scalable and monitored deployment of the AI layer.

  • Deployment Automation: Automate the packaging and deployment pipelines for resource-intensive AI agent services and LLM fine-tuning environments.

  • Resource Monitoring: Set up specific monitoring and alerts to track the performance, resource consumption, and cost of the AI agent compute demands.

Success Metrics & Career Path

Performance will be measured by:

  • Deployment Frequency: Reduction in lead time and increased frequency of stable deployments.

  • Infrastructure Stability: Reliability of provisioned infrastructure (minimal unplanned downtime).

  • Observability Coverage: Completeness and reliability of monitoring, logging, and tracing across all production services.

Mentorship Structure: Reports to the Solution Architect or Head of Technology, working closely with the SRE, DevSecOps, and Backend engineering teams to build a robust platform.

Skills Required

  • Expert proficiency in Terraform (Infrastructure-as-Code)
  • Expert proficiency in Docker and container deployment strategies (Traefik, orchestration concepts)
  • Hands-on CI/CD pipeline experience (Makefile; concepts with Jenkins, GitHub Actions, GitLab CI)
  • Observability implementation experience with Prometheus, Grafana, Loki, Tempo, and OpenTelemetry
  • Linux administration and cloud resource management experience (Oracle Cloud or equivalent)
  • Experience with PostgreSQL 15 backups using pgBackRest
  • Familiarity with Node.js, NestJS, and Next.js application deployment
  • Experience automating packaging/deployment and monitoring for AI agent/LLM fine-tuning workloads
  • Ability to maintain monorepo structure and manage multi-application deployments
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company

What We Do

Metasys is a global management and technology consulting firm that helps organizations improve performance through digital transformation. It develops strategies and technology solutions spanning AI, cloud, emerging technologies, digital engineering, intelligent manufacturing, supply chain, and managed services. The company serves industries including aerospace and defense, automotive, financial services, healthcare, software, retail, and utilities, combining technology, data, and industry expertise to deliver measurable impact.

Similar Jobs

Zscaler Logo Zscaler

Account Executive

Cloud • Information Technology • Security • Software • Cybersecurity
Easy Apply
Remote or Hybrid
Florida, USA
8697 Employees
120K-170K Annually

Zscaler Logo Zscaler

Senior Sales Engineer

Cloud • Information Technology • Security • Software • Cybersecurity
Easy Apply
Remote or Hybrid
3 Locations
8697 Employees
155K-221K Annually

Zscaler Logo Zscaler

Account Executive

Cloud • Information Technology • Security • Software • Cybersecurity
Easy Apply
Remote or Hybrid
Massachusetts, USA
8697 Employees
121K-173K Annually

Zscaler Logo Zscaler

Account Executive

Cloud • Information Technology • Security • Software • Cybersecurity
Easy Apply
Remote or Hybrid
Ohio, USA
8697 Employees
120K-170K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account