The DevOps Engineer is responsible for automating, streamlining, and maintaining the infrastructure and deployment pipelines for our entire integrated platform. You'll ensure rapid, reliable, and consistent delivery of our e-commerce storefront, internal supply chain tools (MES, WMS, OMS), and cutting-edge AI agent services, primarily utilizing Infrastructure-as-Code (IaC) and robust CI/CD practices.
Internship Details
Duration: 3 months
Start Date: Immediate
Location: Remote
Stipend: None initially. Based on your first-quarter performance, you may be offered a paid full-time opportunity, or even be absorbed directly by the client as an FTE.
You will build and maintain the fully automated platform that underpins our entire tech stack.
Infrastructure-as-Code (IaC): Design, implement, and manage infrastructure provisioning across all environments using Terraform for our Oracle Cloud Free VMs (or equivalent cloud resources). Ensure infrastructure is auditable, repeatable, and secure.
CI/CD Pipeline Management: Set up and maintain the Continuous Integration and Continuous Deployment (CI/CD) pipelines, primarily driven by Makefile and automated testing, for the Node.js/NestJS modular monolith and Next.js frontend applications.
Containerization & Orchestration: Manage application containerization using Docker. Define deployment strategies, service discovery, and traffic routing using Traefik for our containerized services.
Observability Implementation: Implement, manage, and optimize the comprehensive logging, monitoring, and alerting system using our selected stack: Prometheus, Grafana, Loki, Tempo, and OpenTelemetry. Ensure end-to-end tracing is functional across the complex business flow (MES → WMS → OMS).
Resilience & Backups: Collaborate with the SRE team to implement high-availability features and maintain automated backup solutions, including pgBackRest for our PostgreSQL 15 database.
Workflow: Maintain the Monorepo structure for streamlined code management and deployment separation across applications (web / admin / API) and domain packages.
Candidates must possess mandatory expertise in our core infrastructure and automation stack:
Infrastructure-as-Code: Expert proficiency in Terraform.
Containerization: Expert proficiency in Docker and deployment strategies (e.g., Traefik, orchestration concepts).
CI/CD: Hands-on experience building and maintaining complex pipelines (Makefile, Jenkins/GitHub Actions/GitLab CI concepts).
Observability: Strong implementation experience with Prometheus, Grafana, Loki, and OpenTelemetry.
Cloud & Linux: Experience with Linux administration and managing cloud resources (Oracle Cloud or equivalent).
AI Agent Focus
You will ensure the scalable and monitored deployment of the AI layer.
Deployment Automation: Automate the packaging and deployment pipelines for resource-intensive AI agent services and LLM fine-tuning environments.
Resource Monitoring: Set up specific monitoring and alerts to track the performance, resource consumption, and cost of the AI agent compute demands.
Success Metrics & Career Path
Performance will be measured by:
Deployment Frequency: Reduction in lead time and increased frequency of stable deployments.
Infrastructure Stability: Reliability of provisioned infrastructure (minimal unplanned downtime).
Observability Coverage: Completeness and reliability of monitoring, logging, and tracing across all production services.
Mentorship Structure: Reports to the Solution Architect or Head of Technology, working closely with the SRE, DevSecOps, and Backend engineering teams to build a robust platform.
Skills Required
- Expert proficiency in Terraform (Infrastructure-as-Code)
- Expert proficiency in Docker and container deployment strategies (Traefik, orchestration concepts)
- Hands-on CI/CD pipeline experience (Makefile; concepts with Jenkins, GitHub Actions, GitLab CI)
- Observability implementation experience with Prometheus, Grafana, Loki, Tempo, and OpenTelemetry
- Linux administration and cloud resource management experience (Oracle Cloud or equivalent)
- Experience with PostgreSQL 15 backups using pgBackRest
- Familiarity with Node.js, NestJS, and Next.js application deployment
- Experience automating packaging/deployment and monitoring for AI agent/LLM fine-tuning workloads
- Ability to maintain monorepo structure and manage multi-application deployments
What We Do
Metasys is a global management and technology consulting firm that helps organizations improve performance through digital transformation. It develops strategies and technology solutions spanning AI, cloud, emerging technologies, digital engineering, intelligent manufacturing, supply chain, and managed services. The company serves industries including aerospace and defense, automotive, financial services, healthcare, software, retail, and utilities, combining technology, data, and industry expertise to deliver measurable impact.






