Cloud Architect Internship

Posted 10 Hours Ago
Be an Early Applicant
Hiring Remotely in United States
Remote
Internship
Artificial Intelligence • Cloud • Information Technology • Consulting
The Role
Design and govern a multi-tenant cloud infrastructure (AWS/GCP/Azure) for an e-commerce platform. Optimize capacity and costs for fluctuating traffic and AI agent compute, implement multi-region HA/DR for PostgreSQL 15 and Redis, enforce Terraform IaC governance, ensure cloud security and data residency compliance, and integrate containerized apps (Docker/Traefik) with monitoring (Prometheus, OpenTelemetry).
Summary Generated by Built In
Overview: Multi-Tenant Cloud Infrastructure Design

The Cloud Architect is responsible for designing, managing, and optimizing the entire multi-tenant cloud infrastructure that hosts our integrated supply chain e-commerce platform. While our current stack utilizes Oracle Cloud Free VMs, this role requires expertise in major public clouds (AWS/GCP/Azure) to architect scalable, resilient, cost-optimized, and compliant solutions that handle fluctuating e-commerce traffic, stable internal tool usage, and intensive AI agent compute demands.

Internship Details

Duration: 3 months
Start Date: Immediate
Location: Remote
Stipend: None initially. Based on your first-quarter performance, you may be offered a paid full-time opportunity, or even be absorbed directly by the client as an FTE.

Key Responsibilities & Core Projects

You will design and govern the platform's foundation, ensuring scalability and compliance across all environments.

  • Cloud Architecture Design: Design and evolve the target cloud infrastructure (utilizing AWS, GCP, or Azure best practices) for maximum scalability, security, and high-availability, ensuring the platform can reliably support the MES $\rightarrow$ WMS $\rightarrow$ OMS flow.

  • Capacity Planning & Optimization: Plan and optimize resource allocation to effectively handle unpredictable e-commerce traffic spikes and the specific compute requirements for AI agent training and inference, driving cost efficiency.

  • High Availability & Disaster Recovery (DR): Implement multi-region/multi-AZ high-availability architectures and define comprehensive Disaster Recovery (DR) strategies for all core services, including PostgreSQL 15 and Redis.

  • Infrastructure-as-Code (IaC) Governance: Establish and enforce best practices for Terraform usage, ensuring configuration consistency, security compliance, and auditable infrastructure changes.

  • Security & Compliance: Conduct regular security reviews of cloud configurations (e.g., IAM, VPC/VNet, Storage) and ensure architecture aligns with data residency and compliance requirements relevant to our multi-tenant operations.

  • Service Integration: Architect the network and service mesh overlay that integrates the containerized applications (Docker/Traefik) with external cloud services and the overall monitoring solution (Prometheus, OpenTelemetry).

Required Technologies & Tools

Candidates must possess deep architectural experience with public cloud providers and infrastructure automation:

  • Cloud Providers: Expert-level proficiency in at least one major public cloud (AWS, GCP, or Azure).

  • Infrastructure-as-Code (IaC): Mandatory expertise in Terraform for cloud resource provisioning.

  • Containerization: Deep knowledge of Docker networking, security, and orchestration principles.

  • Networking & Edge: Experience configuring load balancing, service mesh, and ingress controllers (e.g., Traefik).

  • Data & Storage: Architecting scalable database services (PostgreSQL) and object storage (MinIO).

  • Security: Cloud security best practices, IAM policy design, and network segmentation.

AI Agent Focus

You will optimize the cloud layer for our emerging AI capabilities.

  • Compute Optimization: Design elastic and cost-effective compute clusters (e.g., GPU instances) to efficiently handle the variable demands of LLM fine-tuning and multi-agent system orchestration.

  • Data Residency: Architect the data pipeline and storage solutions to ensure that training data and model artifacts adhere to strict data residency requirements across tenants.

Success Metrics & Career Path

Performance will be measured by:

  • Cost Efficiency: Demonstrable reduction in cloud operational costs (FinOps) while maintaining performance.

  • Availability: Achieving defined SLAs/SLOs for infrastructure uptime and performance.

  • Compliance: Successful implementation and auditing of cloud security and data residency controls.

Mentorship Structure: Reports to the Solution Architect or Head of Technology, collaborating closely with the SRE and DevSecOps teams to operationalize cloud strategy.

Skills Required

  • Expert-level proficiency in at least one major public cloud (AWS, GCP, or Azure)
  • Mandatory expertise in Terraform for infrastructure-as-code and governance
  • Deep knowledge of Docker networking, security, and orchestration principles
  • Experience configuring load balancing, service mesh, and ingress controllers (e.g., Traefik)
  • Experience architecting scalable PostgreSQL (PostgreSQL 15) and Redis deployments with HA/DR
  • Familiarity with object storage solutions such as MinIO
  • Cloud security best practices, IAM policy design, and network segmentation
  • Experience integrating monitoring and observability (Prometheus, OpenTelemetry)
  • Ability to design elastic, cost-effective GPU compute clusters for LLM training and inference
  • Experience designing multi-region/multi-AZ high-availability and disaster recovery strategies
  • Experience implementing data residency controls and compliance for multi-tenant environments
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company

What We Do

Metasys is a global management and technology consulting firm that helps organizations improve performance through digital transformation. It develops strategies and technology solutions spanning AI, cloud, emerging technologies, digital engineering, intelligent manufacturing, supply chain, and managed services. The company serves industries including aerospace and defense, automotive, financial services, healthcare, software, retail, and utilities, combining technology, data, and industry expertise to deliver measurable impact.

Similar Jobs

Liftoff Logo Liftoff

Product Analyst

AdTech • Artificial Intelligence • Big Data • Machine Learning • Marketing Tech • Mobile • Software
Easy Apply
Remote
2 Locations
645 Employees
126K-170K Annually

Liberty Mutual Insurance Logo Liberty Mutual Insurance

Inside Sales Representative

Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Remote or Hybrid
10 Locations
40000 Employees
45K-85K Annually

Affirm Logo Affirm

Senior Product Manager

Big Data • Fintech • Mobile • Payments • Financial Services
Easy Apply
Remote
United States
2200 Employees
173K-255K Annually

RTB House Logo RTB House

Field Marketing Manager

AdTech • Artificial Intelligence • Big Data • Digital Media • eCommerce • Machine Learning • Marketing Tech
Remote
United States
1300 Employees
110K-130K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account