The Cloud Architect is responsible for designing, managing, and optimizing the entire multi-tenant cloud infrastructure that hosts our integrated supply chain e-commerce platform. While our current stack utilizes Oracle Cloud Free VMs, this role requires expertise in major public clouds (AWS/GCP/Azure) to architect scalable, resilient, cost-optimized, and compliant solutions that handle fluctuating e-commerce traffic, stable internal tool usage, and intensive AI agent compute demands.
Internship Details
Duration: 3 months
Start Date: Immediate
Location: Remote
Stipend: None initially. Based on your first-quarter performance, you may be offered a paid full-time opportunity, or even be absorbed directly by the client as an FTE.
You will design and govern the platform's foundation, ensuring scalability and compliance across all environments.
Cloud Architecture Design: Design and evolve the target cloud infrastructure (utilizing AWS, GCP, or Azure best practices) for maximum scalability, security, and high-availability, ensuring the platform can reliably support the MES $\rightarrow$ WMS $\rightarrow$ OMS flow.
Capacity Planning & Optimization: Plan and optimize resource allocation to effectively handle unpredictable e-commerce traffic spikes and the specific compute requirements for AI agent training and inference, driving cost efficiency.
High Availability & Disaster Recovery (DR): Implement multi-region/multi-AZ high-availability architectures and define comprehensive Disaster Recovery (DR) strategies for all core services, including PostgreSQL 15 and Redis.
Infrastructure-as-Code (IaC) Governance: Establish and enforce best practices for Terraform usage, ensuring configuration consistency, security compliance, and auditable infrastructure changes.
Security & Compliance: Conduct regular security reviews of cloud configurations (e.g., IAM, VPC/VNet, Storage) and ensure architecture aligns with data residency and compliance requirements relevant to our multi-tenant operations.
Service Integration: Architect the network and service mesh overlay that integrates the containerized applications (Docker/Traefik) with external cloud services and the overall monitoring solution (Prometheus, OpenTelemetry).
Candidates must possess deep architectural experience with public cloud providers and infrastructure automation:
Cloud Providers: Expert-level proficiency in at least one major public cloud (AWS, GCP, or Azure).
Infrastructure-as-Code (IaC): Mandatory expertise in Terraform for cloud resource provisioning.
Containerization: Deep knowledge of Docker networking, security, and orchestration principles.
Networking & Edge: Experience configuring load balancing, service mesh, and ingress controllers (e.g., Traefik).
Data & Storage: Architecting scalable database services (PostgreSQL) and object storage (MinIO).
Security: Cloud security best practices, IAM policy design, and network segmentation.
You will optimize the cloud layer for our emerging AI capabilities.
Compute Optimization: Design elastic and cost-effective compute clusters (e.g., GPU instances) to efficiently handle the variable demands of LLM fine-tuning and multi-agent system orchestration.
Data Residency: Architect the data pipeline and storage solutions to ensure that training data and model artifacts adhere to strict data residency requirements across tenants.
Success Metrics & Career Path
Performance will be measured by:
Cost Efficiency: Demonstrable reduction in cloud operational costs (FinOps) while maintaining performance.
Availability: Achieving defined SLAs/SLOs for infrastructure uptime and performance.
Compliance: Successful implementation and auditing of cloud security and data residency controls.
Mentorship Structure: Reports to the Solution Architect or Head of Technology, collaborating closely with the SRE and DevSecOps teams to operationalize cloud strategy.
Skills Required
- Expert-level proficiency in at least one major public cloud (AWS, GCP, or Azure)
- Mandatory expertise in Terraform for infrastructure-as-code and governance
- Deep knowledge of Docker networking, security, and orchestration principles
- Experience configuring load balancing, service mesh, and ingress controllers (e.g., Traefik)
- Experience architecting scalable PostgreSQL (PostgreSQL 15) and Redis deployments with HA/DR
- Familiarity with object storage solutions such as MinIO
- Cloud security best practices, IAM policy design, and network segmentation
- Experience integrating monitoring and observability (Prometheus, OpenTelemetry)
- Ability to design elastic, cost-effective GPU compute clusters for LLM training and inference
- Experience designing multi-region/multi-AZ high-availability and disaster recovery strategies
- Experience implementing data residency controls and compliance for multi-tenant environments
What We Do
Metasys is a global management and technology consulting firm that helps organizations improve performance through digital transformation. It develops strategies and technology solutions spanning AI, cloud, emerging technologies, digital engineering, intelligent manufacturing, supply chain, and managed services. The company serves industries including aerospace and defense, automotive, financial services, healthcare, software, retail, and utilities, combining technology, data, and industry expertise to deliver measurable impact.









