Top Tech Jobs & Startup Jobs

5 Days AgoSaved
In-Office
Penang, Daerah Timor Laut, Penang, MYS
Senior level
Senior level
Software
Lead design and operation of CI/CD and MLOps pipelines, cloud-native infrastructure, and observability. Own incident response, security/compliance, IaC, Kubernetes/Docker cluster management, GPU provisioning for AI workloads, high-availability and disaster recovery, and cross-team platform engineering to improve developer productivity and reliability.
Top Skills: Alibaba CloudAnsibleAWSAzureCi/CdDnsDockerElk/EfkGCPGoGpu ClustersGrafanaHelmHTTPInternal Developer PlatformIso27001KubernetesLinuxLoad BalancingMlopsObservabilityPrometheusPythonSecrets ManagementShellSoc2Tcp/IpTerraformTgiTriton Inference ServerVllmVpcsZero Trust
5 Days AgoSaved
In-Office
Singapore, SGP
Junior
Junior
Software
Monitor GPU cluster, network, storage and environmental systems; respond to alerts and follow runbooks; perform hardware triage and standard remediation; collect diagnostics for escalation; manage incident tickets; perform physical data-center tasks (rack, cabling, hardware swaps); execute shift handoffs and maintain runbooks; assist with deployments and firmware under SME guidance.
Top Skills: BmcDcgmGrafanaJira Service ManagementLinuxNagiosPrometheusServicenow
Reposted 5 Days AgoSaved
In-Office
Singapore, SGP
Mid level
Mid level
Software
Design and build global control-plane services for a GPU/AI cloud: tenant management, quota, metering/billing, RBAC, ticketing, customer portal/API, automation, region-aware scheduling, and operational tooling for detection and recovery.
Top Skills: Api DesignBillingBilling/Quota SystemsCloud Control PlaneDistributed SystemsGoJavaMeteringMulti-Tenant SaasOauthOidcPythonRbacReactRegion-Aware SchedulingVue
Reposted 5 Days AgoSaved
Remote
Cyberjaya, Sepang, Selangor, MYS
Junior
Junior
Software
Provide first-line cloud and NOC support for a multi-region GPU cloud: handle tickets, monitor alerts, triage and resolve issues across GPU instances, VMs/bare metal, networking, storage and billing, maintain runbooks, escalate to L3/SRE when needed, and participate in 24/7 follow-the-sun shift rotations.
Top Skills: Bare MetalContainersCudaGpuLinuxNetworkingPublic CloudStorageVirtual Machines
Reposted 5 Days AgoSaved
In-Office
Singapore, SGP
Senior level
Senior level
Software
Own and resolve complex escalations and platform incidents end-to-end. Troubleshoot GPU compute, networking/SDN, storage, drivers/CUDA, control plane, and billing. Drive root-cause analysis with SRE/Compute/R&D, lead incident response and post-incident reviews, improve runbooks and monitoring, mentor L1/L2, and participate in on-call escalation rotation.
Top Skills: Cloud InfrastructureCudaGpuGpu DriversHpcIncident ManagementLinuxNetworkingSdnSreStorage
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
6 Days AgoSaved
In-Office
San Jose, CA, USA
Senior level
Senior level
Software
Design, implement, and maintain CI/CD and MLOps pipelines, provision and scale cloud-native AI infrastructure (Kubernetes, GPU clusters), enforce IaC practices, implement observability and security frameworks, lead incident response and cross-functional platform engineering to ensure high availability and compliance.
Top Skills: Alibaba CloudAnsibleAWSAzureCi/CdDockerElk/EfkGCPGoGpu ClustersGrafanaHelmIacInternal Developer Platform (Idp)KubernetesLinuxMlopsObservabilityPrometheusPythonShellSreTerraformTgiTriton Inference ServerVllmZero Trust
Reposted 6 Days AgoSaved
Remote
Tydal, Trøndelag, NOR
Mid level
Mid level
Software
The AI Infrastructure Engineer is responsible for architecting and managing AI compute environments using GPU clusters and optimizing performance across networking and storage systems.
Top Skills: AnsibleBeegfsCudaGpu ClustersInfinibandKubernetesLustreNcclRoce V2SlurmTerraformTriton Inference ServerWekaio
Reposted 6 Days AgoSaved
Remote
Tydal, Trøndelag, NOR
Mid level
Mid level
Software
The role involves ensuring reliable system integration and operation across various technical systems in a large-scale AI data center while collaborating with multiple engineering disciplines.
Top Skills: AutomationBessBmsCctvElectric HvElectric LvFire AlarmFire SuppressionGeneratorsHvacPlcPmsScadaSecurity SystemsUps
7 Days AgoSaved
In-Office
San Jose, CA, USA
Senior level
Senior level
Software
Lead end-to-end architecture of AI data center networks, DCI, and global backbone for large GPU clusters. Design underlay and overlay networks, IP/VLAN/VXLAN planning, congestion control (PFC/ECN/INT), and optical transport. Produce HLD/LLD, topology diagrams, standards, and SOPs; engage vendors and drive risk mitigation and network improvements.
Top Skills: BgpClosDwdmEcmpEcnEvpnGpu ClustersIn-Band Network Telemetry (Int)InfinibandNvidiaOptical Transport NetworksOspfPfcRocev2Sdn ControllersSegment Routing (Sr-Mpls)Spine-LeafSrv6VlanVxlan
7 Days AgoSaved
In-Office
Singapore, SGP
Mid level
Mid level
Software
Design and produce visual assets and motion graphics for digital, print, and video campaigns. Ensure brand consistency, collaborate with marketing/product teams, manage multiple projects, prepare final files for production, and maintain an organized asset library.
Top Skills: Adobe Creative SuiteAi ToolsFigma
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account