Top Tech Jobs & Startup Jobs

Reposted 9 Days AgoSaved
In-Office
San Jose, CA, USA
180K-260K Annually
Senior level
180K-260K Annually
Senior level
Software
Designs and operates multi-tenant SDN, VPC overlays, and high-performance InfiniBand or RoCE GPU network fabrics. Responsibilities include tenant isolation, cross-region routing, congestion control, network automation, change management, fault localization, architecture standards, security boundaries, and observability for a large-scale multi-region GPU cloud.
Top Skills: EcnEvpnInfinibandL2 NetworkingL3 NetworkingNetwork AutomationPfcRdmaRoceSdnSdn ControllersVpcVxlan
Reposted 9 Days AgoSaved
In-Office
Singapore, SGP
Senior level
Senior level
Software
Design, tune, and maintain GPU hardware and fabric integrations for containerized AI workloads. Optimize host networking (RDMA, SR-IOV, RoCEv2, InfiniBand), implement GPU slicing (MIG, vGPU), build DCGM-based remediation pipelines, profile kernel/driver/runtime (CUDA, NCCL), and collaborate with scheduling/storage teams. Lead investigations, enforce bare-metal provisioning and firmware/OS standards, mentor engineers, and document the AI hardware stack.
Top Skills: Amd GpusAnsibleCCi/CdCudaDcgmGoInfinibandKubernetesKubernetes Device PluginKubernetes OperatorsLinux KernelMigNcclNvidia A100Nvidia H100RdmaRocev2Sr-IovTerraformVgpu
Reposted 9 Days AgoSaved
In-Office
San Jose, CA, USA
145K-260K Annually
Senior level
145K-260K Annually
Senior level
Software
Architect and scale high-cardinality observability and telemetry infrastructure for AI cloud platforms. Build Kubernetes exporters, eBPF diagnostics, hardware and network monitoring, dashboards, alerting, and GPU/network utilization billing pipelines. Establish telemetry standards for AI workloads, support proactive remediation, collaborate with GPU and scheduling teams, lead architecture reviews, and mentor engineers. The role requires expertise in Go, Kubernetes, Prometheus/OpenTelemetry, Linux performance tuning, eBPF, AI hardware metrics, and large-scale cloud or HPC telemetry systems.
Top Skills: BccEbpfGoIpmiKubernetesLinuxMimirNvidia DcgmOpentelemetryPrometheusRedfishThanosVictoriametrics
Reposted 9 Days AgoSaved
In-Office
Singapore, SGP
Senior level
Senior level
Software
Build and maintain cycle-approximate and analytical full-chip performance models (tile array, NoC, memory, IO) in C++/SystemC and Python; analyze LLM/CNN workloads and rooflines; model traffic, contention, memory QoS; run architectural sweeps and provide recommendations; collaborate with compiler and design teams; validate models against RTL/emulation/silicon and deliver reporting/dashboards for leadership.
Top Skills: 3D-DramAi CompilerC++CnnFpga/EmulationHbmLlmNocPythonQuantizationRtlSystemcTransformers
Reposted 9 Days AgoSaved
In-Office
Singapore, SGP
Mid level
Mid level
Software
Design and produce visual assets and motion graphics for digital, print, and video campaigns. Ensure brand consistency, collaborate with marketing/product teams, manage multiple projects, prepare final files for production, and maintain an organized asset library.
Top Skills: Adobe Creative SuiteAi ToolsFigma
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
Reposted 9 Days AgoSaved
In-Office
Austin, TX, USA
Entry level
Entry level
Software
Design and build distributed systems for AI workloads, improve reliability and efficiency, and develop infrastructure automation tools.
Top Skills: C++GoKubernetesPythonRust
Reposted 9 Days AgoSaved
In-Office
San Jose, CA, USA
180K-260K Annually
Senior level
180K-260K Annually
Senior level
Software
Architect and build a highly available Kubernetes control plane for an AI-native cloud platform managing thousands of clusters. Develop Operators and CRDs, optimize etcd, implement multi-tenant security, enable zero-downtime upgrades and federation, and integrate with GPU orchestration systems. Lead distributed systems architecture, reliability engineering, design reviews, and team mentorship while preventing control-plane failures from affecting customer workloads.
Top Skills: Ci/CdCluster ApiCustom Resource Definitions (Crds)EtcdGoKubernetesKubernetes OperatorsTerraform
Reposted 9 Days AgoSaved
In-Office
San Jose, CA, USA
180K-260K Annually
Senior level
180K-260K Annually
Senior level
Software
Own end-to-end reliability for a customer-facing GPU cloud serving workloads across 100–10,000 GPUs. Responsibilities include operating Kubernetes GPU clusters, managing multi-tenant isolation, automating bare-metal provisioning, defining SLIs/SLOs/SLAs, handling incidents and customer communications, building observability with Prometheus and Grafana, implementing infrastructure as code and GitOps, and developing automated remediation and self-service operational tooling.
Top Skills: AlertmanagerArgocdFluxGitopsGoGpu Time-SlicingGrafanaHelmIronicKubernetesMaasMigNvidia Device PluginNvidia Gpu OperatorNvlinkPagerdutyPrometheusPythonTerraform
Reposted 9 Days AgoSaved
In-Office
Singapore, SGP
Expert/Leader
Expert/Leader
Software
Lead a team in verifying complex IC designs, develop verification strategies, and ensure high-quality silicon through advanced methodologies and collaboration.
Top Skills: SvaSystemverilogUvmVerilog
Reposted 9 Days AgoSaved
In-Office
Singapore, SGP
Senior level
Senior level
Software
The Enterprise Systems Consultant will support daily operations of business systems, manage system-related requests, assist in implementation activities, and enhance business processes.
Top Skills: CRMErp
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account