Top Tech Jobs & Startup Jobs

Reposted 22 Days AgoSaved
In-Office
Singapore, SGP
Senior level
Senior level
Software
Own and optimize the performance-critical LLM serving runtime: scheduling, batching, KV cache, speculative decoding, long-context and streaming. Tune and operate runtimes (vLLM, Dynamo, TensorRT-LLM, etc.), profile GPU/network/tokenizer bottlenecks, lead model onboarding (parallelism, quantization, context), define runtime playbooks, and partner with SRE and performance teams to deploy production improvements for latency, throughput, and cost efficiency.
Top Skills: Anthropic-Compatible ApiCudaDynamoGoHbmKv CacheNcclOpenai-Compatible ApiPythonSglangTensorrtTensorrt-LlmTgiTokenizerTritonVllm
Reposted 23 Days AgoSaved
In-Office
Singapore, SGP
Senior level
Senior level
Software
The Power Quality Engineer supports global data centers by monitoring and optimizing power quality, reducing equipment failures, and ensuring compliance with international standards.
Top Skills: DranetzExcelFlukeHiokiMatlabPqubePythonSchneider Ion
Reposted 23 Days AgoSaved
In-Office
Singapore, SGP
Senior level
Senior level
Software
The Enterprise Systems Consultant will support daily operations of business systems, manage system-related requests, assist in implementation activities, and enhance business processes.
Top Skills: CRMErp
Reposted 23 Days AgoSaved
In-Office
Austin, TX, USA
108K-185K Annually
Senior level
108K-185K Annually
Senior level
Software
The Finance Business Partner will assist in budgeting, forecasting, financial performance analysis, contract reviews, and stakeholder support while ensuring compliance and managing risk.
Top Skills: Excel
24 Days AgoSaved
In-Office
San Jose, CA, USA
180K-260K Annually
Senior level
180K-260K Annually
Senior level
Software
Design, deploy, and operate production Kubernetes control planes for large GPU clusters. Implement GPU-specific scheduling, CRDs, multi-tenant isolation, BMaaS provisioning, Terraform-based IaC, monitoring, SLI/SLOs, and automated remediation workflows to enable autonomous AIOps-driven recovery and tenant self-service.
Top Skills: AlertmanagerArgocdBare-Metal As A Service (Bmaas)Custom Resource Definitions (Crds)FluxGitopsGoGpu Device PluginGrafanaHelmIronicKubeflowKubernetesMaasMigNvidia Gpu OperatorNvlinkPagerdutyPrometheusPythonRaySlurmTerraform
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
24 Days AgoSaved
In-Office
San Jose, CA, USA
105K-155K Hourly
Junior
105K-155K Hourly
Junior
Software
Front-line SRE covering 8AM–8PM PST monitoring GPU clusters, networking, storage, and sensors. Execute runbooks, perform hardware triage and physical DC tasks (rack, cable, swap), collect diagnostics for escalation, manage tickets (ServiceNow/Jira), update runbooks, and tag incidents to train the AIOps platform.
Top Skills: AiopsBmcDcgmGpuGrafanaJira Service ManagementLinuxNagiosPrometheusServicenow
24 Days AgoSaved
In-Office
San Jose, CA, USA
180K-320K Annually
Senior level
180K-320K Annually
Senior level
Software
Deploy, operate, and optimize high-performance parallel/distributed storage systems for AI training and inference; implement multi-tenant isolation and GPU Direct data paths; instrument telemetry for storage-fault prediction; convert incidents into automated runbook-as-code; plan capacity, firmware, migrations, and DR for GPU clusters.
Top Skills: CephDdn/LustreFioGpu Direct StorageIorLinuxMdtestNfs Over RdmaNvidia CmxNvme SmartNvme-OfRdmaVast DataWeka
24 Days AgoSaved
In-Office
Penang, Daerah Timor Laut, Penang, MYS
Senior level
Senior level
Software
Lead design and operation of CI/CD and MLOps pipelines, cloud-native infrastructure, and observability. Own incident response, security/compliance, IaC, Kubernetes/Docker cluster management, GPU provisioning for AI workloads, high-availability and disaster recovery, and cross-team platform engineering to improve developer productivity and reliability.
Top Skills: Alibaba CloudAnsibleAWSAzureCi/CdDnsDockerElk/EfkGCPGoGpu ClustersGrafanaHelmHTTPInternal Developer PlatformIso27001KubernetesLinuxLoad BalancingMlopsObservabilityPrometheusPythonSecrets ManagementShellSoc2Tcp/IpTerraformTgiTriton Inference ServerVllmVpcsZero Trust
24 Days AgoSaved
In-Office
Singapore, SGP
Junior
Junior
Software
Monitor GPU cluster, network, storage and environmental systems; respond to alerts and follow runbooks; perform hardware triage and standard remediation; collect diagnostics for escalation; manage incident tickets; perform physical data-center tasks (rack, cabling, hardware swaps); execute shift handoffs and maintain runbooks; assist with deployments and firmware under SME guidance.
Top Skills: BmcDcgmGrafanaJira Service ManagementLinuxNagiosPrometheusServicenow
Reposted 24 Days AgoSaved
In-Office
Singapore, SGP
Mid level
Mid level
Software
Design and build global control-plane services for a GPU/AI cloud: tenant management, quota, metering/billing, RBAC, ticketing, customer portal/API, automation, region-aware scheduling, and operational tooling for detection and recovery.
Top Skills: Api DesignBillingBilling/Quota SystemsCloud Control PlaneDistributed SystemsGoJavaMeteringMulti-Tenant SaasOauthOidcPythonRbacReactRegion-Aware SchedulingVue
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account