Bitdeer Group
Jobs at Bitdeer Group
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Recently posted jobs
Software
As a Senior Front-End Engineer at Bitdeer, you will implement RTL in Verilog for CPU/NPU cores, collaborate on architecture design, and assist with verification and debugging throughout development stages.
Software
Lead a team in verifying complex IC designs, develop verification strategies, and ensure high-quality silicon through advanced methodologies and collaboration.
Software
Develop and execute test plans for digital design verification, manage regressions and coverage, and collaborate with design teams to ensure design correctness.
Software
Integrate and customize a commercial RISC-V core soft IP, perform RTL (Verilog/SystemVerilog) front-end and control-path modifications, tailor microarchitecture for PPA, support ASIC front-end tasks (synthesis, STA, power, floorplanning), collaborate across performance/middle-end/back-end teams, and contribute to pre- and post-silicon verification and debugging.
Software
Design and implement Verilog RTL for PCIe-related IP blocks, drive verification closure and functional coverage, assist pre- and post-silicon debug, and collaborate with verification and software teams to deliver robust SoC/IP solutions.
Software
Operate and harden Kubernetes-based AI/MaaS production environments; define SLOs, alerting, dashboards, and runbooks; improve rollout safety with canaries and fast rollbacks; drive GPU capacity planning; automate operations with Helm, Argo CD, operators and scripts; and partner to debug incidents from API edge to model workers.
Software
Execute CIP and CDD for retail and institutional accounts; screen customers and counterparties against global sanctions, PEPs, and adverse media; use blockchain forensic tools to trace wallets; manage supplier/third-party verifications via ticketing system; collaborate with Tech and Legal on automated workflows; maintain audit-ready due diligence records.
Software
Manage production control (PC) and material control (MC) functions: convert sales forecasts to production plans, release work orders, run MRP/BOM calculations, track production and inventory, resolve material and production issues, coordinate cross-functional teams, conduct stocktakes, and produce performance reports to ensure on-time delivery and inventory optimization.
Software
Lead end-to-end delivery of multi-region AI cloud infrastructure and platform software projects, coordinating cross-functional teams, vendors, and leadership to meet timelines, budgets, and quality standards for large-scale GPU deployments and cloud platform releases.
Software
Design, integrate, and operate distributed/parallel file systems for large-scale GPU AI training and inference. Own provisioning, mounting, multi-tenant isolation, quotas, lifecycle, performance tuning, multi-region architecture, golden images/ drivers/registries, monitoring, capacity planning, and runbooks while partnering with Compute, Network, and Control Plane teams.
Software
Own and optimize the performance-critical LLM serving runtime: scheduling, batching, KV cache, speculative decoding, long-context and streaming. Tune and operate runtimes (vLLM, Dynamo, TensorRT-LLM, etc.), profile GPU/network/tokenizer bottlenecks, lead model onboarding (parallelism, quantization, context), define runtime playbooks, and partner with SRE and performance teams to deploy production improvements for latency, throughput, and cost efficiency.
Software
Lead and develop the after-sales service organization for mining hardware and solutions. Oversee customer support, RMA/repair services, field service, installations, preventive maintenance, troubleshooting with engineering, performance tracking (uptime, hash rate, energy), knowledge-base development, client relationships, and team training to improve service quality and operational excellence.
Software
The Power Quality Engineer supports global data centers by monitoring and optimizing power quality, reducing equipment failures, and ensuring compliance with international standards.
Software
The Enterprise Systems Consultant will support daily operations of business systems, manage system-related requests, assist in implementation activities, and enhance business processes.
Software
Lead architecture, design, and evolution of a global multi-region cloud SRE platform for GPU/AI compute. Author and maintain platform architecture, enforce design invariants, review framework changes, run plugin framework, decide tier placements, coordinate with cloud teams and security, produce pre-flight designs, and shepherd implementations through engineering squads.
Software
Lead design and implement a global public cloud SRE platform for AI and compute workloads. Own architecture and production engineering for observability, cluster health, remediation, lifecycle, secrets, CI/CD, backup/DR, and automation. Collaborate with cross-functional teams to build scalable, reliable multi-region services and run them in production (on-call).
Software
The Finance Business Partner will assist in budgeting, forecasting, financial performance analysis, contract reviews, and stakeholder support while ensuring compliance and managing risk.
Software
Design, deploy, and operate production Kubernetes control planes for large GPU clusters. Implement GPU-specific scheduling, CRDs, multi-tenant isolation, BMaaS provisioning, Terraform-based IaC, monitoring, SLI/SLOs, and automated remediation workflows to enable autonomous AIOps-driven recovery and tenant self-service.
Software
Operate and tune InfiniBand and RoCEv2 fabrics for large GPU clusters (100–10,000 GPUs). Monitor UFM and RDMA telemetry, diagnose link/congestion issues, manage firmware and optics lifecycle, collaborate with vendor support, and feed labeled incidents/metrics into AIOps to enable predictive remediation and runbook automation.
Software
Front-line SRE covering 8AM–8PM PST monitoring GPU clusters, networking, storage, and sensors. Execute runbooks, perform hardware triage and physical DC tasks (rack, cable, swap), collect diagnostics for escalation, manage tickets (ServiceNow/Jira), update runbooks, and tag incidents to train the AIOps platform.


