About Bitdeer:
Bitdeer is a world-leading technology company for Bitcoin mining and AI cloud.
Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers. Apart from designing industry-leading ASIC chips and manufacturing mining rigs, the Group handles complex processes involved in computing across the value chain. This includes equipment procurement, transport logistics, datacenter design and construction, equipment management, and network and facility operations. Bitdeer also offers advanced cloud capabilities to customers with a high demand for artificial intelligence.
Headquartered in Singapore, Bitdeer operates globally with a diversified 3 GW energy portfolio, and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan, Norway, Canada, Malaysia, and Ethiopia.
What you will be responsible for:
- Design and implement advanced batch scheduling architectures using frameworks like Volcano or YuniKorn to support multi-node gang scheduling.
- Develop and manage cluster-wide admission control and sophisticated job queueing mechanisms utilizing Kueue to manage high-volume AI workload traffic.
- Leverage Kubernetes Dynamic Resource Allocation (DRA) and custom scheduler plugins to manage complex accelerator requests natively.
- Architect topology-aware pod placement strategies that optimize for low-latency communication via NVLink and InfiniBand fabrics.
- Implement automated GPU sharing technologies (e.g., MIG, time-slicing) and multi-tenancy isolation policies to maximize cluster-wide utilization.
- Collaborate with the GPU Systems and Storage teams to ensure the scheduling layer is tightly integrated with bare-metal hardware and storage I/O patterns.
- Drive the reliability and scalability of the scheduling stack, resolving resource contention and deadlock scenarios in large-scale HPC environments.
- Mentor junior engineers and conduct design reviews to maintain architectural excellence in our orchestration layer.
How you will stand out:
- Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related field.
- 6+ years of distributed systems engineering, with deep, hands-on expertise in Kubernetes scheduling frameworks and orchestrators.
- Extensive experience with AI workload execution patterns and distributed training frameworks (e.g., PyTorch Distributed, Ray, MPI).
- Proven track record of operating, debugging, and scaling scheduling stacks in high-performance computing (HPC) or large-scale production cloud environments.
- Strong knowledge of GPU hardware architectures and the specific scheduling challenges related to distributed AI training and inference.
- Experience with infrastructure automation and infrastructure-as-code (e.g., Terraform, Go-based Operators).
- Excellent technical communication and leadership skills; ability to influence cross-functional teams and align architectural goals.
- Ability to work in a high-velocity engineering environment and translate complex, ambiguous requirements into concrete, scalable engineering solutions.
What you will experience working with us:
- A culture that values authenticity and diversity of thoughts and backgrounds;
- An inclusive and respectable environment with open workspaces and exciting start-up spirit;
- Fast-growing company with the chance to network with industrial pioneers and enthusiasts;
- Ability to contribute directly and make an impact on the future of the digital asset industry;
- Involvement in new projects, developing processes/systems;
- Personal accountability, autonomy, fast growth, and learning opportunities;
- Attractive welfare benefits and developmental opportunities such as training and mentoring.
--------------------------------------------------------------------
Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, colour, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.
#LI-ST1
Skills Required
- Bachelor's or Master's degree in Computer Science, Electrical Engineering, or related field.
- 6+ years of distributed systems engineering with deep, hands-on expertise in Kubernetes scheduling frameworks and orchestrators.
- Experience designing and implementing batch scheduling architectures (e.g., Volcano, YuniKorn) and gang scheduling.
- Experience developing cluster-wide admission control and job queueing mechanisms (e.g., Kueue).
- Familiarity with Kubernetes Dynamic Resource Allocation (DRA) and building custom scheduler plugins.
- Expertise in topology-aware pod placement and low-latency fabric optimization (NVLink, InfiniBand).
- Experience implementing automated GPU sharing (MIG, time-slicing) and multi-tenancy isolation policies.
- Extensive experience with AI workload execution and distributed training frameworks (PyTorch Distributed, Ray, MPI).
- Proven track record operating, debugging, and scaling scheduling stacks in HPC or large-scale production cloud environments.
- Strong knowledge of GPU hardware architectures and scheduling challenges for distributed AI training/inference.
- Experience with infrastructure automation and infrastructure-as-code (e.g., Terraform, Go-based Operators).
- Excellent technical communication, leadership, and mentoring skills.
- Ability to work in a high-velocity engineering environment and translate ambiguous requirements into scalable solutions.
What We Do
Bitdeer Technologies Group (Nasdaq: BTDR) is a leader in the blockchain and high-performance computing industry. It is one of the world’s largest holders of proprietary hash rate and suppliers of hash rate. Bitdeer is committed to providing comprehensive computing solutions for its customers. The company was founded by Jihan Wu, an early advocate and pioneer in cryptocurrency who cofounded multiple leading companies serving the blockchain economy. Mr. Wu leads the company as Founder, Chairman, and CEO. Linghui Kong serves as Bitdeer’s CBO and provides leadership through deep industry knowledge and technology expertise. Headquartered in Singapore, Bitdeer has deployed mining datacenters in the United States, Norway, and Bhutan. It offers specialized mining infrastructure, high-quality hash rate sharing products, and reliable hosting services to global users. The company also offers advanced cloud capabilities for customers with high demands for artificial intelligence. Dedication, authenticity, and trustworthiness are foundational to our mission of becoming the world’s most reliable provider of full-spectrum blockchain and high-performance computing solutions. We welcome global talent to join us in shaping the future







