SRE Monitoring Platform Software Engineer (Entry Level)

Reposted Yesterday
San Jose, CA, USA
In-Office
105K-155K Annually
Junior
Software
The Role
Build and maintain SRE microservices supporting a global GPU infrastructure platform. Implement GitOps, declarative configuration, and CI/CD automation; monitor metrics, logs, and traces; support operational readiness through on-call participation, runbooks, and incident reviews; and write unit, integration, and end-to-end tests. Work with senior engineers to deliver production-ready platform features while maintaining service reliability and infrastructure consistency.
Summary Generated by Built In

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.

Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.
To learn more, visit https://ir.bitdeer.com/

Position Overview 

Bitdeer is building an AI-operated GPU cloud — a global fleet of self-built and OEM-rented data centers running the world's most valuable compute, run by a platform that observes, protects, and operates the fleet. The SRE Platform team builds the monitoring and automation substrate that every other squad — storage, network, GPU, K8S, and L1 operators — depends on. Their signals become the system you help build.

As an entry-level Software Engineer on the SRE / Monitoring Platform team, you contribute to the NeoCloud SRE platform — the multi-region system that observes, protects, and operates a GPU rental fleet across self-built and OEM-rented data centers. You join a bounded context led by a senior engineer, take well-scoped components from design to production code that ships through GitOps + the CICD release pipeline, follows the Plugin Framework conventions, meets declared SLOs, and stays drift-free.

This is a build + learn role. You write code, write tests, and operate what you build under the guidance of a senior engineer. You participate in on-call as a shadow before taking primary. Within 12 months, you should be delivering components independently within your assigned area and growing toward owning a sub-context.

Key Responsibilities

    Where you'll contribute (guided by a senior engineer)

    • Collection + Storage — help build collection-agent, metrics-store / logs-store / traces-store / profiles-store, enrichment-service, collection-monitor. Write ingestion, query, and storage-path code.
    • Alert + Correlation + SLO — contribute to alert-engine-framework, alert-correlation, slo-framework; implement and tune default alert rules.
    • Topology + Cluster-Health — contribute to topology-service, cluster-health-rollup, OSS-SRE-tool collection plugins for K8s / Slurm / Ray / Volcano / Kueue / KubeRay.
    • Remediation + Workflow + Jobs — help build remediation-actuator, orchestration / workflow components, inspection probes, job-scheduler.
    • Observability instrumentation — instrument services with metrics, logs, and traces via OpenTelemetry; build dashboards; write runbooks an on-call can follow.
    • Test discipline — write unit / integration / contract tests for everything you ship; participate in chaos and soak tests led by senior engineers.

    Why this is a great first role

    • Greenfield with a well-defined vision. The Plugin Framework, GitOps pipeline, and SLO framework are decided; you build components inside them with a clear blueprint — not from a blank page.
    • You learn the full observability stack at production scale — ingest, query, storage — by building it, not just using it.
    • Mentorship-heavy. You work directly with senior and principal engineers who own the architecture; their expertise becomes your growth path.

Job Requirements

    • 0-2 years of software engineering experience (new graduates with strong projects or internships welcome).
    • Solid fundamentals in one programming language — Go (preferred), Python, Java, or Rust. You can write clean, tested, readable code and explain your design choices.
    • CS fundamentals — data structures, algorithms, concurrency, basic networking (TCP / HTTP), and operating-system concepts (processes, threads, I/O). You can reason about correctness and performance.
    • Distributed systems basics — you understand the ideas behind idempotency, retries, back-pressure, caching, and eventual consistency, even if you haven't operated them at scale yet. Eagerness to go deep.
    • Monitoring / observability exposure — some hands-on with Prometheus, Grafana, Loki, or similar; can write a basic PromQL query and instrument a service. Eagerness to learn the ingest, query, and storage path of a real observability stack.
    • Familiarity with Linux and the shell; comfort reading system logs and using standard debugging tools.
    • Kubernetes basics — understand Pods, Services, Deployments; have run something on K8s (a project, lab, or internship).
    • Git + CI basics — branching, pull requests, and have used a CI pipeline (GitHub Actions, GitLab CI, or similar).
    • Test discipline — you write unit and integration tests as a habit, not an afterthought.
    • Communication — clear written and verbal English; can write a good PR description and ask good questions.
    • Curiosity and a learning mindset — the most important qualifier. You're excited to learn GPU / AI infrastructure, AIOps, distributed systems, and observability at production scale.

Nice-to-Haves

    • Internship or project in monitoring / observability, telemetry pipelines, or platform / SRE tooling.
    • Exposure to GPU / AI-infra — DCGM, InfiniBand / RoCE, Kubernetes GPU Operator, Slurm / Ray. Interest counts more than depth.
    • Exposure to AIOps / ML-adjacent tooling (anomaly detection, alert correlation).
    • Contributions to open-source observability or cloud-native projects.

--------------------------------------------------------------------

Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.

Skills Required

  • Bachelor's degree in Computer Science, Computer Engineering, or a related technical field, or equivalent practical experience or internships
  • 0-2 years of hands-on software development experience
  • Proficiency in Go, Java, or Rust
  • Scripting abilities in Python or Bash
  • Knowledge of data structures, algorithms, object-oriented design, and distributed systems concepts
  • Hands-on exposure to Docker, Kubernetes, and Linux fundamentals
  • Experience writing unit and integration tests
  • Strong technical writing and communication skills
  • Experience with Kubernetes Operators, Helm, or GitOps tools such as ArgoCD or Flux
  • Exposure to time-series databases or observability tools including Prometheus, OpenTelemetry, Grafana, or Loki
  • Familiarity with hardware, GPU/AI infrastructure, NVIDIA DCGM, CUDA, or high-performance computing
  • Familiarity with Terraform or Ansible
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Singapore
214 Employees

What We Do

Bitdeer Technologies Group (Nasdaq: BTDR) is a leader in the blockchain and high-performance computing industry. It is one of the world’s largest holders of proprietary hash rate and suppliers of hash rate. Bitdeer is committed to providing comprehensive computing solutions for its customers. The company was founded by Jihan Wu, an early advocate and pioneer in cryptocurrency who cofounded multiple leading companies serving the blockchain economy. Mr. Wu leads the company as Founder, Chairman, and CEO. Linghui Kong serves as Bitdeer’s CBO and provides leadership through deep industry knowledge and technology expertise. Headquartered in Singapore, Bitdeer has deployed mining datacenters in the United States, Norway, and Bhutan. It offers specialized mining infrastructure, high-quality hash rate sharing products, and reliable hosting services to global users. The company also offers advanced cloud capabilities for customers with high demands for artificial intelligence. Dedication, authenticity, and trustworthiness are foundational to our mission of becoming the world’s most reliable provider of full-spectrum blockchain and high-performance computing solutions. We welcome global talent to join us in shaping the future

Similar Jobs

Spectrum Logo Spectrum

Account Executive

Information Technology • Internet of Things • Mobile • On-Demand • Software
In-Office
Indio, CA, USA
100000 Employees
53K-87K Annually

Capital One Logo Capital One

Senior Data Analyst

Fintech • Machine Learning • Payments • Software • Financial Services
Hybrid
5 Locations
55000 Employees
101K-138K Annually

Capital One Logo Capital One

Director, AI Engineering (Remote - eligible)

Fintech • Machine Learning • Payments • Software • Financial Services
Remote or Hybrid
3 Locations
55000 Employees
245K-335K Annually

Capital One Logo Capital One

Senior Manager, Product Management - Consumer and Developer Experience

Fintech • Machine Learning • Payments • Software • Financial Services
Hybrid
5 Locations
55000 Employees
183K-250K Annually

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software • Productivity
US
15 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account