Senior Site Reliability Engineer for Fuse Team

Reposted 5 Hours Ago
Be an Early Applicant
Hiring Remotely in Czechia
Remote
1M-2M Annually
Senior level
Software
The Role
As a Senior Site Reliability Engineer, you will build and maintain a reliable data platform, manage CI/CD pipelines, enhance observability, and ensure security compliance for Bloomreach's AI-driven services.
Summary Generated by Built In
Bloomreach is building the world’s premier agentic platform for personalization.We’re revolutionizing how businesses connect with their customers, building and deploying AI agents to personalize the entire customer journey.
  • We're taking autonomous search mainstream, making product discovery more intuitive and conversational for customers, and more profitable for businesses.
  • We’re making conversational shopping a reality, connecting every shopper with tailored guidance and product expertise — available on demand, at every touchpoint in their journey.
  • We're designing the future of autonomous marketing, taking the work out of workflows, and reclaiming the creative, strategic, and customer-first work marketers were always meant to do.
And we're building all of that on the intelligence of a single AI engine — Loomi — so that personalization isn't only autonomous…it's also consistent.From retail to financial services, hospitality to gaming, businesses use Bloomreach to drive higher growth and lasting loyalty. We power personalization for more than 1,400 global brands, including American Eagle, Sonepar, and Pandora.
Become a Senior SRE for Bloomreach!

Join the Fuse team — the team responsible for the item data management capabilities that connect Bloomreach Data Hub with Marketing, Search, Recommendations, and emerging Loomi agent use cases.

Fuse owns and evolves the systems behind Data Hub item collections: ingesting items data, transforming and validating it, managing data schemas and lifecycle, and distributing data reliably to downstream Bloomreach products. Item collections provide a unified source of data that can be used across all Bloomreach products.

Our current areas of focus include:

  • Unified items data pipelines: processing records into structured items and keeping data synchronized with Marketing and Search destinations.
  • Catalog APIs and lifecycle management: customer-facing and internal APIs, catalog creation and naming, schemas, destinations, migrations, and backward-compatible evolution.
  • Scalable storage and indexing: operating and improving systems built on PostgreSQL, Bigtable, Elasticsearch, Solr.
  • Reliable jobs execution: submission, queueing, execution, progress reporting, retries, cancellation, rate limiting, and operational tooling.
  • Cross-product capabilities: catalog data triggers, multi-dimensional data support, custom item types, catalog data enrichment, recommendations, and semantic catalog profiles for agentic use cases.

As a Senior SRE, you will be the team’s reliability and operability leader. You will work alongside backend engineers, embedded QA, Product, and Engineering Management to make complex product-data systems observable, scalable, safe to release, and straightforward to operate.

Fuse embraces AI-assisted engineering. We expect engineers to use modern coding agents thoughtfully to accelerate investigation, development, testing, documentation, and operational work while retaining full ownership of correctness, security, and production outcomes.

Working from one of our Central European offices (Bratislava, Prague, or Brno), or remotely (Czechia, Slovakia) on a full-time basis, you’ll become a core part of the Engineering organization.

What challenge awaits you?

As a P3 Senior SRE at Bloomreach, you are an independent reliability professional who can turn ambiguous operational problems into measurable improvements and lead initiatives end-to-end with minimal day-to-day guidance.

Your challenge will be to make Fuse’s distributed data platform dependable across the complete data path:

customer or integration → Data Hub API → records and transformations → items → asynchronous jobs execution engine → storage and indexes → Marketing, Search, Recommendations, and Loomi consumers

Your responsibilitiesa. Platform reliability and observability
  • Own and improve the reliability posture of Fuse services, workers, APIs, queues, storage systems, and destination synchronization pipelines.
  • Establish meaningful SLIs, SLOs, and error budgets for customer-facing APIs, asynchronous jobs, catalog data freshness, destination synchronization, and indexing.
  • Build end-to-end observability across Data Hub item collections, from API request and job submission through processing, persistence, indexing, and downstream delivery.
  • Ensure engineers can trace a workspace, item collection, catalog, or job across services without manually correlating disconnected logs and database records.
  • Create and maintain actionable dashboards, alerts, and service health views using Grafana, Prometheus-compatible metrics, OpenTelemetry, PagerDuty, and GCP tooling.
  • Detect missing, stalled, duplicated, or inconsistent processing before customers or downstream teams report it.
  • Improve capacity planning and autoscaling using workload telemetry, queue depth, processing throughput, latency, memory usage, storage growth, and customer-level traffic patterns.
  • Reduce noisy alerts and replace symptom-based monitoring with signals tied to customer impact.
b. Reliability of catalog storage and indexing
  • Improve the availability, scalability, and operability of catalog data across PostgreSQL/Cloud SQL, Bigtable, Elasticsearch, GCS, Kafka, and related storage systems.
  • Support catalog placement, routing, index lifecycle, shard management, safe migration, and recovery across multiple Elasticsearch clusters.
  • Develop safeguards for full replacements, delta updates, deletions, schema changes, destination changes, and catalog reindexing.
  • Define and automate data-consistency checks between source records, transformed items, job state, Bigtable, Elasticsearch, and downstream destinations.
  • Help establish practical platform limits and quotas for catalog size, API traffic, job concurrency, queue depth, payload size, and expensive operations.
  • Partner with engineers on performance testing for large catalogs and high-throughput customer workloads.
c. Infrastructure, deployments, and release safety
  • Own and evolve Kubernetes configuration and operational infrastructure for Fuse components.
  • Improve deployment automation, progressive rollout, rollback, and validation across development and production environments.
  • Make coordinated releases safer when changes span app/app, Fuse workers, Kubernetes configuration, and PostgreSQL migrations.
  • Automate operational procedures that currently depend on manual commands, one-off scripts, or specialist knowledge.
  • Maintain CI/CD pipelines with tests, linters, dependency management, security checks, image publication, and release verification.
  • Create reusable tooling for local development, ephemeral environments, end-to-end testing, load testing, and production diagnosis.
  • Ensure runbooks remain executable and are validated through exercises rather than existing only as documentation.
d. Incident management and L3 support
  • Participate in and help improve the Fuse L3/on-call rotation.
  • Lead incident investigation, mitigation, stakeholder communication, and follow-up for Fuse-owned systems.
  • Use logs, metrics, traces, database state, queue state, and Kubernetes signals to diagnose failures across distributed workflows.
  • Build safe operational tools for common support activities such as job tracing, queue inspection, rate-limit diagnosis, catalog health checks, and index recovery.
  • Facilitate blameless incident reviews and ensure resulting actions address root causes rather than only immediate symptoms.
  • Improve the handoff between customer support, L2, Fuse L3, Infrastructure, and dependent engineering teams.
  • Reduce recurring support demand by turning incident knowledge into safeguards, automation, tests, dashboards, and clear documentation.
e. Security, isolation, and compliance
  • Help Fuse meet Bloomreach security and compliance requirements, including ISO and SOC 2 controls.
  • Enforce least-privilege access, workload identity, service-level authentication and authorization, secret rotation, encryption, and auditability.
  • Protect customer isolation across workspaces, item collections, projects, accounts, databases, indexes, buckets, and asynchronous jobs.
  • Ensure operational tooling and incident procedures respect production-access restrictions and PII-handling requirements.
  • Partner with engineering teams to make security controls observable and testable rather than relying on undocumented assumptions.
f. Reliability by design
  • Participate early in the design of new Fuse capabilities so reliability, recovery, observability, limits, and operational ownership are defined before implementation.
  • Review designs for failure modes, retry behavior, idempotency, backpressure, ordering, consistency, timeout handling, cancellation, and safe rollout.
  • Clarify ownership boundaries and service contracts with teams including Campaigns, Data Pipeline, Integrations, Discovery, Recommendations, Infrastructure, Frontend, and QA.
  • Help teams choose architectures that balance immediate delivery with long-term operability and cost.
  • Coach engineers in production readiness, operational testing, debugging, and sustainable on-call practices.
Our tech stack

Primary languages: Go, Python, SQL
APIs and application: REST APIs, Python application monolith, Go workers and services, Java Infrastructure: GCP, Kubernetes/GKE, internal Kubernetes deployment tooling
Databases and storage: PostgreSQL/Cloud SQL, Bigtable, Elasticsearch, GCS, MongoDB, Redis Messaging and coordination: Kafka, ETCD, asynchronous job queues
Observability: Grafana, Prometheus-compatible metrics, OpenTelemetry, GCP Logging and Monitoring, PagerDuty CI/CD and collaboration: GitLab, Jira, Confluence
Testing: Go and Python unit/integration tests, API and end-to-end automation, performance testing
AI-assisted engineering: Claude Code, Cursor, Copilot, Gemini CLI, or comparable tools

You do not need to have used every technology listed. You should, however, have operated distributed production systems and be comfortable learning unfamiliar components while diagnosing real incidents.

Your qualificationsProfessional experienceImpact
  • You can show how your reliability work improved customer outcomes, engineering velocity, deployment confidence, or operational sustainability.
  • You have introduced practices or tooling that changed how a team builds and operates production systems.
  • You can define meaningful reliability measures and demonstrate improvement using data.
Ownership
  • You embrace the you build it, you run it principle and remain accountable from design through production operation.
  • You can lead ambiguous reliability initiatives without requiring a fully prescribed solution.
  • You take incidents from detection through mitigation, root-cause analysis, and prevention.
  • You are cost-aware and use telemetry, capacity planning, and architecture—not guesswork—to manage cloud spend.
Systematic approach
  • You treat observability, limits, runbooks, rollback, idempotency, and recovery as part of the product design.
  • You design for partial failure in distributed systems.
  • You distinguish symptoms from root causes and prioritize systemic improvements over repeated manual intervention.
  • You are comfortable working in systems where a single customer operation crosses multiple services, queues, databases, and team boundaries.
Data-driven engineering
  • You use metrics, logs, traces, profiling, and workload data to form and validate hypotheses.
  • You can turn operational telemetry into actionable feedback for developers and Product.
  • You are comfortable analyzing throughput, latency, saturation, error rates, queue behavior, database performance, and storage growth.
Technical skills
  • Strong hands-on experience operating services on Kubernetes in a major cloud environment, ideally GCP.
  • Strong experience with observability and incident diagnosis for distributed systems.
  • Experience with Go or Python; practical ability in both is a strong advantage.
  • Experience operating at least one relational database, preferably PostgreSQL or Cloud SQL.
  • Experience with one or more large-scale data or indexing systems such as Bigtable, Elasticsearch/OpenSearch, Kafka, GCS, or comparable technologies.
  • Experience designing or operating asynchronous job-processing systems, queues, workers, and retry mechanisms.
  • Experience with CI/CD, Infrastructure as Code, deployment automation, and safe database migrations.
  • Understanding of API reliability, rate limiting, backpressure, idempotency, and multi-tenant isolation.
  • Comfort participating in an on-call rotation and responding to production incidents.
  • Ability to work effectively in a distributed, remote-first team.
  • Practical use of AI coding tools to accelerate investigation and implementation without outsourcing engineering judgment.
Strongly preferred
  • Experience operating catalog, product-data, ingestion, transformation, or indexing platforms.
  • Experience with large Elasticsearch/OpenSearch clusters, shard management, index lifecycle, routing, or reindexing.
  • Experience with Bigtable or another distributed wide-column database.
  • Experience designing consistency validation across multiple storage or indexing systems.
  • Experience with customer-facing data APIs, high-volume bulk ingestion, or full and incremental synchronization.
  • Familiarity with product catalogs used by search, recommendations, marketing, or personalization systems.
  • Experience coordinating reliability improvements across several engineering teams.
  • Experience working in an environment with ISO, SOC 2, data-isolation, retention, and audit requirements.
Personal qualities
  • Ownership and accountability — you stay with a problem until it is understood, resolved, and less likely to recur.
  • Systematic thinking — you identify patterns and root causes instead of repeatedly treating symptoms.
  • Pragmatism — you balance reliability, delivery speed, complexity, and cost.
  • Clear communication — you explain technical risks and trade-offs to engineers, Product, Support, and other stakeholders.
  • Collaborative leadership — you raise the team’s operational capability rather than becoming the only person who can operate the system.
  • Customer awareness — you connect technical reliability to catalog freshness, data correctness, product availability, and customer trust.
  • Continuous improvement — you are comfortable revisiting assumptions and improving systems incrementally.
  • Remote-first effectiveness — you communicate asynchronously, document decisions, and make progress across time zones.
Your success storyIn 30 days
  • Get to know the Fuse team, its engineers, embedded QA, Product partner, Engineering Manager, and key cross-team collaborators.
  • Complete Bloomreach engineering onboarding and set up your development environments.
  • Understand the primary Fuse domains: Data Hub item collections, Catalogs, job execution and reporting, APIs, storage, indexing, and destinations.
  • Map the core request flows and data paths through Fuse owned and downstream products.
  • Review existing dashboards, alerts, L3 procedures, release practices, recent incidents, and known operational risks.
  • Shadow the L3/on-call rotation and learn the team’s production-access and escalation procedures.
In 90 days
  • Begin contributing to the Fuse L3/on-call rotation with support from experienced team members.
  • Resolve production or pre-production issues using logs, metrics, job state, database state, and distributed traces.
  • Deliver your first meaningful reliability improvement.
  • Define or improve SLIs and SLOs for at least one critical Fuse workflow.
  • Contribute to the production-readiness review of an active Fuse project.
In 180 days
  • Own the reliability posture of at least one major Fuse domain end-to-end.
  • Drive measurable improvement in one or more of:
    • Availability or successful job completion.
    • Overall data freshness.
    • Mean time to detect and recover.
    • Alert signal-to-noise ratio.
    • Deployment and migration safety.
    • Processing throughput or infrastructure efficiency.
    • L3 support effort and recurring incident volume.
  • Lead an incident review or reliability initiative involving multiple teams.
  • Establish reusable operational patterns that Fuse engineers can apply to new services and features.
  • Be a trusted partner in architecture discussions, ensuring new catalog capabilities are observable, scalable, recoverable, secure, and on-call friendly from day one.

#LI-KP1

The pay range actually offered will take into account a variety of potential factors considered in compensation, including but not limited to skills, qualifications, geographic location, accomplishments, experience, credentials, internal equity and business needs, and may vary from the range listed above.

Base Salary Range
1 260 000 Kč1 572 000 Kč CZK
More things you'll like about Bloomreach:Culture:
  • A great deal of freedom and trust. At Bloomreach we don’t clock in and out, and we have neither corporate rules nor long approval processes. This freedom goes hand in hand with responsibility. We are interested in results from day one. 
  • We have defined our 5 values and the 10 underlying key behaviors that we strongly believe in. We can only succeed if everyone lives these behaviors day to day. We've embedded them in our processes like recruitment, onboarding, feedback, personal development, performance review and internal communication. 
  • We believe in flexible working hours to accommodate your working style.
  • We work virtual-first with several Bloomreach Hubs available across three continents.
  • We organize company events to experience the global spirit of the company and get excited about what's ahead.
  • We encourage and support our employees to engage in volunteering activities - every Bloomreacher can take 5 paid days off to volunteer*.
  • The Bloomreach Glassdoor page elaborates on our stellar 4.7/5 rating. The Bloomreach Comparably page Culture score is even higher at 4.9/5
Personal Development:
  • We have a People Development Program - participating in personal development workshops on various topics run by experts from inside the company. We are continuously developing & updating competency maps for select functions.
  • Our resident communication coach Ivo Večeřa is available to help navigate work-related communications & decision-making challenges.*
  • Our managers are strongly encouraged to participate in the Leader Development Program to develop in the areas we consider essential for any leader. The program includes regular comprehensive feedback, consultations with a coach and follow-up check-ins.
  • Bloomreachers utilize the $1,500 professional education budget on an annual basis to purchase education products (books, courses, certifications, etc.)*
Well-being:
  • The Employee Assistance Program -- with counselors -- is available for non-work-related challenges.*
  • Subscription to Calm - sleep and meditation app.*
  • We organize ‘DisConnect’ days where Bloomreachers globally enjoy one additional day off each quarter, allowing us to unwind together and focus on activities away from the screen with our loved ones.
  • We facilitate sports, yoga, and meditation opportunities for each other.
  • Extended parental leave up to 26 calendar weeks for Primary Caregivers.*
Compensation:
  • Restricted Stock Units or Stock Options are granted depending on a team member’s role, seniority, and location.*
  • Everyone gets to participate in the company's success through the company performance bonus.*
  • We offer an employee referral bonus of up to $3,000!
  • We reward & celebrate work anniversaries -- Bloomversaries!*

(*Subject to employment type. Interns are exempt from marked benefits, usually for the first 6 months.)

Excited? Join us and transform the future of commerce experiences!

If this position doesn't suit you, but you know someone who might be a great fit, share it - we will be very grateful!

Any unsolicited resumes/candidate profiles submitted through our website or to personal email accounts of employees of Bloomreach are considered property of Bloomreach and are not subject to payment of agency fees.

#LI-Remote

Skills Required

  • Solid hands-on experience with GCP
  • Experience in Kubernetes
  • Familiarity with data pipeline technologies
  • Fluent use of AI coding agents
  • Comfortable with on-call rotation

Bloomreach Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Bloomreach and has not been reviewed or approved by Bloomreach.

  • Fair & Transparent Compensation Pay is considered competitive versus peers and aligned with tech‑market norms for comparable roles and levels. Many roles cite compensation that feels fair and market‑appropriate.
  • Strong & Reliable Incentives Company‑wide performance bonuses follow a semi‑annual cadence, and go‑to‑market roles feature structured base/OTE plans. This predictable incentive design meaningfully augments base pay.
  • Leave & Time Off Breadth Quarterly company‑wide DisConnect Days, generous PTO practices, and paid volunteer time expand time off beyond standard holidays. These scheduled shutdowns are designed to enable genuine unplugging in a remote‑first setup.

Bloomreach Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Mountain View, CA
600 Employees
Year Founded: 2009

What We Do

Bloomreach is the leader in Commerce Experience™ Our Bloomreach Experience Platform (brX) competes in three core categories: Engagement (CDP and marketing automation), Content (headless content and experience management), and Discovery (e-commerce search, merchandising, recommendations, and SEO). We connect both customer data and product data to personalize all customer touch-points, leveraging our patented AI to recommend, predict, and segment. This empowers the marketer to create individual experiences, increase revenue, strengthen customer loyalty, and improve efficiency. With a global footprint, Bloomreach powers over 25% of all e-commerce experiences across the US and UK, and supports 300+ global enterprises including Neiman Marcus, CapitalOne, Staples, NHS Digital, Bosch, Puma, and Marks & Spencer.

Similar Jobs

GC AI Logo GC AI

Head of EMEA, GTM

Artificial Intelligence • Legal Tech
In-Office or Remote
28 Locations
130 Employees

Deepgram Logo Deepgram

Account Executive

Artificial Intelligence • Machine Learning • Natural Language Processing • Software • Conversational AI
In-Office or Remote
28 Locations
150 Employees

Deepgram Logo Deepgram

Account Executive

Artificial Intelligence • Machine Learning • Natural Language Processing • Software • Conversational AI
Remote
27 Locations
150 Employees
Remote
Czech Republic
575 Employees

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account