Senior Site Reliability Engineer - OneRail’s Technology Hub - Poland

Posted 12 Days Ago
Be an Early Applicant
Kraków, Małopolskie, POL
In-Office
336K-384K Annually
Senior level
Artificial Intelligence • Logistics • Software • 3PL: Third Party Logistics
The Role
Lead design and optimization of a cloud-native SaaS platform for availability, scalability, observability, and resilience. Implement observability, performance and chaos testing, automate operations, tune databases, lead incident response and postmortems, mentor teams, and establish SRE best practices and documentation.
Summary Generated by Built In
OneRail seeks a Senior Site Reliability Engineer to help ensure the availability, performance, scalability, observability, and resilience of our SaaS platform. In this role, you’ll be responsible for designing and optimizing platform architecture, improving system reliability, driving automation initiatives, and enabling engineering teams to build and operate highly scalable services.
This is a hands-on role for someone with deep expertise in cloud-native platforms, distributed systems, observability, and performance optimization. You’ll work closely with Engineering, Product, and Operations teams to improve platform reliability, scalability, and operational efficiency while reducing manual overhead through automation. Over time, you’ll have opportunities to influence platform architecture, lead reliability initiatives, and establish Site Reliability Engineering best practices across the organization.
Responsibilities
 
  • Serve as the platform subject matter expert, mentoring engineering teams on reliability, scalability, security, and operational best practices.
  • Design, implement, and maintain observability solutions covering logs, metrics, traces, APM, and alerting across all platform services.
  • Benchmark, analyze, and optimize application performance, cloud services, integrations, databases, and distributed systems.
  • Build and maintain performance, load, chaos, and resilience testing frameworks to proactively identify system weaknesses.
  • Perform advanced analysis, tuning, and optimization of structured and unstructured databases to improve performance and scalability.
  • Develop automation frameworks and operational tooling that eliminate manual processes and reduce human error.
  • Lead incident response efforts, conduct root cause analysis, and drive continuous improvement through postmortem reviews and reliability initiatives.
  • Optimize the performance of databases, caches, streaming platforms, message brokers, and backend services.
  • Collaborate closely with Engineering, Product, and Operations teams to deliver highly available, production-grade platform solutions.
  • Establish and maintain monitoring, alerting, and reliability standards across the platform.
  • Create and maintain technical documentation, runbooks, architectural diagrams, and operational procedures.
  • Evaluate emerging technologies and recommend solutions that improve platform reliability, scalability, observability, and operational efficiency.

Qualifications
 
  • Bachelor’s degree in Computer Science, Engineering, or a related technical field.
  • 5+ years of experience in Site Reliability Engineering, Platform Engineering or a related role.
  • Strong experience with cloud-native architectures and distributed systems.
  • Hands-on experience implementing observability solutions, including metrics, logging, tracing, and application performance monitoring.
  • Experience designing scalable backend architecture using Node.js, TypeScript, .NET
  • Strong knowledge of database administration, performance tuning, and optimization, including Azure Cosmos DB, MySQL, or similar platforms.
  • Experience building automation frameworks, scripting solutions, and operational tooling.
  • Familiarity with CI/CD pipelines and continuous integration practices using GitHub Actions or similar platforms.
  • Experience with containerization technologies such as Docker.
  • Strong understanding of event-driven architectures, messaging systems, and real-time data processing platforms.
  • Experience with monitoring and observability tools such as Datadog, OpenTelemetry, ELK Stack, App Insights, Grafana, or Prometheus.
  • Excellent troubleshooting, analytical, and problem-solving skills.
  • Strong written and verbal communication skills with the ability to collaborate across technical and business teams.
  • Advanced proficiency in English & Polish, both written and spoken (B2+).
Work Location
Hybrid, Kraków, Poland

Office location: Kraków, Poland
Work Style: Hybrid (2-3 days from the office/week)
Salary levels: 28 000 – 32 000 PLN/month
 

Skills Required

  • Bachelor's degree in Computer Science, Engineering, or related technical field
  • 5+ years of experience in Site Reliability Engineering, Platform Engineering or related role
  • Experience with cloud-native architectures and distributed systems
  • Hands-on experience implementing observability solutions (metrics, logging, tracing, APM)
  • Experience designing scalable backend architecture using Node.js, TypeScript, .NET
  • Knowledge of database administration and performance tuning, including Azure Cosmos DB and MySQL
  • Experience building automation frameworks, scripting solutions, and operational tooling
  • Familiarity with CI/CD pipelines and continuous integration practices using GitHub Actions or similar
  • Experience with containerization technologies such as Docker
  • Strong understanding of event-driven architectures, messaging systems, and real-time data processing platforms
  • Experience with monitoring and observability tools such as Datadog, OpenTelemetry, ELK Stack, App Insights, Grafana, or Prometheus
  • Excellent troubleshooting, analytical, and problem-solving skills
  • Strong written and verbal communication skills; ability to collaborate across technical and business teams
  • Advanced proficiency in English and Polish, both written and spoken (B2+)
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Orlando, FL
217 Employees
Year Founded: 2018

What We Do

OneRail is a leading omnichannel fulfillment solution that pairs best-in-class software with logistics as a service to provide dependability and speed. Their platform utilizes AI to connect shippers with a courier ecosystem, automating and optimizing the delivery supply chain.

Similar Jobs

Samsara Logo Samsara

Senior Software Engineer

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
Poland
4000 Employees

Capco Logo Capco

Senior Java Engineer

Fintech • Professional Services • Consulting • Energy • Financial Services • Cybersecurity • Generative AI
Remote or Hybrid
Poland
6000 Employees

Cencora Logo Cencora

Head of IT

Healthtech • Logistics • Pharmaceutical
In-Office
Kraków, Małopolskie, POL
51000 Employees

Rain Logo Rain

Technical Support

Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3 • Infrastructure as a Service (IaaS)
Remote or Hybrid
28 Locations
100 Employees

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account