[8SN] Senior Site Reliability Engineer (SRE) – Kubernetes

Posted 2 Days Ago
Be an Early Applicant
Hiring Remotely in Montréal, QC, CAN
In-Office or Remote
Senior level
Software
The Role
Senior SRE to ensure reliability of a Kubernetes-based UI/AI service: operate deployments, respond to incidents/on-call, troubleshoot Node.js and Java runtimes, implement observability and GitOps CI/CD, and improve reliability through postmortems and runbooks.
Summary Generated by Built In
Company Description

We are Software Mind, an awesome team of engineers who are ready to ramp up any top-notch company’s projects! Our aim? To always be one step ahead. Become part of a multicultural company in constant growth with an excellent work environment certified by Great Place To Work!
 

About the Client

Our client is a leading enterprise software company building highly scalable cloud-native platforms used by organizations around the world. Their engineering teams focus on delivering reliable, secure, and high-performing services while embracing modern DevOps, Kubernetes, and cloud technologies.

You will join a team responsible for ensuring the stability, reliability, and operational excellence of a critical UI service running in production.

Contract Duration: Initial contract through the end of 2026, extending the engagement to a total 12-month term based on performance.

Job Description

About the Role

This is a Senior SRE role supporting production reliability for a Kubernetes-based UI service / AI experience framework stack.

This is not general infrastructure, and it is not a front-end developer role. The strongest candidates will have production SRE experience across Kubernetes operations, observability, Node.js runtime troubleshooting, JVM / Java service troubleshooting, Splunk, and incident ownership.

What You’ll Do

  • Support the deployment, operation, and reliability of production services running on Kubernetes.
  • Monitor service health and investigate production incidents across distributed applications.
  • Participate in on-call support, incident response, root cause analysis, postmortems, and reliability improvements.
  • Troubleshoot application runtime, networking, and service-to-service issues in collaboration with engineering teams.
  • Support CI/CD, GitOps-based deployments, observability, and production monitoring.
  • Work within a client-directed backlog and established priorities.

Qualifications

Required Qualifications

  • 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Production Engineering, or a closely related role, including strong recent hands-on experience supporting Kubernetes-based production services.
  • 3+ years of hands-on production Kubernetes experience strongly preferred. Kubernetes production operations, including deployment, scaling, rollout / rollback, resource tuning, and service-to-service troubleshooting
  • Strong production incident response experience, including on-call, runbooks, postmortems, and paging hygiene
  • Splunk experience for log aggregation, search, and production troubleshooting
  • Prometheus and Grafana experience, specifically building alert rules and dashboards, not only using existing dashboards
  • CI/CD and infrastructure-as-code for containerized deployments, including Helm and GitOps tools such as ArgoCD or Flux
  • Strong Linux and networking fundamentals, including DNS, load balancing, TCP / HTTP, HTTP/2, and Kubernetes networking
  • Production troubleshooting experience across Node.js and JVM/Java services, with strong depth in at least one runtime environment. Experience may include Node.js heap snapshots, CPU profiling, event-loop and memory analysis, as well as JVM GC log analysis, thread dumps, JVM tuning, and Java service latency investigation.
  • Service-to-service authentication experience, including mTLS, certificate rotation, certificate format conversion, and JWT-based service authentication

Additional Information

Nice to Have

  • Web Components / Lit experience, to perform first-level debugging of UI-related issues
  • Server-side rendering or isomorphic runtime experience
  • Canary rollout / multi-version production operations
  • Distributed tracing and request-context correlation
  • KEDA or event-driven autoscaling
  • Experience with enterprise platform integration layers

What We Offer

  • Competitive salary and laptop
  • Professional development and training opportunities
  • Work with cutting-edge cloud and container technologies
  • Flexible work arrangements and collaborative team environment
  • Impact on organization-wide digital transformation initiatives

Skills Required

  • 5+ years in Site Reliability Engineering, DevOps, Platform or Production Engineering
  • 3+ years production Kubernetes experience (deployment, scaling, rollouts, resource tuning, networking)
  • Production incident response experience including on-call, runbooks, postmortems, and paging hygiene
  • Splunk for log aggregation, search, and production troubleshooting
  • Prometheus and Grafana for building alert rules and dashboards
  • CI/CD and infrastructure-as-code for containerized deployments, including Helm and GitOps tools (ArgoCD or Flux)
  • Strong Linux and networking fundamentals (DNS, load balancing, TCP/HTTP/HTTP/2, Kubernetes networking)
  • Production troubleshooting across Node.js and JVM/Java services with deep expertise in at least one runtime
  • Service-to-service authentication experience including mTLS, certificate rotation/format conversion, and JWT-based auth

Software Mind Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Software Mind and has not been reviewed or approved by Software Mind.

  • Fair & Transparent Compensation Pay is considered competitive for core hiring markets, with “good salary” cited in multiple locales. Public salary snapshots provide a baseline that helps candidates assess offers and negotiations.
  • Flexible Benefits Remote or hybrid options are prominently highlighted, and a remote‑work program is publicly noted alongside positively cited work‑from‑home experiences. Flexibility around schedules and location is presented as part of the package.
  • Wellbeing & Lifestyle Benefits Private medical care, language classes, sports/fitness support, and learning initiatives are listed for several Central/Eastern European locations, with occasional workation perks promoted. These lifestyle‑oriented offerings complement base pay and can enhance perceived total rewards.

Software Mind Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Kraków
1,000 Employees
Year Founded: 1999

What We Do

Software Mind is a global digital transformation partner with operations throughout Europe, the US and LATAM. Driven by tech and empowered by people, we provide companies with software engineers and autonomous, cross-functional development teams who manage software life cycles from ideation to release and beyond. For over 20 years we’ve been enriching organizations with the talent they need to boost scalability, drive dynamic growth and bring disruptive ideas to life. Our top-notch engineering teams combine ownership with leading technologies, including cloud, AI, data science and embedded software to accelerate digital transformations and boost software delivery. A culture, driven by trust, that embraces openness, craves more and acts with respect enables our experts to create evolutive solutions that support scale-ups, unicorns and enterprise-level companies around the world.

Similar Jobs

Block Logo Block

Systems Engineer

Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
In-Office or Remote
8 Locations
12000 Employees
120K-213K Annually

Block Logo Block

Senior Data Engineer

Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
In-Office or Remote
8 Locations
12000 Employees
168K-297K Annually

Toast Logo Toast

Senior Software Engineer

Cloud • Fintech • Food • Information Technology • Software • Hospitality
Remote
Canada
5000 Employees
115K-165K Annually

Zapier Logo Zapier

Sr. Director, Security

Artificial Intelligence • Productivity • Software • Automation
Remote
2 Locations
800 Employees
308K-463K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account