Software Mind develops solutions that make an impact for companies around the globe. Tech giants & unicorns, transformative projects, emerging technologies and limitless opportunities – these are a few words that describe an average day for us. Building cross-functional engineering teams that take ownership and crave more means we’re always on the lookout for talented people who bring passion and creativity to every project. Our culture embraces openness, acts with respect, shows grit & guts and combines employment with enjoyment.
Job DescriptionProject – the aim you'll have
We are the AI Experience Framework team that builds the platform powering ServiceNow's AI-first user interfaces - an SSR runtime (karuna) built on Lit and server-rendered web components, running behind a multi-tier proxy/HTTP2 routing chain with sharded V8 isolate pools, paired with a ServiceNow Glide/Java platform layer (karuna-glide) that supplies metadata, ACLs, and service artifacts. This role owns production reliability for that stack end to end: Kubernetes deployment and operations, observability, and hands-on troubleshooting of both the Node.js and JVM sides of the system - not generalist infrastructure work.
Position – how you’ll contribute
- Own Kubernetes deployment and operational health for Framework services, including scaling, rollout/rollback strategy, and resource tuning
- Build and maintain production observability - Grafana dashboards and Prometheus alerting rules - across the SSR runtime and the Glide platform layer
- Diagnose and resolve Node.js production incidents: event-loop stalls, heap growth, V8 isolate exhaustion, and isolate-pool scheduling issues under concurrent versioned traffic (vN/vN-1)
- Diagnose and resolve JVM production incidents on the Glide/Java side: GC pressure, thread dumps, and platform-service latency
- Own incident response for the team: runbooks, on-call rotation, postmortems, and paging hygiene
- Drive CI/CD and infrastructure-as-code for Kubernetes manifests/Helm and deployment pipelines
- Partner with the framework engineering team to identify reliability gaps before they become incidents - capacity planning, load testing, chaos/failure-injection where useful
- Represent production reliability concerns in architecture reviews for new framework capabilities
Expectations – the experience you need
- Production operations/SRE experience, including hands-on Kubernetes deployment, scaling, and incident response
- Direct operational experience troubleshooting Node.js in production: reading heap snapshots and CPU profiles, diagnosing event-loop blocking, and understanding process/worker isolation models (V8 isolates or equivalent sandboxing)
- Direct operational experience troubleshooting JVM-based services in production: GC log analysis, thread dump analysis, and JVM tuning
- Hands-on experience building and maintaining Prometheus alerting rules and Grafana dashboards from scratch, not just consuming existing ones
- Strong Linux/networking fundamentals: DNS, load balancing, TCP/HTTP semantics (including HTTP/2), and debugging service-to-service networking inside Kubernetes
- Experience with CI/CD and infrastructure-as-code for containerized deployments (Helm, GitOps tooling such as ArgoCD/Flux, or equivalent)
- Track record owning on-call rotations, writing runbooks, and driving postmortems that lead to real reliability improvements
- Experience with Splunk for log aggregation, search, and production troubleshooting.
- Hands-on experience with in-memory caching systems (Valkey/Redis) in production — key design, TTL/eviction tuning, and tenant-scoped cache invalidation
- Experience managing service-to-service mTLS - certificate issuance, rotation, and format conversion (e.g., PKCS/BCFKS↔PEM) - plus JWT-based service authentication
- Very good spoken and written English.
Additional skills – the edge you have
- Familiarity with server-side rendering architectures and the specific failure modes of isomorphic runtimes (markup mismatches, browser-API leakage into server code)
- Experience operating multi-version/canary rollout strategies (two live app versions served concurrently)
- Working knowledge of the ServiceNow Glide platform or a comparable enterprise platform integration layer
- Experience with distributed tracing and request-context correlation across service boundaries
- Familiarity with event-driven autoscaling (e.g., KEDA ScaledObjects driven by PromQL triggers) as a complement to standard HPA-based scaling
Our offer – professional development, personal growth:
- Flexible employment and remote work
- International projects with leading global clients
- International business trips
- Non-corporate atmosphere
- Language classes
- Internal & external training
- Private healthcare and insurance
- Multisport card
- Well-being initiatives
Position at: Software Mind
Skills Required
- Production SRE experience including hands-on Kubernetes deployment, scaling, rollout/rollback strategy, and resource tuning
- Direct operational experience troubleshooting Node.js in production (heap snapshots, CPU profiles, event-loop blocking, V8 isolate models)
- Direct operational experience troubleshooting JVM/Java services (GC log analysis, thread dump analysis, JVM tuning)
- Hands-on experience building and maintaining Prometheus alerting rules and Grafana dashboards from scratch
- Strong Linux and networking fundamentals (DNS, load balancing, TCP/HTTP semantics including HTTP/2) and Kubernetes service-to-service networking debugging
- Experience with CI/CD and infrastructure-as-code for containerized deployments (Helm, GitOps tooling such as ArgoCD/Flux, or equivalent)
- Track record owning on-call rotations, writing runbooks, and conducting postmortems that improve reliability
- Experience with Splunk for log aggregation, search, and production troubleshooting
- Hands-on experience with in-memory caching systems (Valkey/Redis) in production including TTL/eviction tuning and tenant-scoped invalidation
- Experience managing service-to-service mTLS (certificate issuance, rotation, format conversion) and JWT-based service authentication
- Very good spoken and written English
- Familiarity with server-side rendering architectures and isomorphic runtime failure modes (markup mismatches, browser-API leakage into server code)
- Experience operating multi-version/canary rollout strategies (serving two live app versions concurrently)
- Working knowledge of the ServiceNow Glide platform or comparable enterprise platform integration layer
- Experience with distributed tracing and request-context correlation across service boundaries
- Familiarity with event-driven autoscaling (e.g., KEDA ScaledObjects driven by PromQL triggers)
Software Mind Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Software Mind and has not been reviewed or approved by Software Mind.
-
Fair & Transparent Compensation — Pay is considered competitive for core hiring markets, with “good salary” cited in multiple locales. Public salary snapshots provide a baseline that helps candidates assess offers and negotiations.
-
Flexible Benefits — Remote or hybrid options are prominently highlighted, and a remote‑work program is publicly noted alongside positively cited work‑from‑home experiences. Flexibility around schedules and location is presented as part of the package.
-
Wellbeing & Lifestyle Benefits — Private medical care, language classes, sports/fitness support, and learning initiatives are listed for several Central/Eastern European locations, with occasional workation perks promoted. These lifestyle‑oriented offerings complement base pay and can enhance perceived total rewards.
Software Mind Insights
What We Do
Software Mind is a global digital transformation partner with operations throughout Europe, the US and LATAM. Driven by tech and empowered by people, we provide companies with software engineers and autonomous, cross-functional development teams who manage software life cycles from ideation to release and beyond. For over 20 years we’ve been enriching organizations with the talent they need to boost scalability, drive dynamic growth and bring disruptive ideas to life. Our top-notch engineering teams combine ownership with leading technologies, including cloud, AI, data science and embedded software to accelerate digital transformations and boost software delivery. A culture, driven by trust, that embraces openness, craves more and acts with respect enables our experts to create evolutive solutions that support scale-ups, unicorns and enterprise-level companies around the world.









