Software Mind Logo

Software Mind

[MLA] Senior Site Reliability Engineer (SRE) – Kubernetes

Posted 7 Days Ago
Be an Early Applicant
In-Office or Remote
Hiring Remotely in Kraków, Małopolskie
Senior level
In-Office or Remote
Hiring Remotely in Kraków, Małopolskie
Senior level
Own production reliability for an AI Experience Framework: operate Kubernetes deployments, monitor and troubleshoot distributed Node.js and Java/JVM services, run incident response/on-call duties, implement CI/CD and GitOps (Helm/ArgoCD/Flux), build observability with Prometheus/Grafana/Splunk, and collaborate on reliability improvements and root-cause analysis.
The summary above was generated by AI
Company Description

Software Mind develops solutions that make an impact for companies around the globe. Tech giants & unicorns, transformative projects, emerging technologies and limitless opportunities – these are a few words that describe an average day for us. Building cross-functional engineering teams that take ownership and crave more means we’re always on the lookout for talented people who bring passion and creativity to every project. Our culture embraces openness, acts with respect, shows grit & guts and combines employment with enjoyment.

Job Description

Project – the aim you'll have 

We are the AI Experience Framework team that builds the platform powering ServiceNow's AI-first user interfaces - an SSR runtime (karuna) built on Lit and server-rendered web components, running behind a multi-tier proxy/HTTP2 routing chain with sharded V8 isolate pools, paired with a ServiceNow Glide/Java platform layer (karuna-glide) that supplies metadata, ACLs, and service artifacts. This role owns production reliability for that stack end to end: Kubernetes deployment and operations, observability, and hands-on troubleshooting of both the Node.js and JVM sides of the system - not generalist infrastructure work. 

Position – how you’ll contribute

  • Support the deployment, operation, and reliability of production services running on Kubernetes.
  • Monitor service health and investigate production incidents across distributed applications.
  • Participate in on-call support, incident response, root cause analysis, postmortems, and reliability improvements.
  • Troubleshoot application runtime, networking, and service-to-service issues in collaboration with engineering teams.
  • Support CI/CD, GitOps-based deployments, observability, and production monitoring.
  • Work within a client-directed backlog and established priorities.

Qualifications

Expectations – the experience you need

  • 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Production Engineering, or a closely related role, including strong recent hands-on experience supporting Kubernetes-based production services.
  • 3+ years of hands-on production Kubernetes experience strongly preferred. Kubernetes production operations, including deployment, scaling, rollout / rollback, resource tuning, and service-to-service troubleshooting
  • Strong production incident response experience, including on-call, runbooks, postmortems, and paging hygiene
  • Splunk experience for log aggregation, search, and production troubleshooting
  • Prometheus and Grafana experience, specifically building alert rules and dashboards, not only using existing dashboards
  • CI/CD and infrastructure-as-code for containerized deployments, including Helm and GitOps tools such as ArgoCD or Flux
  • Strong Linux and networking fundamentals, including DNS, load balancing, TCP / HTTP, HTTP/2, and Kubernetes networking
  • Production troubleshooting experience across Node.js and JVM/Java services, with strong depth in at least one runtime environment. Experience may include Node.js heap snapshots, CPU profiling, event-loop and memory analysis, as well as JVM GC log analysis, thread dumps, JVM tuning, and Java service latency investigation.
  • Service-to-service authentication experience, including mTLS, certificate rotation, certificate format conversion, and JWT-based service authentication
  • Very good spoken and written English. 

Additional skills – the edge you have

  • Web Components / Lit experience, to perform first-level debugging of UI-related issues
  • Server-side rendering or isomorphic runtime experience
  • Canary rollout / multi-version production operations
  • Distributed tracing and request-context correlation
  • KEDA or event-driven autoscaling
  • Experience with enterprise platform integration layers

Additional Information

Our offer – professional development, personal growth:

  • Flexible employment and remote work  
  • International projects with leading global clients 
  • International business trips  
  • Non-corporate atmosphere 
  • Language classes 
  • Internal & external training 
  • Private healthcare and insurance  
  • Multisport card 
  • Well-being initiatives 

Position at: Software Mind

Similar Jobs

17 Hours Ago
In-Office or Remote
Mid level
Mid level
Cloud • Security • Software • Cybersecurity
As a Penetration Tester, you will assess vulnerabilities in Akamai's products, perform penetration tests, provide remediation guidance, and collaborate with security teams.
Top Skills: AIBurpsuiteIdaproLlmsPenetration Testing
17 Hours Ago
In-Office or Remote
Mid level
Mid level
Cloud • Security • Software • Cybersecurity
Design enterprise and developer-facing products focusing on user-centered design, creating scalable experiences for complex systems, collaborating with cross-functional teams from concept to launch.
Top Skills: AIMlPrototypingSaaSUi DesignUx Design
17 Hours Ago
In-Office or Remote
Senior level
Senior level
Cloud • Security • Software • Cybersecurity
The Senior Software Engineer will design and implement scalable data streaming solutions, develop DevOps tools, and improve system usability and observability.
Top Skills: AnsibleBash ScriptingCi/CdDatadogDevOpsHadoopHelmHTTPJenkinsLinuxPrometheusPuppetRest ApiSparkSQLSshSslTcp/IpTerraform

What you need to know about the Chicago Tech Scene

With vibrant neighborhoods, great food and more affordable housing than either coast, Chicago might be the most liveable major tech hub. It is the birthplace of modern commodities and futures trading, a national hub for logistics and commerce, and home to the American Medical Association and the American Bar Association. This diverse blend of industry influences has helped Chicago emerge as a major player in verticals like fintech, biotechnology, legal tech, e-commerce and logistics technology. It’s also a major hiring center for tech companies on both coasts.

Key Facts About Chicago Tech

  • Number of Tech Workers: 245,800; 5.2% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: McDonald’s, John Deere, Boeing, Morningstar
  • Key Industries: Artificial intelligence, biotechnology, fintech, software, logistics technology
  • Funding Landscape: $2.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Pritzker Group Venture Capital, Arch Venture Partners, MATH Venture Partners, Jump Capital, Hyde Park Venture Partners
  • Research Centers and Universities: Northwestern University, University of Chicago, University of Illinois Urbana-Champaign, Illinois Institute of Technology, Argonne National Laboratory, Fermi National Accelerator Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account