Source Job

Global

  • Own Kubernetes deployment and operational health, including scaling, rollout/rollback, and resource tuning for framework services.
  • Build and maintain production observability with Grafana dashboards and Prometheus alerting across SSR and Glide platform layers.
  • Diagnose and resolve Node.js and JVM production incidents, including event-loop stalls, heap growth, and GC pressure.

Kubernetes Node.js Prometheus Grafana

20 jobs similar to Senior Site Reliability Engineer - AI Experience Framework

Jobs ranked by similarity.

Global

  • Support deployment, operation, and reliability of production services on Kubernetes.
  • Monitor service health, investigate production incidents, and participate in on-call and postmortems.
  • Troubleshoot application runtime, networking, and service-to-service issues across Node.js and JVM.

Software Mind develops innovative solutions for global companies, partnering with tech giants and unicorns on transformative projects. They foster cross-functional engineering teams with a culture of openness, respect, and passion, combining employment with enjoyment.

Global 7w PTO

  • Lead the design and operation of LivePerson's observability platforms across logs, metrics, traces, alerting, and synthetic monitoring.
  • Own large-scale observability pipelines using technologies like Elastic Cloud, Grafana, Prometheus, and Kafka.
  • Provide technical leadership and mentorship while driving best practices in DevOps, cloud engineering, and observability.

LivePerson is a leader in trusted enterprise conversational AI and digital transformation, powering nearly a billion conversational interactions every month. The company is recognized as the #1 Most Innovative AI Company by Fast Company and fosters a diverse, inclusive culture that empowers employees globally.

Global

  • Support the deployment, operation, and maintenance of the Karuna service running on Kubernetes.
  • Monitor production environments to ensure high availability, reliability, and performance.
  • Investigate, troubleshoot, and resolve production incidents, performing root cause analysis.

Software Mind develops solutions that make an impact for companies around the globe. They build cross-functional engineering teams with a culture of openness, respect, grit, and enjoyment.

Europe US LATAM 4w PTO

  • Own observability for critical product journeys, defining SLIs/SLOs and building metrics, dashboards, and alerts.
  • Act as first responder for production incidents, investigating signals and mitigating issues independently.
  • Work within a cross-functional squad of 6-8 engineers to improve reliability, monitoring, and incident response processes.

Feeld is a dating app creating a safer and more inclusive space for exploring relationships and sexuality. They have a distributed engineering team of around 50 people across Europe and the US, working in small autonomous squads.

US Unlimited PTO

  • Serve as the first responder for production incidents, triaging and resolving issues.
  • Monitor application health and system availability using Datadog.
  • Develop automation scripts using Python or PowerShell to improve operational efficiency.

NationsBenefits is a Healthcare Fintech provider of supplemental benefits, flex cards, and member engagement solutions. Recognized as one of the fastest-growing companies in America, it offers a fulfilling work environment with career advancement opportunities across multiple locations in the US, South America, and India.

$235,000–$275,000/yr

  • Define and execute the technical strategy for Observability & Alerting, Platform Infrastructure, and operational excellence.
  • Lead the evolution of reliable, scalable, secure, and efficient cloud platforms and distributed systems.
  • Champion SLIs, SLOs, error budgets, capacity planning, and automation across the service lifecycle.

Filevine is a Legal AI company delivering a unified platform for legal work, powered by LOIS (Legal Operating Intelligence System). The company is rapidly growing, recognized by Deloitte and Inc. as one of the most innovative and fastest-growing technology companies.

US

  • Define and execute the technical strategy for observability, platform infrastructure, and operational excellence.
  • Lead the design and evolution of scalable, secure, reliable cloud-native platforms and distributed systems.
  • Establish reliability best practices including SLIs, SLOs, error budgets, and automation initiatives.

The company is a technology organization that builds and operates large-scale cloud infrastructure. It fosters a collaborative culture centered on innovation, ownership, and impact.

$140,000–$170,000/yr
US

  • Deploy and operate Blitzy's self-hosted platform within a customer-controlled, secure cloud environment.
  • Own the Kubernetes-based deployment, releases, upgrades, capacity planning, and performance benchmarking.
  • Serve as the on-account technical presence, partnering with customer infrastructure and security teams.

We are an AI software development platform that autonomously builds custom software for enterprises. Backed by tier 1 investors and led by two co-founders, we are one of the fastest-growing U.S. companies with a culture of speed and customer focus.

Global

  • Own the reliability of Supabase's deployment and release systems against clear SLOs and error budgets.
  • Standardize and instrument pre-production deployment workflows for trustworthy signal.
  • Drive disaster-recovery readiness and reduce mean-time-to-detect and recover for deploy-related incidents.

Supabase is the Postgres development platform providing a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. They are a globally distributed team of ~400 members across 60+ countries, with over $1B raised and 540,000+ community members.

APAC

  • Design, build, and maintain software, APIs, and automation to enhance platform reliability and observability.
  • Support monitoring, reliability, and continuous improvement in Kubernetes-based environments with a focus on Datadog.
  • Integrate observability into CI/CD pipelines and automate operational tasks using scripting languages like Python.

Europe

  • Lead technical operations for large-scale AI infrastructure environments powered by NVIDIA GPUs and Kubernetes.
  • Act as a senior escalation point for critical incidents and drive root cause analysis and long-term corrective actions.
  • Mentor team members and shape operational standards, automation, and reliability practices for next-generation platform services.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. Mirantis serves many of the world’s leading enterprises, including Adobe, DocuSign, Liberty Mutual, PayPal, Reliance Jio, Societe Generale, Splunk, and Volkswagen.

UK

  • Design, build, and operate Kubernetes infrastructure for AI workloads using Terraform and GitOps.
  • Define SLOs, run incident response, and create runbooks for reliable AI platform operations.
  • Drive AI-specific observability, FinOps, and security practices across the platform.

We are an AI-native consulting partner working with clients like PayPal, adidas, and NatWest to build digital products and services. Our team of over 600 has scaled quickly, earning Great Place to Work-Certified status multiple years in a row.

Global Unlimited PTO

  • Set reliability strategy and SLO culture that scales across engineering teams.
  • Own platform architecture, event-driven messaging, and observability for a global payments platform.
  • Lead chaos engineering, incident response, and mentorship for the most complex production challenges.

Yuno is an AI-native operating system for global commerce, connecting merchants to pay-ins, payouts, fraud prevention, and stablecoins via a single API. It powers payment infrastructure for global brands like McDonald's and GoFundMe, with a culture of remote work and AI-driven innovation.

Brazil

  • Lead technical and managerial direction for the SRE team, defining reliability, observability, and operational excellence strategy.
  • Coordinate critical incident responses and root cause analysis, collaborating with architecture, development, security, and product teams.
  • Drive automation, continuous improvement, and adoption of SRE, DevOps, and Platform Engineering best practices.

Experian is a global data and technology company that drives opportunities for people and businesses worldwide. With 25,200 employees in 32 countries, it has a people-centric, inclusive culture recognized by awards such as World's Best Workplaces™ 2025.

US Unlimited PTO

  • Resolve complex escalations as the final authority, using code-level debugging and architectural investigation.
  • Partner with customer technical leads to diagnose production issues, ensuring transparency and rapid resolution.
  • Own end-to-end P1 resolution and deliver clear, actionable post-incident analysis.

TensorWave delivers a versatile cloud platform for AI compute at scale, eliminating infrastructure barriers. The company fosters a culture of innovation and reliability, empowering builders to focus on breakthrough AI.

Europe US

  • Design and architect observability solutions leveraging OpenTelemetry, Kubernetes, and cloud-native technologies.
  • Develop and execute Proofs of Concept (POCs) that highlight Dash0's differentiated technical capabilities.
  • Deliver engaging technical demos and presentations tailored to engineering and executive audiences.

Dash0 is building an OpenTelemetry-native observability platform that eliminates vendor lock-in and provides transparent pricing. Backed by top-tier investors including Balderton Capital, Accel and Cherry Ventures, the company has a collaborative, fast-moving team culture with a builder mindset.

US

  • Provide solutions to customers to make them successful using our products.
  • Troubleshoot customer environments and engage in active triaging with customers.
  • Participate in on-call rotation for weekend coverage.

Astronomer empowers data teams to bring mission-critical software, analytics, and AI to life with its unified DataOps platform, Astro, powered by Apache Airflow. Trusted by more than 800 enterprises, the company fosters a diverse and inclusive culture as an equal opportunity employer.

North America Unlimited PTO

  • Deploy and scale MCP-based AI agents on Kubernetes for enterprise customers across the US-East and EMEA regions.
  • Lead complex technical engagements, build reusable deployment patterns, and mentor engineers on the team.
  • Shape product roadmap by feeding back field insights from regulated industries and defining regional engagement standards.

Stacklok builds the control plane for enterprise AI agents, enabling organizations to run, govern, and secure them on Kubernetes and private cloud. Founded by two Kubernetes creators, the company is already adopted by leading tech and regulated industries, fostering a collaborative, AI-maximalist culture with deep open-source roots.

US

  • Design, build, and maintain robust, scalable, and secure infrastructure systems supporting Laurel's AI-driven platform.
  • Manage and optimize cloud infrastructure (AWS and Azure), Kubernetes orchestration, and CI/CD pipelines to increase deployment frequency and reliability.
  • Implement comprehensive observability, monitoring, and alerting to maintain system health and partner with engineering teams to optimize performance and cost-efficiency.

Laurel is an AI Time platform for professional services firms, automating work time capture and connecting time data to business outcomes for clients like EY and Crowell & Moring. The company comprises top AI, product, and engineering talent, is VC-backed by Google Ventures and IVP, and fosters an inclusive, ambitious culture.

  • Serve as a trusted technical advisor guiding customers through their observability journey.
  • Design and guide customer observability maturity strategies to improve reliability and operational visibility.
  • Provide expert troubleshooting and technical recommendations to resolve complex challenges.

Jobgether is a platform that uses AI to match candidates with jobs. They focus on remote work and have a collaborative culture built around transparency, autonomy, and trust.