Source Job

US

  • Own the most challenging customer technical problems with deep-dive troubleshooting and analysis.
  • Act as the customer's advocate to ensure their issues influence product and engineering roadmaps.
  • Author technical content and lead internal tooling development with a strong AI focus.

Python Go REST APIs Linux Kubernetes

20 jobs similar to Senior Escalation Engineer

Jobs ranked by similarity.

Brazil

  • Provide technical support experience for Wiz product, troubleshooting issues with debugging, networking, and system administration.
  • Collaborate across teams to own and solve customer technical issues, escalating when necessary.
  • Design and implement automation to scale support offering and participate in on-call rotation.

We are redefining security for the AI era, connecting code, cloud, and runtime into a single platform. As one of the fastest-growing startups, powered by Google, we are trusted by over 65% of the Fortune 100 and scan over 230 billion files daily, with a culture that values world-class talent and collaboration.

$150,000–$165,000/yr
US Unlimited PTO

  • Design and manage high-availability platforms using Kubernetes, Terraform, and Ansible with native-AI capabilities.
  • Develop and operate the observability stack: Grafana, Mimir, Loki, Tempo, and Prometheus on Kubernetes via GitLab CI/CD.
  • Build automation scripts in Python, maintain GitOps pipelines, and mentor mid-level engineers.

Flexential builds and operates critical IT platforms including observability, DevOps, and ITSM technologies. The company fosters a collaborative engineering culture and values diversity.

$170,000–$235,000/yr
US

  • Design and implement backend services for licensing, entitlements, feature access, and usage limits across NodeZero's product and APIs.
  • Build and evolve provisioning, admin experience, MSP/MSSP capabilities, and audit logging for a multi-tenant SaaS platform.
  • Operate production services with monitoring, incident response, and a high bar for design quality and test coverage.

Horizon3 is a fast-growing, remote cybersecurity company that helps organizations proactively find, fix, and verify exploitable attack vectors through its NodeZero autonomous pentesting platform. The team is a fusion of former special operations cyber operators and startup engineers, fostering a culture of respect, collaboration, ownership, and results.

  • Troubleshoot complex technical issues for Premium Support customers, including API failures, integration challenges, and production incidents.
  • Provide end-to-end technical ownership by analyzing logs, reproducing errors, and collaborating with Engineering on root cause analysis.
  • Develop deep understanding of customer architectures to anticipate risks and improve workload resilience through automation and monitoring.

OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. They push the boundaries of AI systems and seek to safely deploy them to the world through their products, embracing diverse perspectives and experiences.

$135,000–$150,000/yr
US

  • Operate, scale, and troubleshoot Bitsight's SaaS cloud infrastructure with focus on reliability, efficiency, and security.
  • Tackle complex system-level designs and proactively anticipate performance and scalability issues.
  • Pioneer self-optimizing infrastructure systems using AI, ensuring manual and staging validation before production deployment.

Bitsight is a cyber risk management leader transforming how companies manage exposure, performance, and risk. Over 3,500 customers and 600 teammates work across Boston, Raleigh, New York, Lisbon, Singapore, and remote locations.

New Zealand

  • Provide expert guidance on deployment and operational best practices for the Wiz platform.
  • Develop trusted advisor relationships with customer stakeholders to drive adoption and satisfaction.
  • Advocate for customer needs across departments and identify opportunities for expanding Wiz usage.

Wiz is redefining security for the AI era by enabling teams to secure cloud and AI applications with a single shared context. They are one of the fastest-growing startups, trusted by over 65% of the Fortune 100, and have a culture that values world-class talent and creative freedom.

Canada USA Unlimited PTO

  • Own the observability, logging and alerting for Kubernetes clusters and critical workloads.
  • Build and maintain automation for lifecycle management of Kubernetes clusters.
  • Identify and root-fix reliability bottlenecks before they become incidents.

Wrapbook is an AI platform for production finance, built for feature films and TV, trusted by Netflix and Paramount. Backed by top investors, our team of over 350 employees uses AI to transform how finance teams work.

$259,200–$388,800/yr

  • Architect and maintain cloud infrastructure for F5's AI security platform, ensuring scalability and reliability.
  • Design CI/CD pipelines, manage Kubernetes workloads, and implement serverless technologies.
  • Build comprehensive monitoring, security, and compliance practices across the B2B SaaS solution.

F5 is a global leader in application delivery and security, helping organizations create, secure, and run applications. With over 6,400 employees and 553 patents, F5 serves more than 23,000 customers across 170 countries and fosters a human-first, inclusive culture.

India

  • Design, build, and ship production services, APIs, and user-facing interfaces.
  • Build and operate production AI systems including RAG, fine-tuning, and inference optimization.
  • Architect AWS/GCP environments with Kubernetes and Terraform and control cloud/AI costs.

Motive empowers people who run physical operations with tools to make their work safer, more productive, and more profitable. Serving nearly 100,000 customers across industries, the company values a diverse and inclusive workplace.

Europe

  • Lead end-to-end technical engagements: Partner directly with engineering teams to diagnose, unblock, and resolve complex infrastructure challenges.
  • Execute critical migrations: Develop reference implementations, tooling, and guidance to transition teams off deprecated systems seamlessly.
  • Accelerate platform adoption: Act as primary technical contact for new teams onboarding to Planet's core infrastructure.

Planet designs, builds, and operates the largest constellation of imaging satellites in history, delivering unprecedented dataset via a cloud-based platform for commercial, environmental, and humanitarian sectors. A global company with offices in the US, Europe, and Slovenia, Planet values a people-centric culture and community.

Europe

  • Operate and improve Linux infrastructure and Kubernetes clusters across bare-metal, virtualized, and on-premise environments.
  • Design and maintain complex networking architectures and automation using Ansible, Bash, Python, and GitOps.
  • Lead incident response, define SLOs, and build observability platforms with Prometheus, Grafana, and ELK.

Jobgether is a platform that connects job seekers with opportunities through an AI-powered matching process. The company fosters a remote-first culture and emphasizes autonomy and ownership for engineers.

$90,000–$110,000/yr
US EMEA

  • Act as senior technical resource and final escalation point for strategic and VIP customers, owning complex issues across Kubernetes, GPU, and enterprise stack.
  • Train and mentor Technical Support Engineers in advanced Linux troubleshooting and customer architectures.
  • Author advanced troubleshooting documentation and drive incident resolution through root cause analysis.

Vultr makes high-performance cloud infrastructure easy to use and affordable for enterprises and AI innovators worldwide. It is the world's largest privately-held cloud infrastructure company with 33 data centers and hundreds of thousands of active customers, committed to growth and employee investment.

Europe

  • Design and build complex, hands-on lab scenarios that showcase Sysdig across containers, Kubernetes, Linux and public cloud.
  • Own the automation and reliability of the training platform, including provisioning and scripting.
  • Design and deliver engaging technical training, workshops and demos for diverse audiences.

Sysdig stops attacks in real-time by detecting changes in cloud security risk with runtime insights and open source Falco. It is a well-funded startup with a large enterprise customer base, recognized as a best place to work.

$143,200–$243,400/yr
North America

  • Contribute to infrastructure automation and operational resilience across hybrid cloud and data center operations.
  • Implement closed-loop auto-remediation systems and SRE tooling to reduce manual intervention and incident resolution time.
  • Develop and maintain SLO frameworks, alerting policies, and Infrastructure-as-Code pipelines for reproducible deployments.

ServiceNow is the AI control tower for business reinvention, helping 85% of the Fortune 500 work smarter, faster, and better. They foster an AI-native culture where technology and talent are unstoppable together.

Europe

  • Own the technical relationship and be the trusted advisor for CTOs and platform teams.
  • Architect real solutions that translate customer constraints into deployable AI infrastructure.
  • Lead proofs of concept under real conditions to demonstrate operational fit.

Mirantis is a Kubernetes-native AI infrastructure company that helps organizations build scalable and secure infrastructure for modern AI workloads. The company has an installed base of 1,500 enterprise customers and values open source innovation, collaboration, and continuous growth.

$12,500–$20,800/mo
Turkey

  • Own the technical architecture and evolution of core infrastructure.
  • Engineer for scale and performance through capacity modeling and bottleneck diagnosis.
  • Participate in on-call rotation and drive technical recovery during incidents.

Sezzle is a fintech company that revolutionizes shopping through interest-free installment plans, blending cutting-edge technology with financial empowerment. It has a dynamic and innovative team culture focused on shaping the future of fintech and retail.

$207,200–$304,600/yr
US

  • Lead engineering teams for Porting, Hosting, and A2P Messaging, owning delivery roadmap and operational excellence.
  • Drive adoption of AI coding assistants and measure impact on team velocity, fostering psychological safety.
  • Recruit, develop, and retain a world-class engineering team, promoting a culture of candid feedback and career growth.

Twilio is shaping the future of communications, delivering innovative solutions to hundreds of thousands of businesses and empowering millions of developers worldwide. They are a remote-first company with a strong culture of connection and global inclusion, employing a diverse team making a global impact daily.

$120,000–$137,000/yr
US Unlimited PTO 18w maternity 12w paternity

  • Triage and resolve complex customer issues involving Chainguard Images using Docker, Kubernetes, GitHub, Helm, and Terraform.
  • Escalate to Engineering when needed, keep customers and SLAs happy, and communicate clearly across technical and non-technical audiences.
  • Document issues, guide customers to resolution, and participate in on-call rotation for after-hours and weekend support.

Chainguard is the trusted source for open source, delivering hardened, secure, and production-ready builds of open source software to help organizations build faster, stay compliant, and eliminate risk. The company is venture-backed by leading investors and serves Fortune 500 enterprises and global industry leaders, with a remote-first culture focused on customer obsession and serious work.

US

  • Investigate and resolve technical issues reported by enterprise customers, owning the issue from triage to resolution.
  • Analyze logs, reproduce issues, and dig into system behavior to identify root causes before escalating to Engineering.
  • Communicate clearly with customers and internal teams throughout the support process.

Unframe is an AI-first startup helping the world’s largest enterprises bring LLM-powered applications to life in days. We are backed by Bessemer, Craft, and TLV Partners with $100M in funding, a fast-growing company working with Fortune 500 customers globally.

$150,000–$175,000/yr
United States

  • Design, build, and maintain automation and tooling to reduce operational toil.
  • Actively participate in the incident-management lifecycle, including detection, escalation, mitigation, and post-incident review.
  • Provide an SRE point of view on capacity planning, resilience testing, and modernization of legacy workloads.

Seismic is the Go-To-Market Performance company, helping organizations turn strategy into revenue through an AI-powered revenue execution platform. Trusted by 2,500 organizations and over 3.5 million users globally, Seismic is headquartered in San Diego with offices across North America, Europe, and Asia-Pacific, fostering an inclusive culture.