Source Job

US

  • Lead the transformation of a diverse operations-heavy organization into a modern, AI-first Production Engineering function.
  • Own end-to-end reliability, performance, scalability, and security of NICE's global cloud, telecom, and datacenter platforms.
  • Drive adoption of software-first operational practices including automated recovery, infrastructure as code, and observability.

Cloud Infrastructure Distributed Systems DevOps SRE Automation

20 jobs similar to Vice President, Production Engineering

Jobs ranked by similarity.

$180,000–$220,000/yr
North America

  • Lead and modernize Sectigo's global infrastructure organization with a focus on reliability and operational maturity.
  • Develop a measurable operating model using SLAs, SLOs, and key metrics to drive improvement.
  • Drive automation, AI-enabled operations, and closer collaboration with engineering teams.

Sectigo is an innovative provider of certificate lifecycle management (CLM) solutions, helping large brands simplify digital trust. With over 700,000 customers including 65% of the Fortune 500, they emphasize a culture of support, excellence, and teamwork.

$114,800–$150,000/yr
US 4w PTO

  • Own complex production support escalations and ticket triage, providing hands-on troubleshooting and resolution alongside reliability work.
  • Partner with Software Engineering to investigate complex production issues, identify root causes and reliability risks, and drive permanent solutions.
  • Build observability across products, services, and critical customer workflows using meaningful metrics, logs, traces, dashboards, and actionable alerts.

Bloomerang provides a powerful giving platform and support for nonprofits to raise more, recruit more, and retain more. The company fosters a mission-driven culture built on core values of Simplify, Care and Act, and is home to innovative and skilled individuals.

$150,000–$185,000/yr
US Unlimited PTO

  • You will lead the reliability and operational evolution of our platform, building and improving system resiliency and establishing SLIs and SLOs.
  • You will partner with product engineering teams to own and operate their services, evolving observability platforms and strengthening incident practices.
  • You will contribute to day-to-day cloud infrastructure work alongside reliability specialty, including on-call rotation.

Rocket Money is a financial technology company that empowers people to live their best financial lives by providing insights and services to save time and money. The company runs hundreds of services in production, processing billions of transactions, and has a culture of reliability and innovation.

Global

  • Lead AI-assisted development as an ongoing experiment, adopting shared standards and building the team's discipline around clear context and human judgment.
  • Build, grow, and retain a strong engineering team while managing complex cross-team programs and holding tight to launch delivery deadlines.
  • Own reliability, cost, and change leadership, ensuring operational excellence and smooth transitions during reorganizations or acquisitions.

Sphera provides enterprise software and services that help companies manage and optimize their environmental, health, safety, and sustainability. They are a rapidly expanding team backed by Blackstone, guided by values of customer centricity, accountability, and collaboration.

US Unlimited PTO

  • Building and coaching a high-performing distributed team with a shared operating model.
  • Owning platform capabilities for provisioning, deployment, and operations of infrastructure.
  • Leading infrastructure migration towards a modern SaaS model with incremental delivery.

Totara is a global learning platform trusted by more than 1,500 organisations and 21 million users worldwide, offering flexible learning, compliance, and talent development solutions. With a distributed team across New Zealand, Australia, the UK, and the US, the company values diverse perspectives and offers flexible, hybrid working.

$240,000–$240,000/yr
US

  • Lead centralization of DevOps, SRE, database reliability, incident management, and developer experience practices.
  • Drive SLOs, observability, alerting, and on-call processes across teams.
  • Build the platform engineering function from the ground up and influence cross-cutting architecture.

First Due provides fire and EMS agencies with transformative, end-to-end software solutions to improve safety and effectiveness. The company offers a fully remote workplace with a comprehensive benefits package and opportunities for advancement.

EMEA

  • Lead and develop a global team of SRE leaders, managers, and engineers, driving reliability strategy and operating model.
  • Own and evolve observability capabilities across metrics, logs, traces, alerting, SLI/SLOs, and service health.
  • Drive cloud modernization initiatives, advancing containerization and Kubernetes-based operating models.

ServiceNow is the AI control tower for business reinvention, bringing together AI, data, and workflows to help 85% of the Fortune 500 work smarter. The company fosters an AI-native culture where technology and talent are unstoppable together.

US

  • Establish performance, throughput, latency, and capacity baselines for critical platform workflows.
  • Define and maintain SLOs, error budgets, dashboards, alerts, and reliability thresholds.
  • Lead load, stress, soak, spike, failure, and recovery testing in representative environments.

Tech Holding is a full-service consulting firm that delivers predictable outcomes and high-quality solutions to clients. The company was founded by experienced industry professionals who have held senior positions at startups to Fortune 50 firms, fostering a culture of deep expertise, integrity, transparency, and dependability.

UK

  • Design, build, and operate Kubernetes infrastructure for AI workloads using Terraform and GitOps.
  • Define SLOs, run incident response, and create runbooks for reliable AI platform operations.
  • Drive AI-specific observability, FinOps, and security practices across the platform.

We are an AI-native consulting partner working with clients like PayPal, adidas, and NatWest to build digital products and services. Our team of over 600 has scaled quickly, earning Great Place to Work-Certified status multiple years in a row.

  • Drive company-wide reliability strategy, standards, and best practices across LinkedIn Engineering.
  • Lead adoption of service criticality models to set reliability expectations based on business impact.
  • Partner with engineering teams to improve system design, reduce incident risk, and strengthen operational readiness.

LinkedIn is the world's largest professional network, built to create economic opportunity for every member of the global workforce. We foster a culture of trust, care, inclusion, and fun, investing in employee growth to transform the way the world works.

$100,000–$145,000/yr
US

  • Own and operate a production agentic AI platform on AWS, ensuring reliability and scaling.
  • Lead infrastructure automation and release management, driving best practices in security and compliance.
  • Collaborate with the platform team on agile ceremonies and proactively communicate status to stakeholders.

Inizio Evoke is a healthcare communications company dedicated to making health more human. As part of the larger Inizio network, it emphasizes a collaborative, inclusive culture where employees are encouraged to be their authentic selves.

$145,000–$177,000/yr
US

  • Build platform capabilities that enable engineering teams to deliver reliable software safely and efficiently.
  • Lead complex technical initiatives spanning cloud infrastructure, Kubernetes, observability, automation, and networking.
  • Design and implement solutions that improve availability, scalability, performance, and resilience of the platform.

Everbridge empowers enterprises and government organizations to anticipate, mitigate, respond to, and recover from critical events. The company focuses on building resilient systems and fostering a culture of ownership, continuous improvement, and operational excellence.

$120,000–$155,000/yr
Global

  • Own infrastructure as code across development, staging, and production environments
  • Build, maintain, and improve CI/CD pipelines for reliable and efficient deployments
  • Manage cloud infrastructure, establish scalable engineering practices, and lead incident response

CelebriOS is a software company building B2B SaaS products that help businesses make better decisions and streamline operations. The company has a remote-first working environment and a benefits package designed to support their team.

$114,700–$195,000/yr
North America

  • Lead enterprise-wide reliability and infrastructure projects with high autonomy, architecting scalable solutions and driving SRE best practices.
  • Partner cross-functionally with Engineering, Product, and Customer Success to align reliability goals with business objectives and communicate complex concepts to diverse audiences.
  • Provide tier 2/3 technical support to enterprise customers, conduct technical onboarding, and act as a trusted advisor for platform architecture.

Veza is the pioneer in identity security, providing an Access Graph platform that maps identity ecosystems across users, groups, roles, policies, and resources. With over 30 billion access permissions under management and now part of ServiceNow, Veza combines enterprise scale with security innovation.

United States

  • Own the CI/CD and configuration management architecture across the production fleet.
  • Define reference architectures and set progressive delivery standards like canary deployments and automated rollbacks.
  • Partner with Cloud Platform, Security, and Compliance to embed AI tooling into delivery workflows.

AlphaSense provides AI-driven market intelligence and search to help professionals make smarter decisions. The company has over 2,000 employees globally and serves a majority of the S&P 500.

US

  • Improve system availability, scalability, and resilience across Flowcode's platforms.
  • Manage and scale core AWS infrastructure through Infrastructure as Code (Terraform) and enhance disaster recovery.
  • Oversee monitoring, logging, and alerting infrastructure, and develop high-signal metrics and dashboards.

Flowcode is a technology company specializing in QR code and smart link solutions for offline-to-online engagement. The company is a growth-stage startup seeking high-performing individuals who thrive in a fast-paced, demanding environment.

India

  • Deliver the Cells and Organizations roadmap including cell provisioning, routing, data migration, and feature parity.
  • Lead the engineering organization by hiring, developing managers and senior ICs, and setting the technical bar.
  • Drive cross-functional leadership across product groups, infrastructure, and security, mainly asynchronously.

GitLab is an intelligent orchestration platform for DevSecOps that enables organizations to increase developer productivity and improve operational efficiency. With over 50 million registered users and a high-performance culture driven by values and continuous knowledge exchange, GitLab embraces AI as a core productivity multiplier.

UK

  • Lead Cloud Platform and SRE teams to scale securely and efficiently.
  • Drive infrastructure strategy, including Kubernetes (GKE) clusters and developer platform.
  • Champion SRE culture with SLOs, error budgets, and observability.

Prolific builds human data infrastructure for AI development. The company is a fast-growing, mission-driven organization with cross-functional teams and a strong ownership culture.

India

  • Lead cloud infrastructure strategy for resilient, secure, and cost-efficient multi-account cloud environments across AWS, Azure, and GCP.
  • Drive Kubernetes excellence as a technical authority for production clusters including EKS and AKS.
  • Advance AI-enabled operations by introducing LLM-based tooling and agentic workflows to improve infrastructure development and operational efficiency.

Jobgether is a platform that uses AI-powered matching to connect candidates with hiring companies. They are a technology company focused on improving the hiring process through automation and data analysis.

UK

  • Enable effective execution with Quality and Speed, in partnership with the team's Product Manager.
  • Ensure 3+ 9s availability of Dedicated infrastructure, ensuring security and automating for maximum scalability.
  • Provide clear direction, meaningful feedback and foster an environment where meaningless toil gets ruthlessly automated.

GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity and improve operational efficiency. With over 50 million users and 50% of Fortune 100, GitLab fosters a high-performance culture driven by values.