Source Job

$180,000–$220,000/yr
North America

  • Lead and modernize Sectigo's global infrastructure organization with a focus on reliability and operational maturity.
  • Develop a measurable operating model using SLAs, SLOs, and key metrics to drive improvement.
  • Drive automation, AI-enabled operations, and closer collaboration with engineering teams.

Cloud Computing Site Reliability Engineering DevOps Automation

20 jobs similar to Vice President, Infrastructure & Operations

Jobs ranked by similarity.

US

  • Lead the transformation of a diverse operations-heavy organization into a modern, AI-first Production Engineering function.
  • Own end-to-end reliability, performance, scalability, and security of NICE's global cloud, telecom, and datacenter platforms.
  • Drive adoption of software-first operational practices including automated recovery, infrastructure as code, and observability.

NICE provides software products used by 25,000+ global businesses to deliver extraordinary customer experiences, fight financial crime, and ensure public safety. With over 8,500 employees across 30+ countries, the company fosters a culture of ambition, game-changing innovation, and high standards.

$100,000–$145,000/yr
US

  • Own and operate a production agentic AI platform on AWS, ensuring reliability and scaling.
  • Lead infrastructure automation and release management, driving best practices in security and compliance.
  • Collaborate with the platform team on agile ceremonies and proactively communicate status to stakeholders.

Inizio Evoke is a healthcare communications company dedicated to making health more human. As part of the larger Inizio network, it emphasizes a collaborative, inclusive culture where employees are encouraged to be their authentic selves.

$282,000–$322,000/yr
US

  • Lead adoption of advanced AI technologies across operations to improve efficiency and accelerate product outcomes.
  • Provide leadership and strategy for automation, stability, and availability of technology infrastructure leveraging cloud and on-premise technologies.
  • Drive innovation across security, data privacy, data and development operations and infrastructure.

Tebra is the only all-in-one EHR+ platform built exclusively for independent healthcare practices. With over 42,000 private practices trusting their platform, they aim to streamline operations, increase revenue, and reduce burnout.

US Unlimited PTO

  • Building and coaching a high-performing distributed team with a shared operating model.
  • Owning platform capabilities for provisioning, deployment, and operations of infrastructure.
  • Leading infrastructure migration towards a modern SaaS model with incremental delivery.

Totara is a global learning platform trusted by more than 1,500 organisations and 21 million users worldwide, offering flexible learning, compliance, and talent development solutions. With a distributed team across New Zealand, Australia, the UK, and the US, the company values diverse perspectives and offers flexible, hybrid working.

$150,000–$185,000/yr
US Unlimited PTO

  • You will lead the reliability and operational evolution of our platform, building and improving system resiliency and establishing SLIs and SLOs.
  • You will partner with product engineering teams to own and operate their services, evolving observability platforms and strengthening incident practices.
  • You will contribute to day-to-day cloud infrastructure work alongside reliability specialty, including on-call rotation.

Rocket Money is a financial technology company that empowers people to live their best financial lives by providing insights and services to save time and money. The company runs hundreds of services in production, processing billions of transactions, and has a culture of reliability and innovation.

$114,700–$195,000/yr
North America

  • Lead enterprise-wide reliability and infrastructure projects with high autonomy, architecting scalable solutions and driving SRE best practices.
  • Partner cross-functionally with Engineering, Product, and Customer Success to align reliability goals with business objectives and communicate complex concepts to diverse audiences.
  • Provide tier 2/3 technical support to enterprise customers, conduct technical onboarding, and act as a trusted advisor for platform architecture.

Veza is the pioneer in identity security, providing an Access Graph platform that maps identity ecosystems across users, groups, roles, policies, and resources. With over 30 billion access permissions under management and now part of ServiceNow, Veza combines enterprise scale with security innovation.

India

  • Lead cloud infrastructure strategy for resilient, secure, and cost-efficient multi-account cloud environments across AWS, Azure, and GCP.
  • Drive Kubernetes excellence as a technical authority for production clusters including EKS and AKS.
  • Advance AI-enabled operations by introducing LLM-based tooling and agentic workflows to improve infrastructure development and operational efficiency.

Jobgether is a platform that uses AI-powered matching to connect candidates with hiring companies. They are a technology company focused on improving the hiring process through automation and data analysis.

EMEA

  • Lead and develop a global team of SRE leaders, managers, and engineers, driving reliability strategy and operating model.
  • Own and evolve observability capabilities across metrics, logs, traces, alerting, SLI/SLOs, and service health.
  • Drive cloud modernization initiatives, advancing containerization and Kubernetes-based operating models.

ServiceNow is the AI control tower for business reinvention, bringing together AI, data, and workflows to help 85% of the Fortune 500 work smarter. The company fosters an AI-native culture where technology and talent are unstoppable together.

$53,300–$119,850/yr
Global Unlimited PTO 16w maternity 16w paternity

  • Lead the discovery, design, and delivery of complex reliability and infrastructure initiatives, translating ambiguous problems into robust technical solutions.
  • Define and operate reliability practices including SLOs, SLIs, error budgets, alerting strategies, and observability standards.
  • Manage and scale production Kubernetes environments, build cloud infrastructure on AWS, and mentor less-senior engineers.

Jobgether is an AI-powered recruitment platform that connects candidates with partner companies through automated matching. They are a globally distributed organization focused on fair, efficient hiring processes.

  • Drive company-wide reliability strategy, standards, and best practices across LinkedIn Engineering.
  • Lead adoption of service criticality models to set reliability expectations based on business impact.
  • Partner with engineering teams to improve system design, reduce incident risk, and strengthen operational readiness.

LinkedIn is the world's largest professional network, built to create economic opportunity for every member of the global workforce. We foster a culture of trust, care, inclusion, and fun, investing in employee growth to transform the way the world works.

UK

  • Design, build, and operate Kubernetes infrastructure for AI workloads using Terraform and GitOps.
  • Define SLOs, run incident response, and create runbooks for reliable AI platform operations.
  • Drive AI-specific observability, FinOps, and security practices across the platform.

We are an AI-native consulting partner working with clients like PayPal, adidas, and NatWest to build digital products and services. Our team of over 600 has scaled quickly, earning Great Place to Work-Certified status multiple years in a row.

$125,000–$135,000/yr

  • Manage patching, updates, and lifecycle for Azure VMs, support on-premises infrastructure transition, and maintain operational documentation.
  • Act as escalation point for complex infrastructure issues, troubleshooting cloud, server, networking, identity, and application problems.
  • Develop PowerShell scripts, automation solutions, and leverage AI-assisted tools to improve operational efficiency.

AIP Publishing is a leading publisher of the physical sciences, accelerating scientific discovery and reimagining scholarly publishing. It is a wholly owned not-for-profit subsidiary of the American Institute of Physics, fostering a collaborative and creative atmosphere that maximizes individual contributions.

North America

  • Lead the technical direction of Device Infrastructure Services, including endpoint management and GitOps transformation.
  • Drive adoption of AI-enabled operations and responsible agentic AI across the device team.
  • Act as escalation point for high-severity endpoint incidents and partner across IT, Security, and product teams.

1Password is a cybersecurity company that provides password management and unified access management solutions. They have over 180,000 businesses as customers, have surpassed $400M in ARR, and emphasize a collaborative, human-centric culture.

Poland

  • Lead the Site Reliability Engineering team to drive reliability and scalability of XTB's systems.
  • Define and drive SRE platform strategy, incident management, and observability engineering.
  • Mentor team members, foster collaboration, and ensure operational excellence.

XTB is a global investment company offering innovative technological solutions for managing finances through an intuitive app, used by over one million users worldwide. It is a certified Great Place to Work with a focus on technical excellence and collaboration.

Canada

  • Manage team performance, career development, and project prioritization while driving a culture of automation.
  • Drive initiatives with partner teams to improve infrastructure reliability and act as crisis management.
  • Analyze existing processes to drive continuous improvement and efficiencies.

ServiceNow is the AI control tower for business reinvention, helping 85% of the Fortune 500 work smarter with an intelligent cloud platform. We are building an AI-native culture where technology and talent are unstoppable together, serving over 8,100 customers.

$205,000–$230,000/yr
US Unlimited PTO

  • Lead security monitoring, incident response, and threat hunting across cloud and AI-enabled environments.
  • Establish operational priorities, metrics, and playbooks based on organizational risk.
  • Drive responsible adoption of AI-assisted detection and response capabilities.

Backblaze is a cloud storage and backup provider that helps customers protect their data across over 175 countries. The company fosters a culture centered on fairness, goodness, and work-life balance, with a strong commitment to diversity and inclusion.

$151,000–$242,000/yr
US

  • Ensure reliability, scalability, and operational excellence of analytics and data systems.
  • Provide technical leadership and direction to an offshore contract team.
  • Drive incident response, automation, and data governance initiatives.

Workiva provides an AI-powered platform that unifies finance, risk, and sustainability for complex organizations. It is a large enterprise with a collaborative and innovative culture centered on data integrity and trust.

US

  • Orchestrate enterprise-wide Agentic AI and AIOps initiatives, tracking milestones and coordinating deployment of multi-agent systems.
  • Lead operational delivery for observability, self-healing automation, and security remediation across multi-cloud environments.
  • Establish executive reporting frameworks communicating delivery health, operational improvements, and measurable business value.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies through automated matching. We are committed to an open, respectful, and inclusive environment where employees are empowered to contribute and grow.

US Australia 18w maternity 18w paternity

  • Own the full technology and security remit, including enterprise infrastructure, security, identity, applications, and workplace technology.
  • Lead and grow three functions: IT & Security Operations, Cloud Platform Engineering, and GRC.
  • Mature security operations with automation and serve as customer zero for the company's own platform.

UpGuard builds a Cyber Risk Posture Management platform that integrates security ratings, threat intel, and agentic AI to help organizations manage cyber risk. They are a certified Great Place to Work with a lean, high-growth culture and a global remote team.

$145,000–$260,000/yr
US Canada Unlimited PTO

  • Design, build, and optimize multi-region, high-availability AWS infrastructure.
  • Drive resiliency and automation using GitOps, modern CI/CD, and Infrastructure as Code.
  • Build end-to-end telemetry and own incident management to harden reliability.

VGS is the world's leader in payment tokenization, trusted by the most innovative AI and Fortune 500 companies. They are a remote-first company with a culture of ownership, collaboration, and continuous learning.