Manage the ticket queue, prioritize and resolve requests, and identify recurring categories for automation.
Participate in rotating on-call and incident response, troubleshooting and documenting issues in real time.
Build and maintain monitoring dashboards (Tableau, Superset, Grafana) to track service health and data quality.
Airbnb is a global community marketplace that connects hosts with guests for unique stays and experiences. With over 5 million hosts and 2 billion guest arrivals, the company fosters a culture of inclusion and belonging, emphasizing innovation and engagement.
Lead the design and operation of LivePerson's observability platforms across logs, metrics, traces, alerting, and synthetic monitoring.
Own large-scale observability pipelines using technologies like Elastic Cloud, Grafana, Prometheus, and Kafka.
Provide technical leadership and mentorship while driving best practices in DevOps, cloud engineering, and observability.
LivePerson is a leader in trusted enterprise conversational AI and digital transformation, powering nearly a billion conversational interactions every month. The company is recognized as the #1 Most Innovative AI Company by Fast Company and fosters a diverse, inclusive culture that empowers employees globally.
Lead automation of release processes including CI/CD, bootstrapping, and configuration management for internal engineering teams.
Work with diverse internal teams to implement requirements and maintain the Internal Engineering Platform.
Participate in an on-call rotation to support platform tooling and ensure system health.
Grafana Labs is the company behind the open source observability platform Grafana, providing a fully managed observability cloud called Grafana Cloud. With over 1,600 team members across 40+ countries, the company fosters a global collaborative culture rooted in open-source values and transparency.
Design, scale, and maintain enterprise monitoring and alerting ecosystems across multi-cloud and native systems.
Bridge development and operations to ensure high availability, performance tuning, and deep visibility.
Automate infrastructure and build robust observability pipelines using cloud-native tools like Prometheus, Grafana, and GCP.
Ontrac Solutions is a leading technology consulting firm specializing in cutting-edge solutions that drive business transformation. Their team is committed to innovation, collaboration, and excellence, empowering clients to succeed in an evolving digital landscape.
Evolve DevOps and platform engineering, including build pipelines, monitoring, infrastructure as code, and cost optimization.
Design and maintain secure, scalable CI/CD pipelines that support rapid iterations without sacrificing stability.
Improve observability via Datadog and GCP, implement FinOps governance, and enable developer productivity through internal tooling.
Sardine is the leading agentic risk platform for fighting financial crime, unifying data across risk teams to stop fraud in real time and prevent AI-driven attacks. It is a remote-first company with hubs in multiple locations, hiring talented self-motivated individuals who value performance over hours worked.
Serve as the primary technical point of contact for a portfolio of Grafana customers, designing and guiding their observability maturity journey.
Conduct regular technical reviews, health checks, and root cause analysis to drive adoption and ensure customer success.
Act as the voice of the customer internally, shaping product feedback and roadmap priorities while building long-term strategic relationships.
Grafana Labs builds the open source observability platform Grafana and its fully managed cloud service. The company has over 1,600 team members across 40+ countries, serving more than 7,000 customers including major enterprises, and fosters a remote-first, transparent, and innovation-driven culture.
Design, implement, and support automation, deployment, monitoring, and operational reliability across cloud and hosted application environments.
Partner with software engineering, site reliability, security, and infrastructure teams to improve continuous integration and continuous delivery processes and infrastructure automation.
Apply AI across CI/CD, observability, incident response, and infrastructure operations to optimize builds, detect failures faster, and recommend cost improvements.
Granicus provides cloud-based solutions for government communications, website design, meeting management, records management, and digital services. They serve 5,500 government agencies and over 300 million citizen subscribers, and have been recognized on the GovTech 100 list and as a best company to work for on BuiltIn.
Define and execute the technical strategy for observability, platform infrastructure, and operational excellence.
Lead the design and evolution of scalable, secure, reliable cloud-native platforms and distributed systems.
Establish reliability best practices including SLIs, SLOs, error budgets, and automation initiatives.
The company is a technology organization that builds and operates large-scale cloud infrastructure. It fosters a collaborative culture centered on innovation, ownership, and impact.
Serve as a trusted technical advisor guiding customers through their observability journey.
Design and guide customer observability maturity strategies to improve reliability and operational visibility.
Provide expert troubleshooting and technical recommendations to resolve complex challenges.
Jobgether is a platform that uses AI to match candidates with jobs. They focus on remote work and have a collaborative culture built around transparency, autonomy, and trust.
Design, build, and operate core cloud infrastructure on AWS, including compute, networking, and container orchestration.
Own the CI/CD platform used across engineering teams, including build pipelines, environment promotion, and progressive rollout.
Build and maintain the observability stack across the organization, including logging, metrics, distributed tracing, and alerting.
RxSense is a healthcare technology company that provides platforms and solutions to improve the management and access of cost-effective pharmacy benefits. As a leader in SaaS technology for healthcare, the company is an Equal Opportunity and Affirmative Action employer committed to diversity and collaboration.
Own and optimize CI/CD pipelines, Kubernetes deployment, and infrastructure for model serving and inference.
Build telemetry, observability, and alerting to catch real problems and reduce noise.
Eliminate toil through thoughtful automation and improve developer and agent productivity.
Obvious is building an AI-native workspace that serves as an operating system for work, putting co-intelligence at the center. They are a small, talent-dense team with founders and leaders from top tech companies.
Design and implement monitoring and alerting systems using tools like Prometheus, Grafana, and DataDog to ensure high availability and reliability.
Optimize performance and reliability of healthcare payment applications, lead incident response, and develop SLOs/SLIs.
Automate CI/CD pipelines, infrastructure provisioning with Terraform, and manage cloud infrastructure on AWS with Kubernetes.
LMI is a digital solutions provider accelerating government impact with innovation and speed, bringing commercial-grade platforms and mission-ready AI to federal agencies. Headquartered in Tysons, Virginia, LMI serves the defense, space, healthcare, and energy sectors, focusing on agility and collaboration to drive impactful results.
Own features end-to-end across the stack, from backend Go services to frontend TypeScript/React.
Design and evolve systems for scheduling, probe lifecycle, and telemetry ingestion at scale.
Build cross-product and AI-assisted workflows that integrate Synthetic Monitoring into the Grafana Cloud platform.
Grafana Labs is the company behind the open source observability cloud, helping organizations see and act on their data. With over 1,600 employees across 40+ countries, the company fosters an open, collaborative culture.
Lead the design and development of automated, resilient platform technologies including Observability, DevOps, and ITSM. - Manage a team of platform engineers, driving technical roadmaps and ensuring platform reliability and security. - Build and operate OpenTelemetry observability platforms using LGTM stack on Kubernetes.
Flexential is a data center and IT services company building next-gen observability platforms for 40+ data center facilities. They value diversity and offer a collaborative culture focused on innovation.
Design and implement enterprise monitoring and observability strategies using AI-driven automation.
Apply machine learning techniques to improve incident detection, prediction, and resolution.
Collaborate with IT teams and stakeholders to optimize event management and operational efficiency.
The partner company specializes in enterprise IT monitoring and observability, leveraging AI and automation. It operates with a global team and offers a fully remote, contract-based work environment.
Design and champion internal Quality Programs, driving engineering-wide process changes and reliability standards.
Build SDLC observability pipelines and enforce hard metrics to measure deployment health.
Architect Shift-Left Pipeline Gates and engineer synthetic data for load testing.
Counterpart Health is transforming healthcare by providing an AI-enabled primary care tool, Counterpart Assistant, to support physicians in diagnosing and managing chronic conditions. The company is a subsidiary of Clover Health and values diversity, with a supportive and empathetic engineering team.
Design and evolve scalable cloud infrastructure on Google Cloud Platform, focusing on reliability and automation.
Strengthen observability platform with metrics, logging, and tracing to improve incident response and reduce recovery time.
Champion reliability practices like SLOs, error budgets, and DORA metrics to drive operational excellence.
They operate at the intersection of geospatial intelligence and environmental technology. They are a growing organization with a collaborative, high-impact engineering culture.
Define and execute the technical strategy for Observability & Alerting, Platform Infrastructure, and operational excellence.
Lead the evolution of reliable, scalable, secure, and efficient cloud platforms and distributed systems.
Champion SLIs, SLOs, error budgets, capacity planning, and automation across the service lifecycle.
Filevine is a Legal AI company delivering a unified platform for legal work, powered by LOIS (Legal Operating Intelligence System). The company is rapidly growing, recognized by Deloitte and Inc. as one of the most innovative and fastest-growing technology companies.
Design, implement, and support automation, deployment, monitoring, and operational reliability across cloud and hosted application environments.
Apply AI to optimize platform operations, including CI/CD, observability, incident response, and infrastructure design.
Provide technical leadership, mentoring, and day-to-day guidance to other engineers, and lead complex projects from planning through implementation.
Granicus provides cloud-based digital solutions for government communications, website design, and records management. With over 5,500 government agency clients and 300 million subscribers, Granicus is a remote-first company that values transparency and inclusion.