Own the day-to-day operation of a monitoring platform, including dashboards and alerts.
Proactively analyze logs, traces, and metrics to detect failures and drive resolution.
Define and measure SLIs, improve reliability, and automate operational tasks.
CI&T helps large enterprises transform AI potential into real business impact with AI deployment and tech-integrated solutions. With 30 years of experience and 8,000 employees across 25 countries, we collaborate to build solutions with real impact.
Lead a two-month observability maturity assessment across metrics, logs, and traces at massive scale.
Drive consolidation to AWS-native observability on OpenTelemetry, including pipelines, dashboards, and alerts.
Act as Pod Leader: set technical direction, stay hands-on, and own the customer relationship.
EverOps is the premier Embedded Service Provider, partnering directly with customer engineering teams to assess and address mission-critical infrastructure, cloud, and delivery challenges. We've been remote since day one, and our culture centers on ownership, technical leadership, and continuous professional growth.
Architect and build the observability platform for metrics, logs, traces, and events across global infrastructure.
Drive instrumentation with OpenTelemetry, building shared libraries and collector deployments for correlated signals.
Run observability as an internal product with published interfaces, versioned clients, and SLOs to ensure adoption.
Smartsheet empowers teams to manage work and scale solutions, uniting human teams with AI agents to automate tasks and uncover insights. With over 20 years of experience, the company fosters a collaborative, innovative culture focused on employee well-being and professional growth.
Act as a subject matter expert for service management tools and practices, creating content across various mediums.
Partner with product engineering teams to build demos and coach internal teams on communication.
Interface with open source communities and contribute to the product through feedback, documentation, or code.
Datadog is the leading observability and security platform for the AI era, providing unified visibility across the technology stack. Trusted by Fortune 500 companies and high-growth AI leaders, Datadog fosters a culture of innovation and collaboration.
Plan and execute Datadog organization/account migration activities, including monitors, dashboards, and log configurations.
Leverage Datadog API, CLI, and Terraform to automate migration and validate monitoring configurations.
Collaborate with engineering teams to ensure migration accuracy and operational readiness.
Miratech is a global IT services and consulting company that helps visionaries change the world through digital transformation. The company retains nearly 1,000 full-time professionals with a culture of Relentless Performance and over 99% project success rate since 1989.
Manage and develop a distributed team of backend and frontend engineers, providing regular feedback and supporting career growth.
Collaborate closely with go-to-market and engineering leadership to ensure seamless integration of Fleet Management, Kubernetes Helm Chart, and Instrumentation Hub.
Foster a psychologically safe environment that encourages learning, experimentation, and continuous improvement.
Grafana Labs is the company behind Grafana Cloud, a fully managed observability platform trusted by over 10,000 organizations for reliability and telemetry optimization. As a 100% remote company with team members across 40+ countries, Grafana Labs fosters a global, collaborative culture that values transparency, autonomy, and meaningful work.
Operate and improve monitoring infrastructure using Zabbix, Prometheus, Grafana, and InfluxDB.
Design alerting, escalation, and network documentation for customer and internal projects.
Share knowledge and participate in a rotating shift schedule.
The hiring partner is a technology-driven team dedicated to building and operating reliable, secure, and high-performance IT environments. The culture emphasizes collaboration, continuous learning, and professional development.
Lead a globally distributed observability team building and operating metrics, logging, alerting, and capacity planning platforms.
Set priorities with Site Reliability Engineering, Product Engineering, and GitLab Dedicated teams while owning reliability, scalability, and cost.
Participate in incident response and on-call rotations, using SLOs, error budgets, and AI tools to improve alerting and sustain operational load.
GitLab is an intelligent orchestration platform for DevSecOps that helps organizations increase developer productivity, improve operational efficiency, and accelerate digital transformation. With more than 50 million registered users and a high-performance culture driven by shared values and AI adoption, GitLab's globally distributed team collaborates to solve complex problems.
Validate and account for energy flows at operational boundaries with other utility providers.
Monitor and verify energy measurements to ensure accuracy and reliability.
Investigate discrepancies and collaborate with teams to maintain smooth operations.
A company operating in Brazil's energy distribution sector, focusing on accurate energy measurement and operational reliability. The environment is structured, results-oriented, and values analytical, detail-oriented professionals.
Own the full incident lifecycle from first alert to root-cause analysis and permanent fix across data pipelines and cloud/on-prem environments.
Support development and business analysis teams while configuring and maintaining monitoring systems.
Opportunity to grow into a DevOps role within a dynamic team.
We are Kyivstar, a Ukrainian telecommunications company providing mobile and data services. We are building new revenue streams through our Big Data platform, with a focus on reliability and innovation.
Monitor and diagnose infrastructure issues, responding to alerts independently or with DevOps support.
Create and improve monitoring rules and perform basic DevOps tasks.
Document all processes and collaborate with global teams.
airSlate is a global SaaS technology company developing no-code workflow automation, electronic signature, and document management solutions. With over hundreds of millions of users and more than one million customers worldwide, the company operates in over 20 countries with a team of diverse professionals.
Design, build, and operate high-scale observability pipelines for logs, metrics, traces, and exceptions.
Lead cross-functional initiatives to resolve scaling bottlenecks and evolve production infrastructure safely.
Partner with engineering teams to improve observability tools and provide technical leadership across teams.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. It emphasizes learning, professional growth, and an inclusive work environment.
Monitor health of products and infrastructure, handling 150-250 alerts per shift.
Analyze alerts using Grafana, Zabbix, and internal docs; fix issues or escalate.
Support senior admins and cover primary duty shifts after onboarding.
Social Discovery Group (SDG) builds social entertainment platforms that connect people online across cultures and regions, addressing loneliness and isolation. The company has an international remote team of digital nomads and is a two-time 'Great Place to Work' winner (USA & Japan, 2024-2025).
Build a strong understanding of Clutch’s platform, architecture, and support workflows to identify systemic friction and recurring issues.
Independently own complex investigations and executive-level escalations, driving them end to end with clear outcomes.
Develop reusable solutions like documentation, scripts, and tooling to reduce repeat issues and raise team effectiveness.
Clutch is a vertical SaaS company backed by Andreessen Horowitz that helps credit unions become FinTech lenders to provide affordable lending solutions. The Support Engineering team is a growing group of 13 experienced engineers, focused on delivering legendary support through collaboration, transparency, and follow-through.
Design and develop ServiceNow reports, dashboards, and Performance Analytics for IT visibility.
Automate ITSM reporting and KPI frameworks to deliver actionable insights.
Partner with stakeholders to define requirements and resolve data quality issues.
Danaher is a global life sciences, biotechnology, and diagnostics innovator helping solve the world's most important health challenges. With over 60,000 associates across more than 15 businesses, it fosters a culture of continuous improvement and belonging.
Monitor the health and performance of products and infrastructure by responding to high volumes of alerts.
Investigate issues using Grafana, Zabbix, and internal documentation, and resolve or escalate incidents.
Collaborate with a distributed international team and take ownership of primary duty shifts after onboarding.
The company provides products and infrastructure monitoring services in a remote-first operational environment. It has a distributed international team and values reliability, collaboration, and continuous learning.
Own and evolve the product roadmap for Observability and Compute/Network teams, from metrics and alerting to Kubernetes and capacity planning.
Partner with engineering and security leaders to prioritize reliability, cost, and developer experience.
Drive self-service adoption and track outcomes like SLO attainment, MTTD/MTTR, and infrastructure cost.
Addepar is a global data and AI platform that empowers investment professionals to turn complex financial information into actionable intelligence. More than 1,500 firms in 60 countries use Addepar to manage nearly $10 trillion in assets, with an inclusive, ownership-driven culture.
Investigate B2B incidents in Engagement-CRM, verifying user data and campaign parameters.
Diagnose via logs (Kibana/OpenSearch), SQL, REST APIs, and DevTools; escalate to engineering.
Guide L1 and operators on configurations, meet SLAs, and maintain runbooks.
GR8_TECH builds B2B iGaming platforms, delivering full-cycle tech, integrations, consulting, and operational support for operators. With 1000+ employees across the globe, the culture is built on trust, ownership, and a growth mindset.
Design and operate cloud infrastructure and observability systems for scalable SaaS and on-premise environments.
Build and maintain CI/CD pipelines, automation, and internal platform tooling using AWS, Kubernetes, Terraform, and GitHub Actions.
Enhance incident response, monitoring, alerting, and SLO management while collaborating with product and engineering teams.
Our partner is a B2B software company providing SaaS and on-premise solutions. The engineering culture emphasizes collaboration, reliability, and automation within a fast-paced environment.
Manage assigned technical projects, guiding teams on Agile/Scrum practices to delight clients.
Remove impediments and build a trusting environment for problem-solving without blame or retribution.
Plan and coordinate project activities, schedules, and budgets to ensure delivery on time and within scope.
AHEAD builds platforms for digital business by weaving together cloud infrastructure, automation, analytics, and software delivery. They prioritize creating a culture of belonging, are an equal opportunity employer, and value diverse perspectives.