Lead observability and monitoring operations integration and workflow support, including event-to-incident patterns and dashboard visualization.
Manage events and incidents tied to monitoring platforms, and support OpenTelemetry implementation and automation using Splunk SOAR/Ansible.
Provide operational reporting views and ensure hands-on experience with enterprise monitoring, observability, or APM engineering for technical and non-technical stakeholders.
20 jobs similar to SRE/Observability/Monitoring/Integration/Event & Incident/Visualization/Open Telemetry SME Support (Senior and Mid-Level Technology Consultant)
Lead technical strategy and solution design for enterprise monitoring and observability platforms, including Splunk, AppDynamics, DataDog, and ELK.
Define implementation patterns for telemetry ingestion, alerting, dashboarding, and ITSM integration with ServiceNow and CMDB.
Provide technical leadership for large-scale federal programs, including proposal and bid strategies, and present results to senior management.
Makpar is a comprehensive professional and technical solutions provider for the Federal government, combining expertise in cloud engineering, data management, cybersecurity and emerging technologies. They pride themselves on a connected and engaged workforce dedicated to delivering success for clients and the American people.
Implement monitoring use cases under senior direction and support data onboarding, alert configuration, and incident workflow alignment.
Participate in troubleshooting and operational support activities, and assist in configuring and tuning dashboards, alerts, and monitoring rules.
Collaborate with team members to validate monitoring coverage across supported systems and escalate unresolved issues through ITSM channels.
Makpar is a comprehensive professional and technical solutions provider for the Federal government, specializing in cloud engineering, data management, cybersecurity, and emerging technologies. They have a connected and engaged workforce dedicated to delivering success for clients and the American people.
Design, operate, and continuously tune platform monitoring across Dynatrace and Splunk.
Own the Single Pane of Glass dashboard and lead all P1/P2 incident responses.
Manage on-call rotation and deliver monthly SLA reports.
We are a growing Service-Disabled Veteran Owned company providing IT solutions to federal clients. We embrace remote work and invest in our employees' growth.
Design and architect observability solutions leveraging OpenTelemetry, Kubernetes, and cloud-native technologies.
Develop and execute Proofs of Concept (POCs) that highlight Dash0's differentiated technical capabilities.
Deliver engaging technical demos and presentations tailored to engineering and executive audiences.
Dash0 is building an OpenTelemetry-native observability platform that eliminates vendor lock-in and provides transparent pricing. Backed by top-tier investors including Balderton Capital, Accel and Cherry Ventures, the company has a collaborative, fast-moving team culture with a builder mindset.
Lead and execute on complex technical troubleshooting and incident resolution from investigation to delivery of permanent solutions.
Design and implement comprehensive monitoring processes, including creating detailed playbooks and runbooks for common scenarios, as well as detailed documentation including post-incident reviews and knowledge base articles.
Leverage AI-powered tools and workflows to automate issue detection, diagnosis, and resolution processes.
Alternative Payments is building the financial operating system for SMBs, consolidating the disconnected tech stack that holds service-based businesses back. We're growing fast, thinking big, and building a global team that wants to be part of something that lasts.
Serve as a trusted technical advisor guiding customers through their observability journey.
Design and guide customer observability maturity strategies to improve reliability and operational visibility.
Provide expert troubleshooting and technical recommendations to resolve complex challenges.
Jobgether is a platform that uses AI to match candidates with jobs. They focus on remote work and have a collaborative culture built around transparency, autonomy, and trust.
Maintain observability platform and introduce observability on new projects.
Implement automated management features and configure solutions per security processes.
Manage CI systems and pipelines, and design and implement infrastructure.
Lingaro is a global technology company providing data, cloud, and DevOps solutions. With over 1,500 employees across 7 sites, they foster a diverse and inclusive culture.
Design and implement enterprise monitoring and observability strategies using AI-driven automation.
Apply machine learning techniques to improve incident detection, prediction, and resolution.
Collaborate with IT teams and stakeholders to optimize event management and operational efficiency.
The partner company specializes in enterprise IT monitoring and observability, leveraging AI and automation. It operates with a global team and offers a fully remote, contract-based work environment.
Manage the ticket queue, prioritize and resolve requests, and identify recurring categories for automation.
Participate in rotating on-call and incident response, troubleshooting and documenting issues in real time.
Build and maintain monitoring dashboards (Tableau, Superset, Grafana) to track service health and data quality.
Airbnb is a global community marketplace that connects hosts with guests for unique stays and experiences. With over 5 million hosts and 2 billion guest arrivals, the company fosters a culture of inclusion and belonging, emphasizing innovation and engagement.
You will define the product vision, strategy, and roadmap for Elastic Agent, Fleet Server, and telemetry collectors.
You will drive data-driven decisions by defining and tracking KPIs for agent deployment success and pipeline efficiency.
You will lead cross-functional initiatives with engineering, UX, and marketing to deliver compelling collector experiences.
Elastic, the Search AI Company, enables everyone to find answers in real time using all their data at scale. Used by more than 50% of the Fortune 500, Elastic's cloud-based solutions for search, security, and observability help organizations deliver on the promise of AI.
Serve as the primary technical point of contact for a portfolio of Grafana customers, designing and guiding their observability maturity journey.
Conduct regular technical reviews, health checks, and root cause analysis to drive adoption and ensure customer success.
Act as the voice of the customer internally, shaping product feedback and roadmap priorities while building long-term strategic relationships.
Grafana Labs builds the open source observability platform Grafana and its fully managed cloud service. The company has over 1,600 team members across 40+ countries, serving more than 7,000 customers including major enterprises, and fosters a remote-first, transparent, and innovation-driven culture.
Architect and maintain critical cloud platform components on AWS EKS with high availability and automated resilience.
Establish SRE standards including SLO/SLI tracking, error budget frameworks, and automated operational tooling.
Design and implement OpenTelemetry capture pipelines for telemetry data feeding downstream platforms.
Inflect is a US-based advisory and marketplace that revolutionizes how companies buy and sell digital infrastructure services. They operate with a focus on high-impact consulting and autonomous work.
Lead the design and development of automated, resilient platform technologies including Observability, DevOps, and ITSM. - Manage a team of platform engineers, driving technical roadmaps and ensuring platform reliability and security. - Build and operate OpenTelemetry observability platforms using LGTM stack on Kubernetes.
Flexential is a data center and IT services company building next-gen observability platforms for 40+ data center facilities. They value diversity and offer a collaborative culture focused on innovation.
Define and execute the technical strategy for observability, platform infrastructure, and operational excellence.
Lead the design and evolution of scalable, secure, reliable cloud-native platforms and distributed systems.
Establish reliability best practices including SLIs, SLOs, error budgets, and automation initiatives.
The company is a technology organization that builds and operates large-scale cloud infrastructure. It fosters a collaborative culture centered on innovation, ownership, and impact.
Lead the design and operation of LivePerson's observability platforms across logs, metrics, traces, alerting, and synthetic monitoring.
Own large-scale observability pipelines using technologies like Elastic Cloud, Grafana, Prometheus, and Kafka.
Provide technical leadership and mentorship while driving best practices in DevOps, cloud engineering, and observability.
LivePerson is a leader in trusted enterprise conversational AI and digital transformation, powering nearly a billion conversational interactions every month. The company is recognized as the #1 Most Innovative AI Company by Fast Company and fosters a diverse, inclusive culture that empowers employees globally.
Lead technical and managerial direction for the SRE team, defining reliability, observability, and operational excellence strategy.
Coordinate critical incident responses and root cause analysis, collaborating with architecture, development, security, and product teams.
Drive automation, continuous improvement, and adoption of SRE, DevOps, and Platform Engineering best practices.
Experian is a global data and technology company that drives opportunities for people and businesses worldwide. With 25,200 employees in 32 countries, it has a people-centric, inclusive culture recognized by awards such as World's Best Workplaces™ 2025.
Design and architect enterprise observability solutions using cloud-native technologies and OpenTelemetry.
Deliver technical demonstrations, proof-of-concepts, and architecture sessions to showcase product value.
Support enterprise sales cycles by partnering with account teams and advising on best practices.
The company specializes in modern observability solutions for cloud-native environments. It is a high-growth technology company with a collaborative, fast-paced culture emphasizing ownership and impact.
You define architecture and operational standards for managed services across AWS. - You serve as the final technical escalation point for complex customer situations. - You shape Honeycomb's open source strategy in the OpenTelemetry ecosystem and mentor engineers.
Honeycomb is a service for observability, defining developer tools for the near and present future. We are a fully distributed company of over 200 talented and inclusive bees, named to Forbes' America's Best Startups of 2022 and 2023.