Remote Software engineering Jobs · Observability

Job listings

  • Design, build, and operate backend services supporting our WebRTC and Voice SDK products.
  • Take ownership of backend systems including authentication, push notifications, and diagnostic logging.
  • Improve reliability, scalability, and observability of customer-facing services.

We are building the future of global connectivity through a private, multi-cloud IP network and intuitive APIs. We are a financially stable and profitable company that fosters continuous learning and growth.

$114,800–$150,000/yr
US 4w PTO

  • Own complex production support escalations and ticket triage, providing hands-on troubleshooting and resolution alongside reliability work.
  • Partner with Software Engineering to investigate complex production issues, identify root causes and reliability risks, and drive permanent solutions.
  • Build observability across products, services, and critical customer workflows using meaningful metrics, logs, traces, dashboards, and actionable alerts.

Bloomerang provides a powerful giving platform and support for nonprofits to raise more, recruit more, and retain more. The company fosters a mission-driven culture built on core values of Simplify, Care and Act, and is home to innovative and skilled individuals.

Global 5w PTO

  • Expand DevDawg and Spice, cloud development environments paired with automated review and root-cause tooling that can approve real pull requests.
  • Build the paved roads for our API layer with consistent patterns for REST, GraphQL, auth, and permission plumbing.
  • Own how we see production with observability, metrics, alerting, and the readiness bar before shipping.

Close builds communication-first, AI-powered sales software to help teams sell more, faster. We are a bootstrapped, profitable 120-person company with a 100% remote team focused on building a CRM that gets out of your way.

  • Provide technical leadership for reliability across a large-scale advertising technology ecosystem
  • Lead reliability initiatives across ad serving, auctions, targeting, reporting, and billing systems
  • Mentor engineers and influence technical decisions to improve system resilience and developer productivity

The company is a partner organization operating a large-scale advertising technology ecosystem. Its size and culture are not detailed, but the role emphasizes reliability and operational excellence in a high-traffic environment.

$145,000–$177,000/yr

  • Build platform capabilities that enable engineering teams to deliver reliable software safely and efficiently.
  • Lead complex technical initiatives spanning cloud infrastructure, Kubernetes, observability, automation, and networking.
  • Design and implement solutions that improve availability, scalability, performance, and resilience of the platform.

Everbridge empowers enterprises and government organizations to anticipate, mitigate, respond to, and recover from critical events. The company focuses on building resilient systems and fostering a culture of ownership, continuous improvement, and operational excellence.

  • Set the technical direction and roadmap for Platform engineering, owning the reliability, performance, and cost of Traild's GCP infrastructure.
  • Improve developer experience by enhancing CI/CD pipelines, build tooling, observability, and reducing friction for other engineers.
  • Lead a small team of platform engineers, staying hands-on while growing and supporting your team members.

Traild is a high-growth SaaS company redefining how finance teams operate by combining AI, automation, and payments infrastructure for B2B finance. They are a rapidly growing global team with a strong culture, as evidenced by an eNPS score of 78.

Europe 6w PTO

  • Build and deliver AI-powered features that help users manage large datasets and enhance analytics-focused AI agents.
  • Rapidly prototype, test, and iterate with real users, shipping LLM- or agent-driven workflows for data engineering.
  • Collaborate across teams to integrate AI components with internal tools, taking full ownership of scalable solutions.

Grafana Labs is the company behind Grafana Cloud, a fully managed observability platform trusted by over 10,000 organizations. It is a 100% remote company with team members across 40+ countries, backed by leading investors and known for its open-source legacy and collaborative culture.

$163,000–$263,670/yr

  • Lead and develop a team of engineers, providing coaching, feedback, and career growth opportunities.
  • Partner with PM and Design to scope, estimate, and sequence roadmap items, making scope tradeoffs in real time.
  • Own the team's delivery of features, ensuring completion within timelines and quality requirements.

LaunchDarkly is a software platform that helps developers innovate on new features faster and manage the software development lifecycle. The company is growing fast and values a humble, open, collaborative, respectful, and kind team culture.

US Unlimited PTO

  • Build production-like test environments and tooling for engineers to spin up test cells resembling production topology.
  • Enable production-representative workloads including stress/load testing and traffic replay for controlled environments.
  • Drive system-level test coverage and build shared frameworks to make failure-mode testing repeatable across the engineering org.

Temporal is an open source programming model that simplifies code and makes applications more reliable. The company values curiosity, drive, collaboration, genuineness, and humility, and is building a team to be the reliable foundation of every developer's toolbox.

  • Build modular, plug-and-play AI agents that integrate into a broader agentic architecture.
  • Design and implement memory capabilities, including short-term and long-term memory, summarization, and retrieval-backed context.
  • Own quality through testing, evaluations, monitoring, and observability.

They build AI-powered applications to transform proposal workflows in the architecture, engineering, and construction industry. They operate in a fast-moving, remote environment with a focus on experimentation and continuous learning.