Source Job

Global

  • Plan and execute Datadog organization/account migration activities, including monitors, dashboards, and log configurations.
  • Leverage Datadog API, CLI, and Terraform to automate migration and validate monitoring configurations.
  • Collaborate with engineering teams to ensure migration accuracy and operational readiness.

Datadog Terraform AWS Kubernetes CI/CD

20 jobs similar to Senior Observability / DevOps Engineer

Jobs ranked by similarity.

Argentina

  • Architect and maintain critical cloud platform components on AWS EKS with high availability and automated resilience.
  • Establish SRE standards including SLO/SLI tracking, error budget frameworks, and automated operational tooling.
  • Design and implement OpenTelemetry capture pipelines for telemetry data feeding downstream platforms.

Inflect is a US-based advisory and marketplace that revolutionizes how companies buy and sell digital infrastructure services. They operate with a focus on high-impact consulting and autonomous work.

Spain

  • Design and optimize AWS cloud infrastructures, managing Kubernetes clusters and implementing Infrastructure as Code with Terraform.
  • Administer observability stacks including Elasticsearch/OpenSearch, VictoriaMetrics, Grafana, Loki, and Vector ingestion pipelines.
  • Automate CI/CD pipelines using Python and Groovy, and resolve complex incidents through Jira and ServiceNow.

Devoteam is a European consulting firm specializing in digital strategy, technology platforms, cybersecurity, and business transformation. With over 10,000 employees across 20 countries in Europe, the Middle East, and Africa, the company combines enterprise-grade technology with a close-knit, professional team culture.

Global

  • Support and optimize AWS infrastructure for Amazon Connect and enterprise cloud platforms.
  • Manage Infrastructure as Code using Terraform and automate tasks with Python and AWS CLI.
  • Develop dashboards, alerts, and monitoring solutions using CloudWatch, Dynatrace, Splunk, Grafana, and OpenTelemetry.

Miratech is a global IT services and consulting company that helps visionaries change the world by bringing together enterprise and start-up innovation. The company retains nearly 1000 full-time professionals, operates in over 25 countries, and has a culture of Relentless Performance with a 99% project success rate since 1989.

$104,000–$166,000/yr
US

  • Build and maintain telemetry pipelines across Dynatrace, Datadog, and Splunk for metrics, logs, and traces.
  • Design observability for distributed systems in AWS/GovCloud, including dashboards and golden-signal monitoring.
  • Build and tune alert definitions, support on-call rotations, and integrate with ServiceNow.

Peraton is a next-generation national security company that drives missions of consequence globally. They are a leading mission capability integrator and enterprise IT provider, serving essential government agencies and supporting the U.S. armed forces.

$6,500–$9,000/mo
Brazil

  • Architect, deliver, and maintain critical cloud platform components on AWS EKS, focusing on production reliability and observability.
  • Establish SRE standards including SLOs, error budgets, and automated tooling to reduce operational friction.
  • Provide technical advisory through code reviews and architecture recommendations to maintain high platform standards.

Inflect is a US-based advisory and marketplace revolutionizing digital infrastructure procurement. They are a small team focused on reducing friction in buying datacenter, cloud, and network services through automation and better deal terms.

US

  • Act as a subject matter expert for service management, including incident response, on-call, and operational automation.
  • Create compelling content across mediums like demos, blogs, and talks to build Datadog's reputation in DevOps and observability.
  • Partner with product engineering teams to build demos and coach internal teams on effective communication and presentation.

Datadog is the leading observability and security platform for the AI era, providing businesses with unified visibility across the technology stack to manage complexity at scale. Trusted by Fortune 500 companies and high-growth AI leaders worldwide, Datadog fosters an inclusive culture that values people from all walks of life.

US

  • Design, build, and maintain robust, scalable, and secure infrastructure systems supporting Laurel's AI-driven platform.
  • Manage and optimize cloud infrastructure (AWS and Azure), Kubernetes orchestration, and CI/CD pipelines to increase deployment frequency and reliability.
  • Implement comprehensive observability, monitoring, and alerting to maintain system health and partner with engineering teams to optimize performance and cost-efficiency.

Laurel is an AI Time platform for professional services firms, automating work time capture and connecting time data to business outcomes for clients like EY and Crowell & Moring. The company comprises top AI, product, and engineering talent, is VC-backed by Google Ventures and IVP, and fosters an inclusive, ambitious culture.

$150,000–$165,000/yr
US Unlimited PTO

  • Design and manage high-availability platforms using Kubernetes, Terraform, and Ansible with native-AI capabilities.
  • Develop and operate the observability stack: Grafana, Mimir, Loki, Tempo, and Prometheus on Kubernetes via GitLab CI/CD.
  • Build automation scripts in Python, maintain GitOps pipelines, and mentor mid-level engineers.

Flexential builds and operates critical IT platforms including observability, DevOps, and ITSM technologies. The company fosters a collaborative engineering culture and values diversity.

US

  • Define and evolve enterprise observability vision, standards, and roadmap.
  • Lead implementation of Dynatrace SaaS and platform capabilities.
  • Design observability solutions for cloud-native infrastructure.

FreedomPay provides commerce solutions. They are a large company with a culture focused on innovation and collaboration.

$152,800–$259,200/yr
Canada United States Unlimited PTO

  • Partner with customers to plan and deliver migrations from self-managed GitLab to GitLab Dedicated, using GitLab Geo for data replication with minimal disruption.
  • Balance hands-on technical delivery with clear communication and project ownership, translating customer goals into practical plans for cutover and validation.
  • Improve tooling, scripts, and runbooks to make migrations repeatable and reduce manual work, sharing insights with Product and Engineering teams.

GitLab is an intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity and improve operational efficiency. More than 50 million registered users and over 50% of the Fortune 100 trust GitLab to ship better software, and the company fosters a high-performance culture driven by values and continuous knowledge exchange.

India

  • Maintain observability platform and introduce observability on new projects.
  • Implement automated management features and configure solutions per security processes.
  • Manage CI systems and pipelines, and design and implement infrastructure.

Lingaro is a global technology company providing data, cloud, and DevOps solutions. With over 1,500 employees across 7 sites, they foster a diverse and inclusive culture.

$74,000–$111,000/yr
Canada Unlimited PTO

  • Define and implement observability strategies, standards, and governance across applications and platforms.
  • Design and maintain monitoring, alerting, dashboarding, and reporting solutions using Dynatrace or equivalent observability platforms.
  • Establish and drive SRE best practices including Service Level Indicators (SLIs), Service Level Objectives (SLOs), error budgets, and symptom-based alerting.

Valtech is the experience innovation company that helps brands unlock new value in an increasingly digital world by blending crafts, categories, and cultures. They have a workplace culture that fosters creativity, diversity, and autonomy, with a borderless global framework enabling seamless collaboration.

Hungary

  • Build and maintain scalable, reliable, and secure environments on AWS using Infrastructure as Code tools.
  • Design and manage CI/CD pipelines, oversee Kubernetes clusters, and ensure GitOps practices.
  • Monitor system health with OpenTelemetry and Grafana, enforce security best practices, and mentor junior engineers.

Deutsche Telekom IT Solutions is a subsidiary of the Deutsche Telekom Group, providing IT and telecommunications services with over 5,300 employees. Recognized as Hungary's most attractive employer, it serves large corporate clients across Europe.

Latin America

  • Design, build, and maintain reliable cloud infrastructure on AWS using CI/CD pipelines and IaC tools.
  • Automate containerized workloads with Docker, implement monitoring and observability solutions, and ensure security best practices.
  • Collaborate with engineering teams to improve system reliability, troubleshoot issues, and drive operational excellence.

GoFasti is a Talent-as-a-Service company that bridges world-class developers and designers from Latin America with first-class companies globally. They are a remote-first organization focused on matching top talent with international opportunities.

$150,000–$250,000/yr
US Europe Singapore

  • Own production uptime, latency, provisioning speed, infrastructure cost, and incident response for core platform services.
  • Build and maintain AWS infrastructure using Terraform, Kubernetes/EKS, Helm, Docker, EC2, CodeBuild, ECR, S3, IAM, networking, and secrets management.
  • Design and improve backend and platform systems for scale — capacity planning, autoscaling, queueing, backpressure, cleanup jobs, retries, and rollback paths.

A fast-growing AI/ML platform startup building infrastructure for training, evaluating, and aligning AI models within reinforcement learning environments. The engineering team of ~15 includes competitive programming medalists, serial AI startup founders, and researchers published at top venues.

US UK Ireland Poland Germany Australia

  • Design and implement a scalable observability platform for Whatnot's growing infrastructure.
  • Work with core infrastructure, platform, and developer tools teams to redesign data collection to visualization.
  • Utilize AI agents and open standards to ensure visibility into software stack performance and reliability.

Whatnot is the largest live shopping platform in North America and Europe, enabling sellers to build businesses across hundreds of categories. They are a remote co-located team anchored in hubs across the US, UK, Ireland, Poland, Germany, and Australia, and were recently named the #1 Best Startup Employer in America by Forbes.

$119,380–$165,100/yr
Spain UK

  • Lead the design of scalable, fault-tolerant, self-healing systems in a multi-region AWS environment.
  • Define SLOs and SLIs to drive architectural decisions and error budget policies.
  • Conduct blameless post-incident reviews and implement long-term preventive measures.

Airalo is the world's first eSIM store, helping travelers access affordable mobile data in 200+ countries. They are a fully remote team of 400+ people across 60+ countries, with a culture of trust, ownership, and freedom.

US

  • Improve system availability, scalability, and resilience across Flowcode's platforms.
  • Manage and scale core AWS infrastructure through Infrastructure as Code (Terraform) and enhance disaster recovery.
  • Oversee monitoring, logging, and alerting infrastructure, and develop high-signal metrics and dashboards.

Flowcode is a technology company specializing in QR code and smart link solutions for offline-to-online engagement. The company is a growth-stage startup seeking high-performing individuals who thrive in a fast-paced, demanding environment.

US

  • Provide day-to-day technical support to internal software engineering teams on cloud infrastructure and platform-related issues.
  • Troubleshoot deployment, networking, infrastructure, and CI/CD challenges while performing root cause analysis.
  • Identify opportunities to automate manual tasks and improve engineering workflows using Terraform, Python, and Bash.

Software Mind is a team of engineers that provides project ramp-up services for top companies. They are a multicultural, growing company with an excellent work environment certified by Great Place To Work.

US

  • Design and build resilient AWS and Kubernetes platforms to improve reliability, scalability, and security.
  • Define SLOs, build observability, automate operational work, and lead incident response and post-incident reviews.
  • Partner with engineering, platform, security, and QA teams to establish reliability standards and optimize cost.

Electric Power Engineers (EPE) provides consulting expertise and energy intelligence software solutions for power and energy clients, focusing on renewable energy and grid modernization. With over half a century in the industry, the company fosters innovation and collaboration, working with industry leaders to build a secure and resilient grid.