Source Job

Global 4w PTO

  • Own and improve service reliability for the product team: design for HA/performance/scale, define SLIs/SLOs
  • Align with org standards, implement DevOps-driven updates: support processes, templates, services, breaking changes, security fixes
  • Build and evolve GitLab CI/CD for build, test, security scans, and progressive delivery; speed up and harden pipelines

AWS Kubernetes GitLab CI/CD Prometheus Python

20 jobs similar to SRE [Cards&Accounts]

Jobs ranked by similarity.

Global

  • Ensure availability, performance, scalability, and resilience of production services in AWS.
  • Automate infrastructure provisioning and management using Infrastructure as Code (IaC).
  • Collaborate with development, architecture, security, and product teams to promote reliability best practices.

Experian is a global data and technology company that drives opportunities for people and businesses worldwide, operating in markets such as financial services, healthcare, automotive, and insurance. The company has over 25,200 employees across 32 countries and is recognized as a Top 25 global workplace by Fortune.

US Unlimited PTO

  • Keep user-facing services and production systems reliable, scalable, and efficient with automation and infrastructure-as-code.
  • Operate and troubleshoot production systems on Kubernetes, and contribute to observability with metrics, logs, and SLOs.
  • Participate in on-call, incident response, and post-incident reviews to drive improvements in automation and processes.

GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity and improve operational efficiency. With more than 50 million registered users and over 50% of the Fortune 100 as customers, GitLab fosters a high-performance, all-remote culture driven by values and continuous knowledge exchange.

Europe 6w PTO

  • Work with a team of DevOps and DBA professionals to improve infrastructure and streamline deployments across countries.
  • Continuously improve Kubernetes platform stability, efficiency, and GitOps-first environment provisioning.
  • Monitor cloud infrastructure, own on-call operations, and define SLIs/SLOs for reliability improvements.

Sporty Group is a remote-first company focused on sustainability in the sports and gaming industry. They maintain a competitive, performance-driven culture with a distributed team across EMEA.

$118,000–$151,000/yr
US 4w PTO

  • Build and improve platform services, including CI/CD pipelines and cloud infrastructure.
  • Collaborate with senior engineers to design scalable solutions and enhance developer experience.
  • Participate in incident response and retrospectives to drive continuous improvement.

Octopus Energy is a tech-powered energy company focused on renewable energy and customer experience. The company culture emphasizes ownership, collaboration, and making a tangible impact across teams.

Latin America

  • Build and operate the self-service infrastructure platform where developers and agents can validate changes in minutes.
  • Build golden paths for CI/CD, GitOps, and IaC to enable self-service provisioning and shipping.
  • Own reliability and observability, carrying on-call and turning recurring toil into automation.

Luxury Presence is building the AI growth platform for real estate. Backed by Bessemer Venture Partners, the company is a Series C firm with over 90,000 real estate professionals and has been ranked on the Inc. 5000 fastest-growing companies list three years in a row.

US Unlimited PTO

  • Own and evolve our SLI/SLO and error-budget frameworks, using them to influence prioritization and product decisions.
  • Lead incident response, drive postmortems, and turn findings into systemic fixes rather than one-off patches.
  • Build and maintain observability across metrics, logs, and traces (Datadog), improving signal and reducing alert fatigue.

MyFitnessPal provides tools, resources and support to enable users to reach their health goals. The company values collaboration, mentorship, and inclusive environments, with a focus on reliability and delivery.

$165,000–$216,000/yr
US Unlimited PTO

  • Develop internal tools and automate infrastructure using AWS, Kubernetes, and programming languages.
  • Research and design solutions to increase website robustness, availability, and cost efficiency.
  • Collaborate on documentation, code reviews, and rollout of new processes.

Angi powers the future of the home services industry, connecting homeowners with skilled pros. With 9 brands in 8 countries and employees worldwide, Angi has helped homeowners with over 300 million home projects.

Brazil

  • Ensure system architecture meets technical requirements by collaborating with IT teams (Architecture, Security, Infrastructure).
  • Maintain and evolve the microservices environment on AWS with a focus on information security.
  • Implement DevOps practices, automation, and monitoring tools to ensure system reliability and scalability.

Experian is a global data and technology company that powers opportunities for people and businesses worldwide. With 25,200 employees across 32 countries, it fosters a people-centric, inclusive culture recognized as a World's Best Workplace.

Canada Europe

  • Developing standards, infrastructure, and automation for dashboards, alerts, and monitors as code.
  • Partnering with development teams to establish production readiness and operational readiness.
  • Building tooling to automate observability and operational workflows, eliminating manual toil.

Glia is the #1 Banking AI platform, empowering community and regional financial institutions with an AI workforce. The company is trusted by over 700 banks and credit unions and has a remote-first culture with offices in Estonia.

Brazil

  • Act as a technical reference for SRE, DevOps, and cloud infrastructure, analyzing cloud environments in GCP/AWS for improvements and cost optimization.
  • Implement FinOps strategies, manage CI/CD pipelines, and maintain infrastructure as code using Terraform and Kubernetes.
  • Provide consultative support and communicate technical recommendations to engineering and business stakeholders.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. The company uses technology to streamline the application process and promote fair evaluation.

$122,000–$129,200/yr
US Unlimited PTO 16w maternity 4w paternity

  • Build, operate, and scale AWS infrastructure supporting a microservices platform.
  • Develop and maintain CI/CD pipelines to make shipping fast, reliable, and safe.
  • Partner with product engineering teams, translating platform concepts into practical guidance.

Storyblocks is a content company that provides video, audio, images, and creative tools through a subscription model, empowering storytellers. They are a remote-first company with a culture driven by data, empowered communication, and ownership, recognized as a top workplace.

Europe Middle East Asia North America

  • Build and operate frameworks to ensure reliable, sustainable solution delivery across Mistral-hosted and customer-hosted environments.
  • Operate Tier-1 customer environments, ensure SLO compliance, manage on-call and incident response.
  • Productize deployment, security, and scaling of Applied AI solutions with automation and security guardrails.

Mistral provides full-stack AI solutions from frontier models to developer tools, applications, and compute, partnering with enterprises across high-stakes industries. It is a dynamic, collaborative team with a diverse workforce distributed globally, known for being creative, low-ego, and team-spirited.

US

  • Design, build, and scale reliable infrastructure for Klover's fintech platform using modern technologies like Kubernetes, Terraform, and Istio.
  • Use AI agents as force multipliers to automate manual processes and improve developer experience.
  • Collaborate with engineering teams to ensure system reliability, performance, and security across production systems.

Attain powers Klover, a fast-growing fintech platform serving over one million active users monthly, processing over $1.5 billion annually. The company emphasizes collaboration, reliability, and innovation, with a culture of automation and AI-driven development.

Global

  • Own and evolve Webshare's production infrastructure by leading migration from Docker Swarm to Kubernetes and maintaining high availability across hundreds of servers and ~50 services.
  • Drive observability, establish IaC practices, CI/CD pipeline reliability, and participate in on-call rotation alongside backend developers.
  • Contribute platform tooling to improve developer experience and reduce infrastructure toil, ensuring no silos and shared infrastructure ownership.

We develop cutting-edge proxy and web data scraping solutions for thousands of the world's best known businesses, including Fortune 500 companies. We are a team of 500+ professionals with a culture focused on growth, learning, and shared infrastructure ownership.

$160,000–$180,000/yr
US

  • Own the infrastructure and platform powering the marketplace, focusing on reliability, observability, security, and automation.
  • Manage production AWS and EKS clusters, infrastructure as code with Terraform and GitOps, and CI/CD pipelines via GitHub Actions.
  • Build automation and internal tooling in Python, Bash, Go, and Node.js/TypeScript, and operate PostgreSQL, MongoDB, and Temporal.

Office Hours is an on-demand expert network that connects leading organizations with trusted experts across various knowledge domains. The company is hyper-growth, profitable, and expanding quickly, backed by top marketplace investors.

Poland

  • Design, write and deliver software to implement and support large web-scale, highly-performant, highly-available infrastructure on GCP/AWS.
  • Monitor infrastructure, respond to incidents, correct and improve systems to prevent incidents, and plan capacity.
  • Tune large-scale clusters for optimal performance and efficiency and support system deployments and product releases.

OpenX develops digital advertising marketplaces and technologies to optimize ad delivery for publishers and advertisers. The company operates a large-scale cloud infrastructure in Poland and values teamwork, customer centricity, and continuous learning.

US

  • Improve deployment reliability, reduce operational risk, and modernize AWS infrastructure toward Kubernetes.
  • Manage large-scale, high-availability distributed systems on AWS with Terraform and CI/CD pipelines.
  • Implement robust monitoring and observability, practicing SRE and security compliance best practices.

Peek is the operating system powering the experiences industry, helping museums, attractions, and tours increase revenues and deliver seamless guest experiences. Recognized by Forbes as a Best Startup Employer and by Built In as a Best Place to Work, we are a global remote-first team of Peeksters who obsess over customers and collaborate with purpose.

$152,000–$195,000/yr
US Unlimited PTO

  • Design, build, and scale Kubernetes infrastructure for secure, multi-tenant, high-availability applications.
  • Build and operate AI tooling infrastructure, including MCP servers and secure AI access.
  • Optimize CI/CD pipelines, implement progressive delivery, and advance Infrastructure as Code.

SecurityScorecard is the global leader in cybersecurity ratings, rating over 12 million companies across 64 countries. Headquartered in New York, it is recognized as a best workplace and funded by top investors.

$120,000–$120,000/yr
Canada Unlimited PTO

  • Design and maintain scalable, reliable infrastructure using AWS services like ECS and RDS
  • Build and manage infrastructure as code with Terraform, and optimize CI/CD pipelines for secure software delivery
  • Improve system observability using Datadog, Sentry, and Coralogix, and enable a seamless developer experience

Fullscript provides a platform for healthcare practitioners to prescribe and manage supplements and lab tests for patients. They serve over 125,000 practitioners and 10 million patients across North America, with a mission to help people get better.

Global

  • Embed with product and platform teams from early stages to ensure reliability is designed in from the start.
  • Define production-readiness standards and measurable SLIs/SLOs to guide operational excellence.
  • Build tooling and infrastructure across AWS, GCP, and Azure using Terraform, and share on-call rotation.

We build WebContainers and Bolt.new, an AI-powered app builder that lets you create, edit, and deploy full-stack apps instantly in your browser. We are a fully remote, globally distributed team of passionate engineers serving over 1 million developers monthly.