Similar Jobs

See all

Key Responsibilities:

  • Own the reliability strategy including SLOs, SLIs, and error budgets across engineering teams.
  • Drive architectural decisions for event-driven systems and asynchronous communication.
  • Manage AWS infrastructure with Infrastructure as Code and ensure Kubernetes scalability.

Qualifications:

  • Extensive experience in SRE, Platform Engineering, or DevOps with ownership of production systems.
  • Deep expertise in event-driven architectures and messaging systems like Kafka or RabbitMQ.
  • Strong AWS, Kubernetes, and observability skills with tools like Datadog.

Compensation & Benefits:

  • Competitive compensation and stock options.
  • Fully remote work with home office allowance and company equipment.
  • Flexible days off and access to professional development courses.

Partner Company

This partner company builds a globally scaled, AI-native platform with a focus on reliability and event-driven systems. They offer a collaborative international culture with significant technical ownership and continuous improvement.

Apply for This Position