Design, build, and scale reliable infrastructure for Klover's fintech platform using modern technologies like Kubernetes, Terraform, and Istio.
Use AI agents as force multipliers to automate manual processes and improve developer experience.
Collaborate with engineering teams to ensure system reliability, performance, and security across production systems.
Attain powers Klover, a fast-growing fintech platform serving over one million active users monthly, processing over $1.5 billion annually. The company emphasizes collaboration, reliability, and innovation, with a culture of automation and AI-driven development.
Own and evolve observability strategy including monitoring, alerting, dashboards, logging, and distributed tracing.
Define and manage SLIs, SLOs, and reliability metrics, improving MTTD and MTTR through automation.
Build and maintain reliable cloud infrastructure on AWS and Kubernetes while mentoring engineers on SRE best practices.
Filevine is a Legal AI company delivering Legal Operating Intelligence for legal work. Fueled by a team of exceptional collaborators and innovators, Filevine’s rapid growth has earned AI awards and recognition from Deloitte and Inc. as one of the most innovative and fastest-growing technology companies in the country.
Define and implement SLIs/SLOs for critical services, lead incident response, and conduct blameless postmortems to drive systemic improvements.
Design and improve monitoring and alerting with Prometheus and Grafana, build internal tooling, and automate operational workflows to reduce toil.
Partner with cross-functional engineering teams to improve system resilience, contribute to architectural discussions, and strengthen production readiness standards.
Runpod provides a cloud platform for AI development, used by over one million developers for training, fine-tuning, and deploying AI models. The company is a small, remote-first team that closed a $100M Series A in June 2026, emphasizing ownership, speed, and impact at scale.
Design and deploy a modern observability stack including logging, metrics, and distributed tracing across hundreds of services.
Build alerting policies and incident response workflows to reduce manual escalations and improve mean time to detect and resolve.
Automate toil and set SRE standards while mentoring engineers on observability tooling.
WeightWatchers is a global digital health company and the world's #1 doctor-recommended behavioral weight health program. They have a cloud architecture supporting hundreds of services and are expanding their platform engineering team.
Developing standards, infrastructure, and automation for dashboards, alerts, and monitors as code.
Partnering with development teams to establish production readiness and operational readiness.
Building tooling to automate observability and operational workflows, eliminating manual toil.
Glia is the #1 Banking AI platform, empowering community and regional financial institutions with an AI workforce. The company is trusted by over 700 banks and credit unions and has a remote-first culture with offices in Estonia.
Design, build, and run distributed cloud architectures and large-scale production systems.
Ensure reliability, observability, performance, and cost efficiency of the platform.
Collaborate with product and backend teams to design system architecture and optimize resource use.
Tinybird helps developers and data teams unlock the power of real-time data, enabling them to build data pipelines and innovative data products quickly. They are a remote-first company with a culture of ownership, transparency, and clear communication.
Design and build automated reliability and self-healing systems to protect production at scale.
Own and improve incident management tooling and on-call health, reducing alert noise and empowering teams.
Develop and evolve observability infrastructure, including monitoring, alerting, SLOs, and performance regression detection.
Samsara is the pioneer of the Connected Operations Cloud, enabling organizations to harness IoT data for improved safety, efficiency, and sustainability. As a public company with over 2.3 million IoT devices deployed, we foster a culture of growth mindset and collaboration.
Define and lead the end-to-end observability strategy covering logging, metrics, tracing, and alerting.
Architect and evolve a unified observability platform ensuring scalability and reliability.
Build and lead a high-performing observability engineering team with strong technical standards.
The company operates a high-scale developer-facing platform focused on reliability and performance. It is a remote-first organization with a globally distributed engineering team committed to building best-in-class developer infrastructure.
Design and maintain Grafana dashboards and telemetry visualizations to monitor system performance and platform health.
Develop and maintain modular Ansible playbooks to automate infrastructure provisioning and configuration.
Configure observability solutions with Prometheus monitoring and alerting, and participate in Agile ceremonies.
Miratech is a global IT services and consulting company that helps visionaries change the world by supporting digital transformation for large enterprises. With nearly 1,000 full-time professionals across 5 continents and 25 countries, the company has a culture of Relentless Performance with a 99% project success rate and over 25% annual growth.
Design and implement high-quality, scalable integrations for observability solutions.
Collaborate with cross-functional teams to deliver features aligned with product strategy.
Participate in on-call rotations and contribute to open-source communities.
Grafana Labs provides an open-source observability platform, Grafana Cloud, that integrates metrics, logs, and traces. With over 1,600 team members across 40+ countries, they maintain a remote-first, collaborative culture backed by leading investors.
Work with a team of DevOps and DBA professionals to improve infrastructure and streamline deployments across countries.
Continuously improve Kubernetes platform stability, efficiency, and GitOps-first environment provisioning.
Monitor cloud infrastructure, own on-call operations, and define SLIs/SLOs for reliability improvements.
Sporty Group is a remote-first company focused on sustainability in the sports and gaming industry. They maintain a competitive, performance-driven culture with a distributed team across EMEA.
Harden, simplify, and operationalize the production platform on Google Cloud Platform for enterprise customers.
Own and evolve core infrastructure, CI/CD pipelines, and infrastructure as code practices using Terraform or Pulumi.
Drive observability, developer productivity, and engineering culture to raise the bar across the team.
GC AI is the fastest-growing legal AI platform for in-house legal teams, building the future of legal work. With over 1,700 companies using the platform, including 150+ public companies and 25+ unicorns, the team has 10x'd revenue in 12 months and raised a $60 million Series B.
Ensure reliability, performance, and scalability of Backcountry's multi-cloud platform.
Drive incident resolution, postmortems, and automation to reduce operational toil.
Leverage AI-assisted engineering tools and collaborate with teams to build and maintain observability and SLI/SLO instrumentation.
Backcountry is an online retailer of outdoor gear and apparel, rooted in adventure and the outdoor lifestyle. The company fosters a culture of recognition, wellbeing, and connection, with a lean, fast-paced engineering team.
Lead client discovery, architecture workshops, and solution design across observability, telemetry, reliability, and operational intelligence initiatives.
Define scalable standards for telemetry onboarding, naming, tagging, RBAC, service ownership, dashboards, alert governance, runbooks, and operational handoff.
AHEAD builds platforms for digital business by weaving together cloud infrastructure, automation, analytics, and software delivery to help enterprises achieve digital transformation. The company prioritizes a culture of belonging and is an equal opportunity employer that values diversity and inclusion.
Design and evolve cloud architecture on GCP using Terraform and GitOps.
Build CI/CD pipelines for IaC with Policy-as-Code and drift detection.
Strengthen observability stack with Prometheus, Grafana, Loki, and Tempo.
Alpaca is a US-headquartered self-clearing broker-dealer and brokerage infrastructure for stocks, ETFs, options, crypto, fixed income, and 24/5 trading. With over $320 million in total investment and a diverse global team of 380+ members spanning multiple countries, we are committed to opening financial services to everyone on the planet.
Own and operate customer-facing managed infrastructure across multiple AWS accounts and regions.
Serve as the senior technical escalation point for production incidents and complex configurations.
Contribute to OpenTelemetry distributions and maintain open source projects like Refinery.
Honeycomb provides observability for developer tools, helping companies like HelloFresh and Slack understand their software. They have over 200 employees and were named to Forbes' Best Startups in 2022 and 2023, with a culture that values inclusion and autonomy.
Work with teams to define SLIs and SLOs, and create systems for observability.
Analyze failure scenarios, create runbooks, and reduce work that does not add value.
Participate in incident management and facilitate on-call duty to ensure reliable production environments.
Valtech is an experience innovation company that helps brands unlock new value in an increasingly digital world. They are a global team with a values-driven culture that fosters creativity, diversity, and autonomy.
Own and evolve our SLI/SLO and error-budget frameworks, using them to influence prioritization and product decisions.
Lead incident response, drive postmortems, and turn findings into systemic fixes rather than one-off patches.
Build and maintain observability across metrics, logs, and traces (Datadog), improving signal and reducing alert fatigue.
MyFitnessPal provides tools, resources and support to enable users to reach their health goals. The company values collaboration, mentorship, and inclusive environments, with a focus on reliability and delivery.
Build and maintain observability across the platform in Datadog, including dashboards, monitors, APM, and log pipelines.
Participate in on-call rotation and incident response, driving blameless post-incident reviews and automating toil.
Leverage AI tools to accelerate debugging, generate runbooks, and build automation for operational efficiency.
IPSY is a beauty subscription platform that connects brands and consumers through curated beauty products. It is a remote-first company with a focus on community and engagement.
Lead the design, development and operation of large-scale, secure observability systems to keep services online and performant.
Deploy and scale Prometheus, ElasticSearch clusters, and high-throughput Kafka data pipelines for millions of customer devices.
Collaborate with the Observability team to build alerting systems, APIs, and self-service monitoring tools using Terraform and multiple languages.
ItD is a new generation consulting and software development company that blends diversity, innovation, and integrity with real business results. It is a woman- and minority-led firm with a global community, empowering employees and offering benefits like medical, dental, vision, 401(k), and career development.