Embed with product teams to improve operational maturity through on-call, monitoring, and alerting practices.
Run game day exercises and implement reliability techniques in Haskell & TypeScript code.
Champion reliability practices through design reviews and advocate for SLOs tied to customer outcomes.
Mercury is a fintech company that provides banking services for startups. They are building a modern banking platform and value reliability and innovation.
Design and manage high-availability platforms using Kubernetes, Terraform, and Ansible with native-AI capabilities.
Develop and operate the observability stack: Grafana, Mimir, Loki, Tempo, and Prometheus on Kubernetes via GitLab CI/CD.
Build automation scripts in Python, maintain GitOps pipelines, and mentor mid-level engineers.
Flexential builds and operates critical IT platforms including observability, DevOps, and ITSM technologies. The company fosters a collaborative engineering culture and values diversity.
Participate in a structured training track to develop expertise in AI benchmarking, profiling, and performance tuning.
Build and scale benchmarking infrastructure for evaluating AI systems in enterprise settings.
Design agent evaluation pipelines that measure reasoning, accuracy, alignment, and user outcomes.
DevRev is building Computer, an AI teammate that unifies data, tools, and workflows into a single AI-ready platform, giving employees real-time insights and proactive suggestions. Backed by Khosla Ventures and Mayfield with over $150M raised, the company is trusted by global companies across industries and fosters a culture of innovation and collaboration.
Drive technical strategy and roadmap for adaptive telemetry databases.
Lead end-to-end delivery of large, cross-functional projects.
Own architecture, reliability, performance, and cost for critical systems.
Grafana Labs is the company behind Grafana Cloud, the fully managed observability platform trusted by more than 10,000 organizations. We are a 100% remote company with team members across 40+ countries, backed by leading investors, and we foster a collaborative, open-source culture.
Work collaboratively with a team to create and maintain the foundational platform for Reddit's infrastructure.
Deliver software to improve the availability, scalability, latency, and efficiency of observability components.
Contribute upstream changes to open source projects and share on-call responsibilities.
Reddit is a community of communities, built on shared interests, passion, and trust. With 100,000+ active communities and approximately 130 million daily active unique visitors, Reddit is one of the internet’s largest sources of information, employing a flexible-first workforce that values open-source contributions.
Own Kubernetes deployment and operational health, including scaling, rollout/rollback, and resource tuning for framework services.
Build and maintain production observability with Grafana dashboards and Prometheus alerting across SSR and Glide platform layers.
Diagnose and resolve Node.js and JVM production incidents, including event-loop stalls, heap growth, and GC pressure.
Software Mind develops solutions for global companies, working with tech giants and unicorns on transformative projects. The company fosters a culture of openness, respect, grit, and enjoyment, building cross-functional engineering teams.