Design, scale, and maintain enterprise monitoring and alerting ecosystems across multi-cloud and native systems.
Bridge development and operations to ensure high availability, performance tuning, and deep visibility.
Automate infrastructure and build robust observability pipelines using cloud-native tools like Prometheus, Grafana, and GCP.
Ontrac Solutions is a leading technology consulting firm specializing in cutting-edge solutions that drive business transformation. Their team is committed to innovation, collaboration, and excellence, empowering clients to succeed in an evolving digital landscape.
Own and improve service reliability for the product team: design for HA/performance/scale, define SLIs/SLOs
Align with org standards, implement DevOps-driven updates: support processes, templates, services, breaking changes, security fixes
Build and evolve GitLab CI/CD for build, test, security scans, and progressive delivery; speed up and harden pipelines
Plata Card is a fintech company focused on cards and accounts services. They foster a high-tech environment with a supportive team and innovative spirit.
Design, build, and scale reliable infrastructure for Klover's fintech platform using modern technologies like Kubernetes, Terraform, and Istio.
Use AI agents as force multipliers to automate manual processes and improve developer experience.
Collaborate with engineering teams to ensure system reliability, performance, and security across production systems.
Attain powers Klover, a fast-growing fintech platform serving over one million active users monthly, processing over $1.5 billion annually. The company emphasizes collaboration, reliability, and innovation, with a culture of automation and AI-driven development.
Productize deployment, security, and scaling of Applied AI solutions with automation and security guardrails.
Mistral provides full-stack AI solutions from frontier models to developer tools, applications, and compute, partnering with enterprises across high-stakes industries. It is a dynamic, collaborative team with a diverse workforce distributed globally, known for being creative, low-ego, and team-spirited.
Define and implement SLIs/SLOs for critical services, lead incident response, and conduct blameless postmortems to drive systemic improvements.
Design and improve monitoring and alerting with Prometheus and Grafana, build internal tooling, and automate operational workflows to reduce toil.
Partner with cross-functional engineering teams to improve system resilience, contribute to architectural discussions, and strengthen production readiness standards.
Runpod provides a cloud platform for AI development, used by over one million developers for training, fine-tuning, and deploying AI models. The company is a small, remote-first team that closed a $100M Series A in June 2026, emphasizing ownership, speed, and impact at scale.