Design, build, and own shared infrastructure for Interpretability research environments, data systems, and compute tooling.
Lead cross-team efforts with agentic engineering, security, compute, and storage platform teams.
Discover and resolve major organization-wide developer experience issues and help take interpretability methods from research code to dependable audit pipelines.
PythonCloud InfrastructureKubernetesSecurityDistributed Systems
Design and write high-performing scalable software for training models.
Develop new tools to support and accelerate research and LLM training.
Collaborate with engineering teams and scientific teams to implement experiments on cluster and data infrastructure.
Cohere is a security-first enterprise AI company building cutting-edge foundation AI models and end-to-end products for real-world business problems. The company is a global team of researchers, engineers, and designers passionate about AI, headquartered in Toronto with offices worldwide.
Design and build sandboxed evaluation environments for AI models to safely execute code and interact with tools.
Build backend services and infrastructure supporting large-scale AI and agentic evaluations.
Develop agent scaffolding, evaluation harnesses, and systems for provisioning isolated environments using Docker, Kubernetes, and cloud infrastructure.
10a Labs is the safety and threat-intelligence layer trusted by frontier AI labs, AI unicorns, Fortune 10 companies, and leading global technology platforms. The company focuses on adversarial red teaming and model evaluations to help teams deploy AI systems safely.
Lead and mentor a team of engineers specializing in GPU infrastructure, fostering a culture of technical excellence and continuous improvement.
Define and execute the technical roadmap for GPU cluster deployment, optimization, and scaling.
Partner with AI researchers to understand emerging infrastructure needs and translate them into robust solutions.
Cohere is a security-first enterprise AI company that builds cutting-edge foundation models and end-to-end products for enterprises. It is a global technology company with a team of researchers, engineers, and designers passionate about their craft.
Design and build the AI execution platform with event-triggered workflows and model-agnostic runtimes.
Develop evaluation layers with golden test suites and safety checks to ensure AI reliability.
Implement governance mechanisms and build AI agents supporting business workflows.
Jobgether is a platform that connects talent with opportunities using AI-powered matching. The company has a globally distributed team and a remote-first culture.
Design and deploy scalable AI infrastructure and agent systems for enterprise customers.
Work on Kubernetes cluster design, multi-agent system architecture, and CI/CD pipelines.
Engage directly with customers to assess needs and present technical recommendations.
LangChain builds the foundation for agent engineering, helping developers create production-ready AI agents. With $125M raised at Series B from top venture firms and 100M+ monthly open source downloads, they have a strong engineering culture and meaningful team impact.
Architect, design, build, deploy, and maintain Model Serving infrastructure using industry-standard AI tools.
Own projects that scale model serving and data processing services to handle 10x traffic.
Collaborate closely with MLE and Data Science teams to distill feedback and execute on strategy.
Abnormal AI protects the humans behind the world's most critical organizations from AI-powered cybercrime. Over 4,500 enterprises trust their behavioral AI platform, fostering a culture of innovation and security.
You own the platform runtime, voice infrastructure, enterprise security, and evaluation layer for all clients.
Work directly with the founder on architecture, from design through production and incident response.
Build self-serve infrastructure, lead incident response, and keep the platform fast and reliable on Kubernetes.
Rifa AI builds an AI agents platform for contact centers in regulated industries, helping enterprises deploy trustworthy voicebots. They are a small, passionate engineering team with paying enterprise clients and growing revenue, backed by Seaborne Capital.
Own and optimize CI/CD pipelines, Kubernetes deployment, and infrastructure for model serving and inference.
Build telemetry, observability, and alerting to catch real problems and reduce noise.
Eliminate toil through thoughtful automation and improve developer and agent productivity.
Obvious is building an AI-native workspace that serves as an operating system for work, putting co-intelligence at the center. They are a small, talent-dense team with founders and leaders from top tech companies.
Lead end-to-end deployment of North in private cloud and on-premises environments, including planning, configuration, testing, and rollout.
Partner with enterprise IT teams to assess infrastructure, security requirements, and data management practices.
Design and implement deployment strategies tailored to client needs, ensuring compliance with data privacy and security standards.
Cohere is a security-first enterprise AI company that builds cutting-edge foundation models and end-to-end products. It is a global team of researchers, engineers, and designers passionate about their craft.
Partner directly with customers to architect, build, and deploy high-scale production AI applications on Baseten's platform.
Own the journey from initial exploration to production deployment, translating ambiguous goals into reliable services.
Work across product, software development, performance engineering, and customer-facing implementations.
Baseten powers mission-critical inference for dynamic AI companies like Cursor and Notion. They are rapidly growing, recently raised a $1.5B Series F, and foster a collaborative, forward-thinking culture.
Architect and deploy autonomous AI agents and multi-agent workflows for privacy-first systems.
Build scalable backend services using FastAPI and orchestrate agentic workflows with LangGraph in AWS or Azure.
Develop rigorous evaluation pipelines for accuracy, citation adherence, latency, and reliability.
Osano is a leading data privacy platform that helps organizations comply with global privacy regulations like GDPR and CCPA. Backed by top-tier investors and recognized as a Great Place to Work for four years running, the company has a 97% employee satisfaction rate and a mission-driven, fast-growing culture.
Lead end-to-end development of advanced AI solutions including agentic systems, RAG pipelines, and multi-agent workflows.
Architect AI infrastructure with guardrails, access controls, and compliance protocols for ethical deployment.
Guide multi-disciplinary teams, setting technical playbook and best practices.
Nava is a consultancy and public benefit corporation that makes government services simple and effective. Since 2015, federal, state, and local agencies have trusted Nava to solve technology modernization challenges, and they maintain a collaborative, remote-friendly team environment.
Build and operate Lumin's security fabric engineered for reliability and scale across hundreds of environments.
Design and maintain agentic AI workflows using tools like Claude Code and MCP integrations to automate security platform engineering.
Write production-quality Python applications and tooling supporting platform operations and AI-assisted workflows.
Lumin Digital is a trailblazer in digital banking solutions, providing cloud-native digital experiences for credit unions and banks. We foster a culture of curiosity, innovation, trust, and respect, with a focus on collaboration and continuous improvement.
Deploy and operate Blitzy's self-hosted platform within a customer-controlled, secure cloud environment.
Own the Kubernetes-based deployment, releases, upgrades, capacity planning, and performance benchmarking.
Serve as the on-account technical presence, partnering with customer infrastructure and security teams.
We are an AI software development platform that autonomously builds custom software for enterprises. Backed by tier 1 investors and led by two co-founders, we are one of the fastest-growing U.S. companies with a culture of speed and customer focus.
Design and ship security workflows combining deterministic analysis with LLM reasoning to find real vulnerabilities across languages and frameworks.
Engineer agentic pipelines and prompts that are precise, cost-aware, and trustworthy for security-critical work.
Push on hard problems in automated triage and validation to close the gap between finding and actionable fix.
Semgrep is a code security platform that helps teams catch and fix vulnerabilities before they ship. They are a venture-backed startup with a transparent culture that values respect and honesty.
Build and operate the infrastructure behind AI-powered products, improving reliability, security, scalability, and cost efficiency.
Write code, automate infrastructure, investigate production issues, and design systems that reduce operational complexity.
Take ownership of unfamiliar systems, identify highest-leverage improvements, and balance immediate production needs with long-term platform investments.
Zencoder builds and orchestrates AI agents that ship real work across code, research, and operations. It is a growing platform where people and agents collaborate, with a high-caliber team and a culture that values individual contributors.
Own the reliability, security, and infrastructure for the AI operations platform running sandboxed agents.
Join a newly formed SRE team to build reliability practice from scratch on real infrastructure.
Manage distributed systems, observability, incident response, and automation with a security-first mindset.
Duvo builds an AI operations platform for retail and CPG enterprises to automate data workflows across systems. They are a fast-moving, humble team focused on solving real customer problems with strong traction.
Build agent-driven workflows for cloud security platforms like FedRAMP High Cloud-Native SIEM and Tier 1 Agentic SOC.
Develop autonomous AI tools including triage, investigation, and reporting agents using agentic AI frameworks.
Translate complex systems requirements into production-ready AI solutions following AWS Well-Architected and NIST security controls.
First Due provides transformative, end-to-end software solutions for fire and EMS agencies to improve safety and effectiveness. The company offers a comprehensive benefits package and promotes a fully remote, inclusive work environment.
Design, build, and operate Kubernetes infrastructure for AI workloads using Terraform and GitOps.
Define SLOs, run incident response, and create runbooks for reliable AI platform operations.
Drive AI-specific observability, FinOps, and security practices across the platform.
We are an AI-native consulting partner working with clients like PayPal, adidas, and NatWest to build digital products and services. Our team of over 600 has scaled quickly, earning Great Place to Work-Certified status multiple years in a row.