Build and deploy production-grade LLM inference systems from scratch, owning the pipeline from query to response.
Optimize inference workloads for latency, throughput, and cost using tools like vLLM, SGLang, and TensorRT-LLM.
Collaborate with the CTO and Product to define the technical roadmap for inference infrastructure as the organization scales.
The company is an open-source-oriented startup building production-grade inference infrastructure for large language models. It is a remote-first, globally distributed team that values engineering ownership, speed, and customer impact.
Deploy LLMs into production across GPU infrastructure, owning the full pipeline from customer query to served response.
Stand up and operate serving infrastructure using vLLM, SGLang, or TensorRT-LLM.
Apply quantization, batching, caching, and routing to optimize latency and cost at scale.
vCluster Labs is a venture-backed tech startup pioneering Kubernetes virtualization for the AI era, enabling AI Cloud providers and AI factories to operate GPU infrastructure with hyperscaler-like experiences. We raised over $30M from top-tier VCs like Khosla Ventures, are in a hyper-growth phase, and maintain a remote-first, distributed global team with headquarters in San Francisco.
Deploy and commission GPU cloud infrastructure across regional and core datacenters, covering hardware, networking, storage, and platform software.
Act as a hands-on technical escalation point, troubleshooting complex issues across physical and software layers and driving incidents to resolution.
Establish deployment standards, validation procedures, documentation, and operational practices for a rapidly evolving AI infrastructure environment.
The company builds and operates cutting-edge GPU cloud and AI infrastructure at significant scale. It is a fast-growing international scale-up with a diverse and flexible working environment.
Own the cost and performance of the inference stack, improving throughput and latency without compromising reliability.
Optimize through KV-cache management, continuous batching, speculative decoding, and quantization.
Work within serving engines like vLLM, SGLang, and TensorRT-LLM, profiling performance down to kernel level.
Adaption builds efficient AI that evolves in real-time, making intelligence flexible, personalized, and accessible to everyone. They focus on talent density, bringing together driven individuals to push the boundaries of continual adaptation.
Own the infrastructure layer for AI workloads including inference serving, Kubernetes, and agent-sandboxing platforms.
Manage the serving tier for open-weight models, Kubernetes operators, and stateful data planes.
Oversee the sandbox runtime, control-plane services, and observability tooling.
AZX accelerates positive impact in critical industries through AI transformation, specializing in physics-informed ML and enterprise AI solutions for climate and sustainability. Founded in 2024, the company is a profitable public benefit corporation with a growing team working with category leaders in real estate, energy, logistics, and utilities.
Develop and deploy machine learning and AI systems.
Work with LLMs, generative AI, and modern ML frameworks.
Optimize model performance, latency, and cost.
A fast-growing technology company building critical infrastructure that powers high-volume, real-time business operations across multiple systems and platforms. It is a collaborative, fast-moving environment where engineers have meaningful influence on architecture and product direction.
Participate in customer and sales meetings to explain technical capabilities and develop proof-of-concept solutions.
Troubleshoot and resolve complex customer issues through code analysis, scripting, and log analysis.
Collaborate with internal teams across engineering, sales, and product to improve platform and customer experience.
The company specializes in AI infrastructure and cloud environments, helping customers adopt and scale AI workloads. It is a remote-first, collaborative organization with a fast-growing team that values ownership and quick execution.
Own the AI architecture across Flip Fusion, Ask AI agent gateway, and Flip Flows.
Lead and grow a team of three AI engineers, staying hands-on with the hardest problems.
Shape a young AI engineering discipline with enormous design space and no inherited playbook.
Flip is the leading AI-powered employee experience platform for frontline workers in retail, manufacturing, and logistics. The company is a young tech company with a relaxed working atmosphere and highly motivated people.
Design and build production AI agent systems for diagnosing and remediating infrastructure issues in large-scale GPU environments.
Develop distributed services, orchestration frameworks, knowledge graphs, and retrieval systems to power AI agents.
Own services end-to-end from architecture through production, collaborating with infrastructure and engineering teams.
The company builds AI agents that operate and automate large-scale GPU infrastructure. The engineering team is highly collaborative and remote, fostering ownership and autonomy.
Lead product engagement with strategic AI infrastructure customers to define requirements and drive execution from discovery to production readiness.
Collaborate cross-functionally with engineering, infrastructure, and operations teams to deliver customer-ready solutions.
Translate complex customer needs into clear product priorities, technical specifications, and scalable AI infrastructure offerings.
Vultr provides high-performance cloud infrastructure solutions, including Cloud Compute, GPU, Bare Metal, and Storage, with 33 global data centers. It is a privately-held company valued at $3.5 billion, trusted by hundreds of thousands of customers across 185 countries, and known for its self-funded growth and inclusive culture.
Design and build production AI agent systems that diagnose, investigate, and remediate infrastructure issues across one of the world’s largest GPU fleets.
Build the distributed services, orchestration framework, knowledge graph, and retrieval systems that power infrastructure agents.
Own services end to end, including architecture, implementation, testing, deployment, observability, and production operations.
Together AI is a research-driven artificial intelligence company that builds open and transparent AI systems. The company has contributed to leading open-source research like FlashAttention and RedPajama, and aims to lower the cost of modern AI through co-designed software, hardware, algorithms, and models.
Act as the technical bridge between product and customer setups, translating business requirements into production-ready implementations.
Lead end-to-end rollouts of AI assistant configurations in production, including workflows, logic, data flows, prompts, and guardrails.
Implement integrations with third-party systems via APIs and webhooks, covering CRMs, calendars, ticketing, payments, and common auth patterns.
This company builds AI phone assistants for small and medium-sized businesses. The team is lean (~50 people) and growing rapidly with a strong engineering culture and a proprietary orchestration layer built in-house.
You will work directly within client environments, translating real operational needs into production-ready AI systems rather than proof-of-concept solutions.
The role spans technical discovery, architecture, software development, deployment, optimization, and ongoing delivery ownership.
You will design and build RAG systems, agentic AI workflows, backend services, and cloud-native architectures that operate at enterprise scale.
Jobgether uses an AI-powered matching process to connect candidates with roles at partner companies. They operate remotely and focus on efficient, fair hiring through automated shortlisting, with a commitment to data privacy and human oversight in final decisions.
Own the technical path from customer interest to working deployment, integrating the platform into production AI environments.
Build and operate AI/MLOps pipelines, debug complex environments, and create prototypes and demos.
Translate customer needs into product improvements, partnering with Sales, Product, and Engineering.
Neuromorphic Labs is a Seed-stage AI startup building a trust layer for production AI, making security, governance, and control intrinsic to every model and deployment. Backed by top-tier VCs, the team is small and fast-paced, emphasizing ownership, high standards, and collaboration with exceptional builders.
Bridge business needs and AI engineering to deliver production-ready solutions for enterprise customers.
Partner with cross-functional teams to identify high-value opportunities and turn them into scalable deployments.
Take complex workflows from discovery and prototyping through production adoption and optimization.
The company is a technology firm specializing in AI-powered enterprise solutions. It fosters a collaborative culture focused on innovation and continuous learning, with a team of professionals dedicated to driving customer success.
Design, develop, and deploy AI models to solve complex business challenges.
Collaborate with cross-functional teams to integrate AI solutions into production systems.
Continuously improve and optimize AI pipelines for performance, scalability, and reliability.
We are passionate about helping businesses navigate the rapidly changing and complex world of emerging technologies. We create well-structured, secure, scalable solutions at speed to provide the foundation for groundbreaking change.
Design and build reusable platform capabilities for LLM applications, AI agents, RAG, and tool calling.
Develop scalable data pipelines and secure integrations connecting AI applications to enterprise systems.
Establish technical standards for AI deployment, observability, evaluation, and lifecycle management.
Jobgether uses AI-powered matching to connect candidates with roles. They operate as a global, fully distributed international team with colleagues across multiple countries, offering a collaborative and inclusive environment.
Design, build, and own shared infrastructure for Interpretability research environments, data systems, and compute tooling.
Lead cross-team efforts with agentic engineering, security, compute, and storage platform teams.
Discover and resolve major organization-wide developer experience issues and help take interpretability methods from research code to dependable audit pipelines.
Anthropic is an AI safety company focused on building reliable, interpretable, and steerable AI systems. They are a quickly growing team of researchers, engineers, policy experts, and business leaders committed to beneficial AI.
They are a technology company focused on integrating AI capabilities into customer-facing products and internal workflows. The team is fast-moving and product-focused, with a culture of ownership and continuous iteration.
Build AI-powered tools and copilots across the SDLC to reduce cognitive load and eliminate manual steps.
Research and deploy GenAI solutions to improve delivery pipelines and system reliability.
Collaborate with Platform, SRE, and DevOps teams to integrate intelligent automation into the core engineering platform.
Coderio designs and delivers scalable digital solutions for global companies. They combine strong technical expertise with a product mindset and value autonomy and clear communication.