Develop and improve core AI methods and systems for reliable AI agents across the full lifecycle.
Create novel approaches for simulation, evaluation, and optimization of agent behavior in production.
Turn research ideas into working prototypes and production-facing capabilities.
This is an early-stage AI infrastructure company focused on making AI agents reliable in production. The company values innovation and practical deployment, with a small team driving frontier AI research and product development.
Build and ship AI methods for federal research data with LLMs, NLP, RAG, knowledge graphs.
Benchmark AI against baselines and human reviewers to prove gains in accuracy, reliability, efficiency, cost.
Set standards for trustworthy AI and mentor data scientists through code review and guidance.
ARI builds AI methods for federal research-funding and administrative data, with clients including the NIH. It is a small multidisciplinary team of scientists, analysts, economists, and data scientists experienced with NIH data.
Own the build out of new agents, skills, and platform capability for teams across TLDR.
Build and deploy agents end to end, from design through implementation, evals, and rollout to internal users.
Partner with stakeholders across sales, editorial, and people ops to find where an LLM belongs in their process.
TLDR runs the largest network of tech newsletters in the world, with over 8 million subscribers covering startups, software engineering, AI, and more. Our 31-person full-time team is bootstrapped, profitable, and on track for $35M in revenue this year, with a culture of owning functions rather than slices.
Build reinforcement learning and agent environments for real customer and Frontier lab use cases, including task specifications, scoring, and evaluation.
Develop benchmarks and evaluation harnesses to measure model and data quality across accuracy, robustness, safety, latency, and cost.
Run fine-tuning, adapter, and other model experiments to evaluate how data and methods influence model behavior, and deploy local or self-hosted models for evaluation and inference.
Appen has been a leader in AI training data for over 30 years, specializing in human-generated data to train, fine-tune, and evaluate models across generative AI, LLMs, computer vision, and speech recognition. They support model development through an AI-assisted data annotation platform and a global crowd of over 1 million contributors in more than 200 countries, fostering a culture of innovation, collaboration, and humility over ego.
Develop, configure, deploy, and optimize AI agents using Cresta's AI platform and tools.
Build AI agent integrations with external systems such as APIs, databases, and CRMs to ensure seamless workflow integration.
Collaborate with customers and internal stakeholders to gather technical requirements and translate business needs into AI agent solutions.
Cresta builds a unified AI platform that combines conversational AI agents, real-time human agent augmentation, and conversation intelligence to drive revenue and efficiency for customer experience channels. Born from Stanford AI Lab, the company has raised over $270 million from top investors and is led by AI industry veterans.
Develop and deploy cutting-edge AI/ML solutions to enhance the platform and improve student learning experiences.
Design, develop, and optimize LLM-powered agentic systems and APIs for real product use cases.
Collaborate with senior engineers and contribute to evaluation frameworks and MLOps pipelines.
Interview Kickstart specializes in interview preparation and career transitions into high-demand tech fields like AI, ML, and Data Science. Over 17,000 tech professionals have been guided by current and former hiring managers to land coveted positions at companies like Google and Amazon.
Build RL environments, agentic systems, LLM pipelines, and evaluation frameworks for real-world AI use cases.
Develop benchmarks and evaluation harnesses to assess model quality across accuracy, safety, latency, and cost.
Conduct fine-tuning and model experiments, deploy self-hosted models, and document reproducible methodologies.
The employer is an organization focused on applied AI research, building practical and reusable AI systems for real-world use cases. It values curiosity, accountability, innovation, collaboration, and continuous learning in a remote environment.
Design, develop, and maintain AI-powered software solutions across the full development lifecycle.
Work with backend stacks like Python, Node.js, .NET, and Java, plus modern frontend frameworks such as React or Angular.
Build and integrate AI agents, RAG pipelines, and LLM-driven products into production environments.
Cresteo is a nearshore tech services company on a mission to be the world leader in people-first and honest software development. Our small, high-trust team values transparency, profit-sharing, and fearless innovation while working with US-based clients and international teams.
Collaborate with client teams to diagnose operational bottlenecks and develop testable hypotheses.
Build and test AI-powered prototypes using coding agents like Claude Code or Cursor.
Drive adoption by working directly with users and iterating on solutions until they are effective in live workflows.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. The platform uses AI to review applications and provide a shortlist to employers, emphasizing efficiency and objectivity.
Design and build production-grade agentic AI systems for marketing and customer experience.
Architect agent workflows with reasoning, tool use, retrieval, and performance monitoring.
Evaluate emerging frameworks like LangGraph and LangChain for enterprise applications.
The company specializes in developing production-grade agentic AI systems for enterprise marketing and customer experience. It fosters a remote-first, collaborative culture focused on innovation and technical excellence.
Design, build, and ship full-stack applications for admissions, events, and internal operations.
Develop semantic search and AI features using embeddings, natural language queries, and large datasets.
Work directly with operational teams to understand needs and make practical product and engineering tradeoffs.
A venture capital and startup accelerator organization that builds software for founders and internal teams. It operates as a small product engineering team with a collaborative, low-ego culture.
Collaborate directly with customers as an embedded technical expert, designing and deploying agentic AI solutions.
Quickly understand new industries, data, and systems to identify high-value AI and agentic workflows.
Build proof-of-concepts and production solutions using MCPs, agent frameworks, and tool-using LLMs.
We are a fast-growing company that helps customers turn AI ambitions into real production outcomes. We are fully remote with a 32-hour workweek and a collaborative, trust-based culture.
Drive end-to-end program management for complex, cross-functional AI and autonomy initiatives with engineering teams and external partners.
Act as a technical thought partner, translating between business stakeholders, R&D engineers, and partner teams on AI systems, LLMs, and agentic frameworks.
Establish success criteria, manage risks, and optimize processes to deliver measurable program value across the AI & Autonomy function.
Rockwell Automation is a global technology leader in industrial automation and digital transformation, helping manufacturers be more productive and sustainable. With approximately 28,000 employees across more than 100 countries, they foster a culture of problem solvers and forward thinkers.
Lead the engineering strategy and execution for evaluations of AI agents, owning the core evaluation platform.
Design scalable evaluation infrastructure, APIs, workflows, and production systems across software categories.
Mentor and develop a team of engineers as the technical authority on agentic evaluation.
This company focuses on building credible, scalable evaluations of AI agents from software vendors. It operates as a fully remote, inclusive team with a flexible culture and a focus on professional growth.
Architect and build scalable Generative AI and agentic AI applications from prototype to production.
Design LLM-powered workflows, prompt strategies, and multi-agent architectures using LangChain, LangGraph, or similar.
Collaborate directly with customers, product leaders, and engineering teams to translate business requirements into robust AI solutions.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. They are a fast-moving technology environment with a focus on AI-driven recruitment processes and a culture of innovation and objectivity.
Own product features end-to-end, from problem statement to production.
Write specs with AI agents and critically review their code in complex systems.
Deliver reliable features handling financial transactions and evolve AI-first development.
Social Discovery Group creates social entertainment platforms that connect people online across cultures, addressing loneliness and disconnection. The company's international remote team works from everywhere, and it earned 'Great Place to Work' recognition in 2024-2025.
Design and develop specialized AI agents using Generative AI, LLMs, and agentic frameworks like LangChain and Semantic Kernel.
Build MCP servers and integrate AI applications with event-driven architectures, ensuring secure and traceable decisions.
Implement Human-in-the-Loop workflows and apply PromptOps practices for scalable, secure, and high-impact AI solutions.
The hiring company is a technology organization in Brazil focused on building AI-powered applications and agentic systems. It fosters a collaborative and international culture, with fully remote work and opportunities to work on innovative AI technologies.
Design and implement autonomous agents and multi-agent orchestration systems for complex, open-ended tasks.\n- Build and optimize production-grade LLM applications with a focus on reliability, observability, and low-latency performance.\n- Architect advanced RAG pipelines and vector database strategies to provide agents with accurate, real-time context.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies through objective evaluation. The company operates with a distributed, global-first team culture, emphasizing fairness and innovation in recruitment.
Own product features end-to-end, from problem identification to production.
Work with AI agents to create specs and delegate development tasks.
Critically review AI-generated code to ensure reliability and quality.
Our partner company is seeking a Product Engineer to own features end-to-end in an AI-native environment. They are a remote, international team with an AI-first culture and a focus on using AI agents as primary development partners.
Design and develop AI-powered applications using LLMs, Generative AI, and Machine Learning.
Build AI Copilots, AI Agents, and RAG solutions to improve software engineering and operational workflows.
Integrate AI capabilities into enterprise applications, APIs, Azure cloud services, and DevSecOps pipelines.
Planned Systems International (PSI) is a technology solutions provider specializing in IT services for government agencies. They foster a collaborative culture focused on innovation and professional growth.