Build tools for understanding traces, evaluating agent behavior, and turning results into improvements.
Run experiments on Mastra's primitives to find where they fall short and improve configuration and behavior.
Help turn findings into capabilities developers can use in their own applications.
Mastra is the open-source TypeScript framework for building AI agents. We're a small, fully remote team backed by Y Combinator and Spark Capital with $35M raised.
Own the agentic development environment, ensuring AI agents can operate autonomously in cloud-based environments and execute full test suites.
Build MCP server integrations to connect agents to internal systems like CircleCI, Slack, Datadog, and GitHub.
Drive enablement by owning repo-wide agent documentation and developing tooling for engineers and non-engineers.
Hightouch is a data activation platform that helps companies sync customer data to their business tools. It is a remote-first startup with a fast-paced, high-ownership culture.
Build reinforcement learning and agent environments for real customer and Frontier lab use cases, including task specifications, scoring, and evaluation.
Develop benchmarks and evaluation harnesses to measure model and data quality across accuracy, robustness, safety, latency, and cost.
Run fine-tuning, adapter, and other model experiments to evaluate how data and methods influence model behavior, and deploy local or self-hosted models for evaluation and inference.
Appen has been a leader in AI training data for over 30 years, specializing in human-generated data to train, fine-tune, and evaluate models across generative AI, LLMs, computer vision, and speech recognition. They support model development through an AI-assisted data annotation platform and a global crowd of over 1 million contributors in more than 200 countries, fostering a culture of innovation, collaboration, and humility over ego.
Own the AI coding agent's context and harness system design.
Own online and offline evals for the agent.
Leverage 100k+ users and 200k+ repos to guide experimentation.
Hercules builds an AI coding agent designed to make developers more productive. It is a startup with a founding team in San Francisco that values speed, excellence, and a bias for action.
Evaluate software engineering tasks for technical accuracy, realism, and reproducibility.
Investigate codebases, tests, and integration issues to identify technical weaknesses.
Provide clear, actionable feedback that directly improves AI training and evaluation workflows.
Jobgether is an AI-powered job platform that connects candidates to roles through objective, skill-based matching. It focuses on remote and freelance opportunities, with a data-driven recruitment process and a global candidate pool.
Direct the agent array on production workstreams by decomposing problems into tasks and integrating agent output into shipped software.
Review agent-generated pull requests at volume and depth, identifying correctness, security, and accessibility defects.
Author evaluation suites that make quality measurable using eval-driven development and own end-to-end quality within a FedRAMP-authorized environment.
Granicus provides cloud-based solutions for government communications, website design, meeting management, and records management, serving over 5,500 agencies and 300 million citizens. With a globally distributed team and a culture of transparency and inclusion, Granicus has been recognized on the GovTech 100 list for the past 5 years.
Lead technical evaluations and proofs of concept for an enterprise AI coding platform in real customer environments.
Engage directly with CTOs, VPs of Engineering, and senior developers on AI agent architecture and integration.
Manage enterprise security and deployment reviews, troubleshoot integration issues, and feed customer insights into product.
The company is an enterprise AI coding platform provider helping engineering organizations accelerate development with AI agents and LLM infrastructure. It operates as a remote-first global team with a fast-paced, innovative culture focused on collaboration and rapid product iteration.
Architect and lead the development of an AI-first coding platform where agents and developers collaborate in real time.
Build agentic systems that plan tasks, write code, run tests, and deploy applications autonomously.
Integrate LLMs like Claude into production workflows, designing tool-use and agent orchestration layers.
Ubiminds is a GPTW-certified, people-first company that scales development teams for American software product companies. We connect Brazil's top 5% talent with U.S. companies, fostering a culture of improvement and teamwork.
Evaluate AI-generated coding interactions end to end for correctness and engineering judgment.
Assess whether outputs reflect strong engineering taste and provide clear, opinionated feedback.
Help define what great looks like for AI coding tools like Codex, Claude Code, and Cursor.
G2i Inc. is a technology staffing company that connects software engineers with remote contract opportunities. The company values engineering excellence and provides flexible, ongoing projects for senior-level developers.
Participate in a remote video interview about your experience with AI coding tools.
Discuss how tools like GitHub Copilot and Cursor fit into your development workflow.
Share feedback on the strengths and limitations of AI-assisted coding.
Our partner company is a research organization focused on AI-assisted software development. It is currently seeking feedback from practicing software engineers through paid interviews.
Own product features end-to-end, from problem identification to production.
Work with AI agents to create specs and delegate development tasks.
Critically review AI-generated code to ensure reliability and quality.
Our partner company is seeking a Product Engineer to own features end-to-end in an AI-native environment. They are a remote, international team with an AI-first culture and a focus on using AI agents as primary development partners.
Design and deploy AI agents that automate sales, operations, and support workflows.
Own technical delivery end-to-end, from architecture to production and post-launch improvements.
Work directly with clients to understand workflows and explain technical tradeoffs.
Stello is an AI transformation firm that partners with mid-market and enterprise companies to automate repetitive work with AI agents. They work across sales, operations, and customer support, deploying solutions inside existing client tools.
Run open-ended research projects using the internet as the primary tool.
Turn messy findings into clean outputs: matrices, briefs, landscape maps.
Use AI tools aggressively and build lightweight tools when needed.
A private design studio building and supporting a portfolio of companies across software, hardware, hospitality, and more. An affiliate of Expa, it also runs an early-stage investment fund and a founder cohort program.
Architect, build, and optimize high-performance production LLM systems while maintaining a strong personal technical presence on the team.
Spearhead strategic technological changes and champion code refactoring efforts to keep the core codebase cutting-edge and performant.
Lead technical story breakdowns, architectural design, and mentor engineers across the department.
Appian provides AI automation for mission-critical work, automating complex processes in large enterprises and governments. With over 25 years of experience, the company is known for its reliability and scale, and fosters an inclusive culture with employee-led affinity groups.
Develop and deploy cutting-edge AI/ML solutions to enhance the platform and improve student learning experiences.
Design, develop, and optimize LLM-powered agentic systems and APIs for real product use cases.
Collaborate with senior engineers and contribute to evaluation frameworks and MLOps pipelines.
Interview Kickstart specializes in interview preparation and career transitions into high-demand tech fields like AI, ML, and Data Science. Over 17,000 tech professionals have been guided by current and former hiring managers to land coveted positions at companies like Google and Amazon.
Own product features end-to-end, from problem statement to production.
Write specs with AI agents and critically review their code in complex systems.
Deliver reliable features handling financial transactions and evolve AI-first development.
Social Discovery Group creates social entertainment platforms that connect people online across cultures, addressing loneliness and disconnection. The company's international remote team works from everywhere, and it earned 'Great Place to Work' recognition in 2024-2025.
Build production Agentforce agents including topics, actions, grounding, guardrails, and evaluation harnesses.
Implement actions across Flow, Apex, prompt templates, and external APIs, grounding answers in Data 360.
Design guardrails against prompt injection and data leakage, and run evaluation and regression tests before changes.
AspenView Technology Partners builds high-performing nearshore IT teams for North American clients, focusing on innovation and efficiency. They are a people-first, purpose-driven company that believes great culture drives great outcomes.
Curate high-quality code examples and datasets for LLM training and evaluation.
Develop and assess AI-generated software across multiple programming languages and the full SDLC.
Collaborate with research teams to design verification mechanisms and improve coding benchmarks.
This partner company specializes in evaluating large language models and improving AI systems through rigorous engineering benchmarks. It offers a remote, collaborative culture where engineers and researchers advance AI evaluation workflows together.
Maintain and evolve platform foundations, tools, and processes for agentic development to keep n8n operating as an AI-native engineering organization.
Build and integrate agentic workflows across Linear, GitHub, CI, and developer review processes while ensuring safety and security.
Define quality metrics, benchmarks, and adoption signals to measure agent output and improve engineering productivity.
n8n is an open workflow orchestration platform built for the new era of AI, giving technical teams the freedom of code with the speed of no-code. Since 2019, the company has grown to over 260 employees across Europe and the US, with a community of 650,000+ developers, 190K+ GitHub stars, and a $5.2bn valuation.
Participate in a paid remote video call about your AI-assisted development workflow.
Share how you use tools like GitHub Copilot, Cursor, or Claude Code to write and debug code.
Discuss strengths, limitations, and prompt strategies of current code generation models.
Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.