Turning measurements of automated AI R&D into concrete policy proposals for decision makers.
Designing and running independent audits of frontier AI developers from scoping to defensible findings.
Working with high agency and comfort with ambiguity to advance AI safety.
P-Zero Research is a public benefit corporation improving the long-term trajectory of AI by forecasting and mitigating risks from automated AI R&D. The team values high agency, clear communication, and alignment with its mission.
Contribute to the research agenda for enterprise AI economics, addressing investment prioritization, budgeting, and value realization.
Advise CIOs and CFOs on AI portfolio decisions, including build or buy choices, funding, and scaling.
Produce syndicated and custom research, brief clients, and represent IDC in executive discussions and industry events.
IDC is the premier global provider of trusted technology intelligence, equipping business and technology leaders with evidence for confident decisions. With over 1,000 analysts worldwide, IDC has been recognized as Analyst Firm of the Year for five consecutive years, fostering a culture of growth, collaboration, and long-term career development.
Run open-ended research projects using the internet as the primary tool.
Turn messy findings into clean outputs: matrices, briefs, landscape maps.
Use AI tools aggressively and build lightweight tools when needed.
A private design studio building and supporting a portfolio of companies across software, hardware, hospitality, and more. An affiliate of Expa, it also runs an early-stage investment fund and a founder cohort program.
Build reinforcement learning and agent environments for real customer and Frontier lab use cases, including task specifications, scoring, and evaluation.
Develop benchmarks and evaluation harnesses to measure model and data quality across accuracy, robustness, safety, latency, and cost.
Run fine-tuning, adapter, and other model experiments to evaluate how data and methods influence model behavior, and deploy local or self-hosted models for evaluation and inference.
Appen has been a leader in AI training data for over 30 years, specializing in human-generated data to train, fine-tune, and evaluate models across generative AI, LLMs, computer vision, and speech recognition. They support model development through an AI-assisted data annotation platform and a global crowd of over 1 million contributors in more than 200 countries, fostering a culture of innovation, collaboration, and humility over ego.
Own Merlin's foundation and world-model work, including architecture selection, post-training, and capability roadmap.
Lead and mentor a small team of world-model engineers, setting the technical bar and review culture.
Design model interface to the autonomy stack with structured, schema-constrained plan outputs and build evaluation harnesses.
Merlin is a publicly traded aerospace and defense company building a non-human pilot for full-stack aircraft autonomy. Headquartered in Boston, it is expanding its organization to accelerate the deployment of its autonomy platform.
Study how engineers and customers use Archie in the field, identify where it succeeds or fails, and turn observations into actionable evaluations.
Recreate real-world engineering tasks and failure modes in repeatable environments for development teams.
Build and refine evaluation methodologies, including automated judges and human evaluation processes, to align with expert judgment.
P-1 AI is building Archie, an AI engineer agent for the physical world that works alongside human engineering teams. The company recently raised a $50 million Series A led by NEA and is driven by the mission of building superintelligence for engineering.
Build RL environments, agentic systems, LLM pipelines, and evaluation frameworks for real-world AI use cases.
Develop benchmarks and evaluation harnesses to assess model quality across accuracy, safety, latency, and cost.
Conduct fine-tuning and model experiments, deploy self-hosted models, and document reproducible methodologies.
The employer is an organization focused on applied AI research, building practical and reusable AI systems for real-world use cases. It values curiosity, accountability, innovation, collaboration, and continuous learning in a remote environment.
Develop and improve core AI methods and systems for reliable AI agents across the full lifecycle.
Create novel approaches for simulation, evaluation, and optimization of agent behavior in production.
Turn research ideas into working prototypes and production-facing capabilities.
This is an early-stage AI infrastructure company focused on making AI agents reliable in production. The company values innovation and practical deployment, with a small team driving frontier AI research and product development.
Develop agentic research systems to automate mechanistic model development for biochemical and physiological processes.
Design and train machine learning models for in-context learning and few-shot adaptation on molecular and tabular data.
Collaborate with chemists and biologists to drive drug discovery decisions using AI models.
PostEra is building an AI-first biotech, using its Proton AI platform to accelerate the discovery of new medicines for patients. The company has a minimalist organizational structure, celebrating individual achievement through proportional compensation and collaboration toward delivering cures.
Build tooling for capturing and processing data from agents and humans at significant scale.
Solve hard problems around compute, orchestration, scaling, security, and reliability.
Help develop approaches for training, benchmarking, and evaluating AI agents.
Prolific builds human data infrastructure for AI development, connecting researchers with a global pool of participants to collect high-quality, ethically sourced behavioral data. They are a mission-driven company at the forefront of AI innovation, with a remote culture and a focus on impactful work.
Set the technical direction for AI engineering across the team: agent architecture patterns, evaluation methodology, deployment and monitoring strategies
Design the AI platform layer, including shared agent frameworks, tool integrations, and evaluation infrastructure
Work directly with clients on the most complex engagements, identifying new problem domains and ensuring production quality
Addepar is a global data and AI platform empowering investment professionals to turn complex financial information into actionable intelligence. With over 1,500 firms in 60 countries managing nearly $10 trillion in assets, Addepar fosters a culture of ownership, collaboration, and innovation.
Conduct applied research in machine learning, focusing on time-series forecasting and generative AI for supply chain planning.
Design and run experiments using real-world and benchmark datasets, evaluating models against strong baselines.
Collaborate with researchers and engineers to build prototypes and assess product potential for real-world applications.
Kinaxis is a global leader in modern supply chain orchestration, powering complex global supply chains with an AI-infused platform. We are a global team of over 2,000 employees with a best-in-class HQ in Ottawa, Canada, and we take our culture seriously.
Create realistic PowerPoint tasks based on private equity or corporate development experience.
Complete each task to establish benchmark outcomes for AI-generated work.
Develop objective evaluation criteria and provide source materials for AI assessments.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. This role is posted on behalf of a partner company, and the hiring process is managed by their internal team.
Own the technical win with AI labs and inference platforms, getting in front of researchers and ML engineers.
Build demos and integrations in customer stacks, running evaluations and pilots to a verdict.
Own the platform channel and feed real customer needs back to product and engineering.
Keenable builds AI infrastructure for web search and knowledge access at web scale. We are a small, high-talent team with high autonomy, low process, and direct exposure to frontier customers.
Craft creative prompts and multi-turn scenarios to stress-test AI guardrails across diverse risk categories.
Discover ways around safety filters and restrictions using jailbreak, evasion, and prompt injection techniques.
Evaluate and score model responses against structured harm taxonomies and severity rubrics.
Handshake AI partners with leading AI research labs to make models safer and more robust. Our red teaming operations help identify vulnerabilities before they reach users, contributing directly to the responsible development of frontier AI systems.
Build Eon's growth engine using AI, automation, and data pipelines.
Identify and build technical tools and distribution channels to create demand.
Work closely with marketing and revenue operations to test new GTM ideas.
Eon is building a new category around cloud backup data, helping enterprises turn their backup infrastructure into searchable, accessible, and AI-ready data. They are a high-growth company backed by Sequoia, Lightspeed, and others, with a team including founders of CloudEndure, where small teams have outsized impact.
Lead a high-performing team to design and deliver innovative, scalable learning experiences for Zscaler's GTM organization.
Define and evolve the learning and experience design strategy and operating model to drive measurable impact on business OKRs.
Oversee the entire learning lifecycle, from intake and prioritization to measurement and continuous improvement.
Zscaler accelerates digital transformation with cloud security, protecting thousands of customers from cyberattacks. With a global presence and a culture of ownership and innovation, the company is committed to an AI-native future.
Design high-converting landing pages using our AI engine and own the creative process.
Collaborate with marketing managers and designers to deliver polished, timely work.
Give feedback to shape the AI landing page engine.
Uplane builds AI technology for the marketing agency of the future, helping creative teams generate ads and high-converting landing pages. It is a fast-growing, VC-backed startup with an ambitious, fun, and humble culture.
Drive down proving cost for zero-knowledge proofs of AI inference.
Build production-quality proof systems for frontier-scale models.
Own the path from research prototype to deployable verification tools.
SASH's Verification team builds and tests tools to verify international agreements about AI, prototyping new verification mechanisms and working with policymakers. They are a small, globally distributed team combining technical expertise with international AI policy experience.
Create original graduate- and PhD-level problems within your STEM expertise.
Develop rigorous, step-by-step solutions and review AI-generated responses for accuracy.
Work fully asynchronously through an online platform to contribute to AI reasoning research.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. They use technology to streamline recruitment and are part of a network of partner companies.