Applied Machine Learning Scientist - Agent Evaluation and Harness Engineering
POSITION SUMMARY
As an Applied Machine Learning Scientist, Agent Evaluation and Harness Engineering, you will lead applied research on evaluation, observability, stress-testing, and systematic improvement of AI agents. The role focuses on assessing agent performance and safety across long-horizon, multi-step tasks, and on building methods and tools to help organizations understand whether those systems are working, why they fail, and how to make them measurably better.
A core objective is developing adaptive evaluation approaches tailored to Canadian organizations, moving beyond static public benchmarks towards rigorous, organization-specific test environments of end-to-end agentic systems. Working alongside Vector researchers, research professionals, and external partners, the role balances high-quality applied research with the creation of practical technical systems that improve the reliability, safety, security, and effectiveness of deployed agents.
KEY RESPONSIBILITIES
Research and implement state-of-the-art methods for evaluating agents operating over long horizons, multiple tools, changing environments, and partially observable states;
Develop evaluations that assess complete agent trajectories, including planning quality, tool selection, intermediate decisions, state transitions, recovery behaviour, verification, termination decisions, resource consumption, and final outcomes;
Develop methods for creating organization-specific evaluations from production traces, human feedback, incidents, near misses, support interactions, domain-expert knowledge, and synthetic scenario generation;
Create techniques for converting discovered failures into durable regression evaluations that can be rerun across model, prompt, policy, tool, and harness changes;
Partner with Vector researchers, Applied ML Specialists, research professionals, and external collaborators – including Vector industry partners and members of the Canadian AI Safety Institute – to identify consequential agent use cases and create tools, reference agents, and evaluations required for trustworthy deployment;
Develop schemas and infrastructure for capturing structured traces of active agents;
Research representations of agent trajectories, such as event streams, causal graphs, tool-call graphs, state-transition graphs, and compact trajectory embeddings;
Develop approaches for identifying recurrent failure patterns and attributing outcomes to specific components or decisions within an agent system;
Build privacy-preserving and security-conscious methods for collecting and analyzing traces in sensitive organizational environments;
Research and build agent harnesses incorporating tools, memory, retrieval, sandboxes, permissions, validators, execution loops, recovery strategies, state management, and human approval mechanisms;
Develop automated or semi-automated methods for optimizing agent harnesses based on evaluation results and execution traces;
Develop safe mechanisms for agents to propose modifications to their own prompts, tools, policies, memory structures, workflow logic, or evaluation criteria while preserving auditability and human control;
Lead or contribute to peer-reviewed publications, technical reports, open-source software, benchmark releases, and reference implementations;
Contribute to training programs and technical workshops that help Vector partners and external stakeholders design, evaluate, debug, and govern agent systems;
Serve as a Vector expert on emerging methods in agent evaluation and harness engineering and connect external stakeholders with relevant members of the Vector research community; and,
Other related duties as assigned from time to time.
KEY SUCCESS MEASURES
Development of novel, scientifically rigorous evaluation methods for tool-using and long-horizon agents;
Release of novel evaluation tooling in collaboration with the Canadian AI Safety Institute;
Creation of adaptive evaluation systems that discover materially important failures beyond those captured by static benchmarks;
Adoption of evaluation artefacts by Vector partners and stakeholders across Canada;
High-quality technical outputs, including peer-reviewed publications, technical reports, open-source tools, benchmarks, datasets, and reference architectures;
Effective collaboration with internal research and engineering teams and with external partners across sectors; and,
Meaningful contribution to Vector’s technical leadership in agent evaluation, agent observability, harness engineering, and safe recursive improvement.
PROFILE OF THE IDEAL CANDIDATE
PhD in computer science, computer engineering, machine learning, or a related discipline, or equivalent demonstrated research or engineering experience;
Research expertise in one or more of: evaluation of AI agents or language-model systems; automated red-teaming; program synthesis or automated software improvement; AI safety, security, or robustness; multi-agent systems;
Strong ability to design controlled experiments and reason about confounding variables, stochasticity, statistical power, evaluator reliability, and reproducibility;
Experience evaluating systems whose behaviour unfolds across multiple steps, tool interactions, or environmental state changes;
Strong knowledge of Python and experience building high-quality research software;
Experience working with modern language models and tool-using agent architectures;
Understanding of the distinction between model evaluation and evaluation of the broader model–harness–environment system;
Familiarity with open-source machine-learning and agent frameworks such as PyTorch, JAX, Google ADK, LangGraph, the OpenAI Agents SDK, or comparable systems;
Comfortable working at the boundary between open-ended research and production-quality engineering; and,
Able to communicate complex findings clearly to technical researchers, engineering leaders, domain specialists, and senior organizational stakeholders.
TOTAL REWARDS: The expected salary for this position will be $125,800 - $157,300 per year, plus benefits if applicable. The final salary offer will reflect the successful candidate's experience, skills, and qualifications, in alignment with the Vector Institute's Compensation Policy and may differ from above.
The Vector Institute’s Total Rewards approach extends beyond traditional compensation and benefits. Full-time employees are eligible for a comprehensive suite of supports that recognize and value employees, including vacation time, floater days, GRRSP, a Health Spending Account, a Summer Hours program, and flexible work arrangements.
POSITION STATUS: This posting is for an existing vacancy.
USE OF ARTIFICIAL INTELLIGENCE: Vector may use both internal and external third party AI-based tools to assist in the screening of applications for this posting. Any data collected will be used solely for recruitment purposes and handled in accordance with Vector’s External Privacy Policy and Use of AI-Based Tools in Recruitment and Selection Policy.
INCLUSION AND EQUAL OPPORTUNITY EMPLOYMENT: Vector believes AI powers possibility by advancing cutting-edge research and translating it into real-world impact through collaboration with research, industry, and government. Vector is committed to fostering a diverse and inclusive culture that reflects its values.
The Vector Institute welcomes applications from all qualified candidates, including those who are Indigenous, 2SLGBTQIA+, racialized persons/visible minorities, women, and people with disabilities.
If you require an accommodation at any stage of the recruitment or selection process, please contact [email protected]. The Vector Institute team will be happy to work with you to ensure your experience is as inclusive and accessible as possible.
JOIN OUR COMMUNITY: Check out the Vector Institute’s Careers Page to explore open opportunities at Vector and Follow Vector on X, LinkedIn, and Bluesky to stay connected with the latest developments in Ontario's AI ecosystem and the Vector Institute.