Product Quality Engineer (AI & Agentic Systems)
About This Role
Collective OS needs a product-minded engineer who can turn business intent into executable evidence. You will define what good behavior means, create realistic personas and scenarios, build automation, and evaluate agent outputs where exact-text assertions are not sufficient. This is not a downstream manual-testing function. It is a hands-on engineering role that reports into the CTO organization and works as an embedded partner to Product, Design, and Engineering from discovery through release and production learning.
What You Will Own
- Translate product decisions, business rules, and agent specifications into clear behavioral contracts and acceptance criteria.
- Create representative firm and operator personas, together with their contacts, relationship networks, target organizations, integrations, consent states, and data conditions.
- Build and maintain critical-path automation across browser, API, worker, integration, and data layers, including Playwright-based journeys.
- Design AI evaluation suites that combine deterministic checks, structured rubrics, repeated trials, calibrated graders, and periodic human review.
- Test signal detection, relationship-sensitive routing, scoring, recommendations, grounding, memory, permissions, and graceful failure across realistic and adversarial conditions.
- Define release evidence and production quality signals; turn operator corrections and escaped failures into durable regression cases.
- Improve the shared simulation, scenario, trace, and evaluation tooling used by Product and Engineering.
What Success Looks Like
Quality becomes an observable product discipline rather than a final approval gate.
- High-value journeys and high-risk behavioral boundaries are covered by reusable personas, scenarios, and gold datasets.
- Model, prompt, routing, scoring, and data-source changes are evaluated through measurable evidence rather than isolated demos or intuition.
- Failures are diagnosable across ingestion, detection, routing, agent decision, persistence, and presentation.
- Release decisions include concise evidence about coverage, quality deltas, known risks, and explicit exceptions.
- Production feedback becomes a fast learning loop for the product and a permanent part of the regression system.
What We Are Looking For
- Strong experience in software quality, test engineering, product engineering, developer productivity, or AI evaluation for complex production systems.
- Hands-on experience evaluating LLM applications or agentic systems, including tool use, retrieval, grounding, memory, and non-deterministic outputs.
- Strong TypeScript and/or Python skills and experience building maintainable test or evaluation infrastructure.
- Practical depth in Playwright or equivalent browser automation, plus API, integration, contract, worker, and data-pipeline testing.
- Ability to design gold datasets, rubrics, sampling strategies, automated graders, and human calibration workflows.
- Strong product and business judgment: you can decide whether an output is useful to an agency owner, not only whether it is syntactically valid.
- Comfort diagnosing asynchronous and distributed systems with retries, partial failure, eventual consistency, and multiple data stores.
- Clear cross-functional communication and the ability to create alignment in ambiguous, fast-moving product work.
Helpful Experience
- B2B SaaS, sales or relationship intelligence, professional services, graph analytics, recommendation systems, or calibrated scoring.
- Establishing a quality or evaluation discipline in an early-stage company and mentoring others in quality practices.
- Model and prompt versioning, offline and online evaluation, observability, and cost-quality tradeoffs.
FIRST 90 DAYS
Map the current quality surface and baseline the highest-risk flows; establish the first canonical persona,
scenario, and gold-dataset library; introduce repeatable AI evaluations and release evidence; define the
production feedback-to-regression loop and next-stage roadmap.
Pay: $30.00-$40.00 per hour
Benefits:
- Casual dress
- Company events
- Work from home
Work Location: Remote