Tools I build to make
quality measurable.
Applied quality engineering — open-source reporting, multi-agent test automation, and LLM evaluation pipelines running in production.
playwright-spec-doc-reporter
A Playwright reporter with BDD annotations, AI failure analysis, and interactive HTML dashboards. Generates living specification documents from your test runs. 100+ weekly npm downloads.
Multi-Agent QA Framework
Four specialised agents with a DOM Memory architecture for persistent selector knowledge — spec parsing, test generation, execution, and triage working together.
DeepEval Production Pipelines
LLM evaluation for RAG systems — faithfulness, relevancy, and hallucination detection wired directly into Azure DevOps CI/CD.
QA AI Toolkit
Open-source QA skills, agents, and MCP servers organized by objective — installable into Claude Code, Cursor, or Windsurf with one command.
Frameworks & resources
QE Maturity Framework
A five-level assessment model for benchmarking quality engineering organisations across automation, DORA metrics, and AI adoption.
Prompt Engineering for QE
A practical field guide to using LLMs across the testing lifecycle — from test design to failure triage.
QA Transformation Roadmap
A step-by-step roadmap for taking a QA organisation from manual gatekeeping to embedded quality engineering.
Building something?
If you're working on quality engineering, AI testing, or reliability at scale, I'd like to hear about it.