Agent PM: OpenAI Agents-powered product management orchestrator with automated PRDs, tickets, and comms.
-
Updated
Sep 2, 2026 - TypeScript
Agent PM: OpenAI Agents-powered product management orchestrator with automated PRDs, tickets, and comms.
Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild
Hermes plugin: route LLM calls through the EvalOps gateway and report agent registration + tool spans to the EvalOps platform.
RAG pipeline evaluation for .NET — LLM-as-judge scoring for Faithfulness, Answer Relevance, Context Precision, and Context Recall. The evaluation layer your .NET RAG pipeline is missing.
CI-ready LLM regression harness for document-grounded RAG assistants: hard checks, pairwise judging, bootstrap CIs, and artifact reports.
AI Agent EvalOps dashboard for regression testing, drift detection, run comparison, and quality analysis.
Evidence-aware benchmark and PR advisory prototype for AI coding-agent context safety.
To associate your repository with the evalops topic, visit your repo's landing page and select "manage topics."