Skip to content
#

groundedness

Here are 25 public repositories matching this topic...

pip install gauntlet-evals · v0.1.0. Merge-blocking evaluation gates for generative AI features: YAML suites run against any HTTP endpoint or Python callable, fail the build on a miss, and emit both a diffable JSON pack and a reviewer document cross-referenced to California's published GenAI risk framework. Aligned to, never approved by.

  • Updated Sep 12, 2026
  • Python

v0.2.0. Fail-closed evaluation harness for government-facing chat systems: reproducible, provenance-stamped audit verdicts, byte-identical across Python 3.11 to 3.14, with no third-party dependencies. A silent or unreadable target scores zero rather than passing by absence. Two public projects of my own pin it by exact commit.

  • Updated Sep 10, 2026
  • Python
sprout

In-build reference implementation: an offline-first plant-care assistant and public evaluation harness with cited-corpus answers, calibrated abstention, toxicity guardrails, EN/ES parity, photo plant ID, and local reminders.

  • Updated Sep 10, 2026
  • Python
fare-policy-assistant

Beta. Reduced-fare policy assistant citing dated corpus passages in English and Spanish; the bilingual-parity gate is currently failing (see EVALS.md). Corpus of eighteen California transit agencies, public 385-case evaluation harness. Deployed demo serves five agencies; published evidence run lags the repository.

  • Updated Sep 9, 2026
  • HTML

Add this topic to your repo

To associate your repository with the groundedness topic, visit your repo's landing page and select "manage topics."

Learn more