Framework for evaluating and improving agents
-
Updated
Sep 11, 2026 - Python
Framework for evaluating and improving agents
A curated, non-BS library of the best resources for building and evaluating AI agents — papers, blogs, talks, tools, benchmarks. Maintained by BenchFlow.
HuggingEnvs — RL Environments 101: building and scaling RL environments in the age of LLMs
A Universal Platform for Training and Evaluation of Mobile Interaction
A graphical interface for reinforcement learning and gym-based environments.
Interoperating between (Deep) Reiforcement Learning libraries
Gymnasium-style API standard for RL environment creation in JAX
Workspace manager for coding agents. Interactively solve and develop Harbor tasks.
An agentic auditor for RL environments and eval suites. Emits an Environment Card: every claim tied to a probe result, and every probe that could not run, with the reason.
Create new gridworld gym environments easily
Awesome Environment Scaling for AI Agents — a periodically updated survey and curated list of papers, projects, RL environments, world models, agent sandboxes, and open research problems
Turn any real software into a replayable RL environment for training AI agents — deterministic replay, verifiable rewards, TRL & verifiers adapters.
Adversarial QA for LLM-RL environments: find out what reward an empty answer earns. Model-free, zero API cost.
A configurable 3D jump-chess framework with A/B geometries, multi-player rule switches, and RL/self-play training support.
A lightweight, open-source framework that turns historical GitHub pull requests into reproducible, verifiable software-engineering tasks for training and evaluating coding agents.
Agent-evaluation environments: planted-truth worlds, ungameable graders, calibrated difficulty
A Harbor-format RL environment for pass@k estimation, with a verifier built to resist reward hacking. 17 tests, 10 adversarial solutions rejected, amd64 CI.
Comprehensive AI agent evaluation platform — searchable benchmark catalog, comparison matrices, automated scanner, interactive dashboards, and community-curated best practices for LLM evaluation.
A playbook for the execution and sandboxing layer beneath agent / RL environments on on-prem hardware — isolation tiers, snapshot and branch, run where the author could run it and marked where he could not.
To associate your repository with the rl-environments topic, visit your repo's landing page and select "manage topics."