Codex Auditor — a watchdog that scores AI-agent-authored commits for intent accuracy and concurrency safety before you merge them.
Agentic coding tools (Codex, and agents like it) now author production commits autonomously. The bottleneck has shifted from "can the agent write code" to "can I trust the code it wrote without reading every line myself." Two specific trust gaps matter most: intent drift (does the diff actually do what the commit message claims?) and concurrency correctness (race conditions and deadlocks that are notoriously hard to catch). Codex Auditor closes both gaps: it is a watchdog that runs on every commit/PR, produces a human-readable verdict, and assigns a trust score before a human has to decide whether to merge.
- Checks whether a diff's actual behavior matches its stated commit message intent
- Detects three specific, well-defined concurrency bug patterns (missing-await, lock-order cycles, unsynchronized shared state)
- Auto-generates and runs property-based tests for touched functions
- Produces a transparent 0-100 trust score with a full breakdown
git clone <repo-url>
cd CodexAuditor
uv sync
export OPENAI_API_KEY=<your-key>uv run codex-auditor audit \
--diff demo_repo/fixtures/mismatched_intent.diff \
--message "Fix off-by-one error in pagination" \
--repo-root demo_repo- Generated the foundational CLI orchestration scaffolding (
apps/cli/src/codex_auditor_cli/main.pyandaudit.py) usingclick. - Drafted the
unidiffboilerplate and AST visitor classes for parsing unified diffs (packages/core/src/codex_auditor_core/diff_ingest.py). - Auto-generated the Jinja2 template structure (
report.md.jinja2) and logic inrender.py. - Wrote extensive test scaffolding and pytest configurations across
packages/coreandapps/cli, accelerating our TDD loop.
- Intent-vs-implementation reasoning (
concurrency/explain.py:check_intent) - Concurrency finding explanation / false-positive filtering (
concurrency/explain.py:explain_findings) - Property test generation (
property_tests.py)
- Scope limitation: We scoped concurrency detection to three specific patterns (missing-await, lock-order cycle, unsynchronized shared state) rather than building a general race detector to maximize feasibility and strictly control false-positives.
- LLM acting as a Filter: GPT-5.6 is used as a filter/explainer on top of deterministic static analysis rather than the sole detection mechanism to guarantee reliability and explainability.
- Sandboxing generated tests: Generated property tests are run in a sandboxed subprocess with a strict timeout (10s) to guarantee safety and prevent an LLM hallucination from hanging the CI pipeline indefinitely.
- Heuristic shared-state scanner may miss bugs behind indirection, or (rarely, after the LLM filter pass) flag safe patterns.
- Single-language (Python) MVP; cross-file coroutine resolution is best-effort only.
- Not a substitute for code review or formal verification — it serves as a triage signal.
┌─────────────────────┐
diff + commit │ │
message (input) │ Codex Auditor │
───────────────► │ Engine │ ──► Trust Report (Markdown + JSON)
│ │
└─────────┬───────────┘
│
┌─────────────────────┼─────────────────────┐
▼ ▼ ▼
┌───────────────┐ ┌─────────────────┐ ┌──────────────────┐
│ Intent Checker │ │ Concurrency │ │ Property Test │
│ (GPT-5.6 call) │ │ Pattern Scanner │ │ Generator + │
│ │ │ (AST + graph) │ │ Runner │
└───────────────┘ └─────────────────┘ └──────────────────┘
│
┌─────────────────┐
│ Trust Score │
│ Aggregator │
└─────────────────┘
- Fast unit tests: run
uv run pytest(skips the network requests by mocking) - Real API integration tests: run
uv run pytest -m integration(slower, hits the actual OpenAI APIs)
Codex Auditor can be used directly as a GitHub action in your repository to seamlessly scan PRs!
Example .github/workflows/audit.yml to set this up:
name: Codex Audit
on: [pull_request]
jobs:
audit:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Run Codex Auditor
uses: ./apps/github-action
with:
github-token: ${{ secrets.GITHUB_TOKEN }}Testing the Action without Rebuilding:
You can test the GitHub action easily without needing to modify your codebase by targeting the included demo_repo/ sample data folder.