Skip to content

Latest commit

 

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Codex Auditor — a watchdog that scores AI-agent-authored commits for intent accuracy and concurrency safety before you merge them.

The Problem

Agentic coding tools (Codex, and agents like it) now author production commits autonomously. The bottleneck has shifted from "can the agent write code" to "can I trust the code it wrote without reading every line myself." Two specific trust gaps matter most: intent drift (does the diff actually do what the commit message claims?) and concurrency correctness (race conditions and deadlocks that are notoriously hard to catch). Codex Auditor closes both gaps: it is a watchdog that runs on every commit/PR, produces a human-readable verdict, and assigns a trust score before a human has to decide whether to merge.

What It Does

  • Checks whether a diff's actual behavior matches its stated commit message intent
  • Detects three specific, well-defined concurrency bug patterns (missing-await, lock-order cycles, unsynchronized shared state)
  • Auto-generates and runs property-based tests for touched functions
  • Produces a transparent 0-100 trust score with a full breakdown

Setup Instructions

git clone <repo-url>
cd CodexAuditor
uv sync
export OPENAI_API_KEY=<your-key>

How To Run (Sample Data)

uv run codex-auditor audit \
  --diff demo_repo/fixtures/mismatched_intent.diff \
  --message "Fix off-by-one error in pagination" \
  --repo-root demo_repo

How Codex Accelerated Our Workflow

  • Generated the foundational CLI orchestration scaffolding (apps/cli/src/codex_auditor_cli/main.py and audit.py) using click.
  • Drafted the unidiff boilerplate and AST visitor classes for parsing unified diffs (packages/core/src/codex_auditor_core/diff_ingest.py).
  • Auto-generated the Jinja2 template structure (report.md.jinja2) and logic in render.py.
  • Wrote extensive test scaffolding and pytest configurations across packages/core and apps/cli, accelerating our TDD loop.

Where GPT-5.6 Is Used At Runtime

  • Intent-vs-implementation reasoning (concurrency/explain.py:check_intent)
  • Concurrency finding explanation / false-positive filtering (concurrency/explain.py:explain_findings)
  • Property test generation (property_tests.py)

Key Decisions

  • Scope limitation: We scoped concurrency detection to three specific patterns (missing-await, lock-order cycle, unsynchronized shared state) rather than building a general race detector to maximize feasibility and strictly control false-positives.
  • LLM acting as a Filter: GPT-5.6 is used as a filter/explainer on top of deterministic static analysis rather than the sole detection mechanism to guarantee reliability and explainability.
  • Sandboxing generated tests: Generated property tests are run in a sandboxed subprocess with a strict timeout (10s) to guarantee safety and prevent an LLM hallucination from hanging the CI pipeline indefinitely.

Known Limitations

  • Heuristic shared-state scanner may miss bugs behind indirection, or (rarely, after the LLM filter pass) flag safe patterns.
  • Single-language (Python) MVP; cross-file coroutine resolution is best-effort only.
  • Not a substitute for code review or formal verification — it serves as a triage signal.

Architecture

                    ┌─────────────────────┐
   diff + commit    │                     │
   message (input)  │   Codex Auditor     │
  ───────────────►  │      Engine         │ ──► Trust Report (Markdown + JSON)
                    │                     │
                    └─────────┬───────────┘
                              │
        ┌─────────────────────┼─────────────────────┐
        ▼                     ▼                     ▼
 ┌───────────────┐   ┌─────────────────┐   ┌──────────────────┐
 │ Intent Checker │   │ Concurrency     │   │ Property Test    │
 │ (GPT-5.6 call) │   │ Pattern Scanner │   │ Generator +      │
 │                │   │ (AST + graph)   │   │ Runner           │
 └───────────────┘   └─────────────────┘   └──────────────────┘
                              │
                     ┌─────────────────┐
                     │ Trust Score      │
                     │ Aggregator       │
                     └─────────────────┘

Testing

  • Fast unit tests: run uv run pytest (skips the network requests by mocking)
  • Real API integration tests: run uv run pytest -m integration (slower, hits the actual OpenAI APIs)

GitHub Action

Codex Auditor can be used directly as a GitHub action in your repository to seamlessly scan PRs! Example .github/workflows/audit.yml to set this up:

name: Codex Audit
on: [pull_request]
jobs:
  audit:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - name: Run Codex Auditor
        uses: ./apps/github-action
        with:
          github-token: ${{ secrets.GITHUB_TOKEN }}

Testing the Action without Rebuilding: You can test the GitHub action easily without needing to modify your codebase by targeting the included demo_repo/ sample data folder.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages