Skip to content

feat: add native E2B backend for cloud-based agent evaluation - #11

Merged
Odysseusqs merged 17 commits into
mainfrom
feat/e2b-native
Sep 8, 2026
Merged

Odysseusqs merged 17 commits into
mainfrom
feat/e2b-native

Conversation

@Blizzard-cyber

Copy link
Copy Markdown
Collaborator

Summary

This PR adds native E2B support to SForge, allowing EdgeBench tasks to run in managed cloud Sandboxes without requiring a local Docker daemon or Kubernetes cluster.

Work, Judge, and Game environments are created from pre-built official E2B Templates, while SForge preserves the existing task definitions, iterative evaluation workflow,
scoring semantics, and result layout used by the Docker and Kubernetes backends.

What changed

E2B backend

  • Add an E2BBackend implementation for Sandbox creation, command execution, file transfer, service exposure, and cleanup.
  • Resolve official Work and Judge Templates from task image references through a configurable Template namespace.
  • Support Work, ephemeral Judge, and interactive Game Sandboxes.
  • Keep the Judge Server on a user-managed, publicly reachable host so the evaluation flow remains consistent with the existing backends.
  • Apply E2B-native network isolation for tasks that disable public internet access.

Long-running evaluation reliability

  • Add Sandbox TTL renewal for long-running tasks.
  • Use provider-side on_timeout=kill, explicit deletion, and bounded cleanup fallback.
  • Harden command timeout handling, file transfers, and run-state release.
  • Support independent concurrency control for Work and Judge/Game Sandboxes.
  • Preserve auto-eval, stop hooks, auto-resume, submission cooldowns, and best-result selection on the E2B path.

Agent and CLI support

  • Add --backend e2b and the related configuration options.
  • Add versioned Claude Code and Codex variants.
  • Add an OpenCode agent integration and visualizer parsing support.
  • Add a unified --effort option translated to each supported agent.
  • Preserve existing Docker and Kubernetes behavior.

Documentation

  • Add English and Chinese E2B setup guides.
  • Add a complete single-task E2B example covering:
    • installation and credentials;
    • official Template resolution;
    • Judge Server deployment;
    • third-party model endpoints;
    • network isolation;
    • plan limits and run sizing;
    • troubleshooting.
  • Update CLI, environment variable, agent, experiment, and backend documentation.

Architecture

The evaluation flow remains consistent with the existing SForge model:

  1. sforge run creates an E2B Work Sandbox from the task's published Work Template.
  2. The agent works in the isolated Work environment and submits an archive to the external Judge Server.
  3. The Judge Server creates a fresh Judge Sandbox from the corresponding Judge Template for each evaluation.
  4. Interactive tasks use dedicated Game Sandboxes exposed through the E2B service gateway.
  5. SForge collects the run history, selects the best completed result, and cleans up all managed Sandboxes.

No managed Controller Sandbox is required.

Compatibility

  • Docker remains the default backend.
  • Kubernetes behavior is preserved.
  • Existing task definitions, parsers, grading, selection, and result formats are unchanged.
  • E2B support is installed through the optional e2b dependency extra.
  • Official Templates are pre-built and published by maintainers; users do not need to build Templates locally.

Validation

The E2B path was validated with the complete 51-task open-source EdgeBench suite using the official leaderboard resource specifications.

  • 51/51 tasks created Work Sandboxes and produced final result files.
  • 50/51 tasks completed the full 12-hour evaluation budget.
  • Work, Judge, and Game execution paths were covered.
  • The public-leaderboard aggregation for completed tasks produced an @12h score of 32.1.
  • Lifecycle cleanup, TTL renewal, auto-eval, agent resume, and concurrent task execution were exercised during the long-running evaluation.

Additional checks:

  • git diff --check
  • Python package build
  • Documentation build
  • CLI and E2B optional dependency installation
  • Single-task E2B end-to-end workflow
  • Multi-task long-running evaluation

User-facing example

pip install "sforge[e2b]"

export E2B_API_KEY="..."
export SFORGE_E2B_TEMPLATE_NAMESPACE="edgebench"

# On a publicly reachable host
sforge serve --host 0.0.0.0 --port 8080

# On the run host
sforge run \
  --backend e2b \
  --task ad_placement_optimization \
  --agent claude-code \
  --judge-url http://YOUR_PUBLIC_HOST:8080

Copilot AI lite review requested due to automatic review settings September 8, 2026 08:10

This comment was marked as off-topic.

@Odysseusqs
Odysseusqs merged commit ce685f2 into main Sep 8, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants