Skip to project details
Oney Erge
Agent evaluationSimulation platform

AgentGymnasium

Bring AI agents and physics into the same experiment.

Give an agent a simulated world, explicit tools, and physical challenges, then measure what it builds, what happens, and how its next attempt changes.

AgentGymnasium tiny city simulation previewCITY_PULSE / SCORE 100%

What AgentGymnasium does.

Aim

An AI agent builds something in a simulated world, sees what happened, and gets another chance to improve it.

What it does

Validated tools produce engine-neutral episode traces. Named reward functions and paired model-by-seed trials make behavior replayable and statistically comparable.

Good at

  • 24 explicit tools
  • Explainable scores
  • Replayable trials

The AgentGymnasium flow.

  1. 01

    Configure

    Choose the world, task, protocol, model, tools, and resource bounds.

  2. 02

    Build

    Require every design mutation to pass one validated tool boundary.

  3. 03

    Simulate

    Run the candidate design and capture an engine-neutral episode trace.

  4. 04

    Explain

    Convert named trace metrics into a scorecard and improvement hint.

  5. 05

    Compare

    Replay paired runs with configuration, tokens, latency, and score deltas.

Choose a reference challenge, run one agent trial, then replay the trace and inspect the score.

Setup and examples

Explore another project.

View all projects