Interactive agent evaluation

Can AI build simulations that teach?

121 simulations test behavior, scientific fidelity, and classroom usefulness.

Live benchmark surface
Energy conservation in motionOpen lab ->
121Interactive simulations
242Benchmark variants
4Benchmark models
0Verified outputs

Generation coverage

Live results from published, browser-verified candidates.

ModelVerifiedCoverageAverage check scoreMedian timeStatus

Coverage is measured against 242 tasks per model. Only candidates passing all automated browser checks are published.

4 Signals separate working science from a convincing animation.
35%

Behavior

Controls produce continuous, visible change.

30%

Science

Relationships remain physically defensible.

20%

Teaching

Evidence supports a classroom inquiry.

15%

Robustness

Repeated use and reset states hold up.

Benchmark task catalog

Search every simulation, inspect its tasks, and run generated model outputs.

PhET Benchmark Interactive STEM evaluation prototype