> ## Documentation Index
> Fetch the complete documentation index at: https://modal-computer-use.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Run benchmarks

> Run credential-free checks and authorized Modal benchmarks.

Start with a credential-free report.

## Run a local report

```bash theme={"system"}
uv sync --extra dev
uv run computer-use benchmark report --mock-local --iterations 5
```

Run the ordered batch comparison:

```bash theme={"system"}
uv run computer-use benchmark action-batch \
  --mock-local \
  --four-click-only \
  --warmup-iterations 1 \
  --iterations 5
```

These commands create no Modal resources.

## Run promotion gates

Each live script requires an explicit authorization flag, Modal credentials, a clean source revision, fixed placement and resources, retries set to zero, and a fresh output directory.

| Gate | Command |
| - | - |
| Raw binary SDK default | `uv run python scripts/run_optimized_default_promotion.py --help` |
| Computer Step | `uv run python scripts/run_step_promotion.py --help` |
| Input-work capacity | `uv run python scripts/run_input_capacity_gate.py --help` |
| Managed Image lifecycle | `uv run python scripts/run_image_lifecycle_benchmark.py --help` |

Read the [canonical benchmark procedure](https://github.com/ashtonchew/modal-computer-use/blob/b60c1cb7495200e36a738c0f6e07961b1d2db93c/docs/benchmarking.md) before a live run. It defines the exact flags, timer boundaries, cost ceiling, cleanup, sanitization, and promotion rules.

<Warning>
  Live commands create billable Modal resources. Record an approved cost ceiling. Poll each run to a terminal state. Check the provider console after cleanup.
</Warning>

## Keep measurements comparable

* Use one exact source commit and a clean worktree.
* Record requested and observed placement.
* Hold resources, Image, ingress, HTTP version, input backend, screenshot options, and warm capacity constant.
* Interleave matched arms according to the preregistered schedule.
* Keep failures. Use no replacement samples.
* Report cleanup failures and survivors.
* Separate cold allocation, startup, Function dispatch, borrow, and warm operations.

## Store evidence safely

Write raw output under ignored `benchmark-results/`. Remove endpoint URLs, resource IDs, tokens, screenshots, typed text, clipboard text, and raw failure content before promotion.

Publish only a sanitized artifact accepted by its repository validator. Keep dated reports and artifacts immutable.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.