> ## Documentation Index
> Fetch the complete documentation index at: https://modal-computer-use.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Run benchmarks

> Run credential-free and Modal benchmarks with the repository benchmark commands.

Start with a credential-free run. Use a live Modal run only when you need provider evidence.

## Run a local report

Install the development environment:

```bash theme={"system"}
uv sync --extra dev
```

Run the release report against the in-process mock daemon:

```bash theme={"system"}
uv run computer-use benchmark report --mock-local --iterations 5
```

Run the action-batch comparison:

```bash theme={"system"}
uv run computer-use benchmark action-batch \
  --mock-local \
  --four-click-only \
  --warmup-iterations 1 \
  --iterations 5
```

These commands do not create Modal resources.

## Run against an existing daemon

Replace `--mock-local` with a daemon URL. Add a token when the daemon requires one.

```bash theme={"system"}
uv run computer-use benchmark report \
  --base-url http://127.0.0.1:8080 \
  --token dev \
  --iterations 5 \
  --output benchmark-results/report.json
```

The report command does not create a Sandbox. The optional `--include-sandbox-exec` mode attaches to an existing Sandbox when you provide its ID.

## Run a live Modal SDK benchmark

Install the Modal extra and authenticate:

```bash theme={"system"}
uv sync --extra modal
uv run modal setup
```

Run the SDK benchmark:

```bash theme={"system"}
uv run computer-use benchmark sdk \
  --create-modal-sandbox \
  --surfaces daemon-http \
  --browser chromium \
  --resource-profile browser \
  --iterations 30 \
  --output benchmark-results/modal-sdk.json
```

This command creates a billable Sandbox. It waits for readiness. It attempts termination and detachment after the run.

<Warning>
  Live commands can create billable resources. Record a cost ceiling before the run. Inspect the provider console after cleanup.
</Warning>

## Run the current Modal optimized path

Use a clean commit. Publish revision-addressed images before you run the benchmark.

```bash theme={"system"}
evidence_harness_sha="$(git rev-parse HEAD)"
test -z "$(git status --porcelain)"

uv run python scripts/publish_modal_images.py --revision "$evidence_harness_sha"

uv run computer-use benchmark modal-optimized-provider \
  --modal-region us-west-2 \
  --image-revision "$evidence_harness_sha" \
  --modal-cpu 1 \
  --modal-memory-mib 2048 \
  --runner-cpu 1 \
  --runner-memory-mib 2048 \
  --browser chromium \
  --iterations 30 \
  --warmup-iterations 1 \
  --output benchmark-results/modal-optimized-provider.json
```

The command measures fresh lifecycle samples and six warm-operation cases. It fails when a required sample, placement check, or cleanup gate fails.

## Run a focused experiment

Use the command that owns the question:

| Question | Command |
| - | - |
| Does one request reduce four-click overhead? | `computer-use benchmark modal-action-batching-ab` |
| How does Modal placement affect the caller path? | `computer-use benchmark modal-region-ab` |
| How does a colocated runner behave? | `computer-use benchmark modal-colocated-client` |
| How do provider-default SDK paths compare? | `computer-use benchmark compare` |
| How does ingress affect the optimized path? | `computer-use benchmark modal-optimized-ingress-ab` |

Read the exact flags and gates in the [canonical benchmark procedure](https://github.com/ashtonchew/modal-computer-use/blob/4425402dbc681133252dbc54d971ea4c95bc0ffc/docs/benchmarking.md) before you run a publishable experiment.

## Keep raw output private

Write raw output under `benchmark-results/`. Do not write it at the repository root. Do not commit credentials, endpoint URLs, resource identifiers, screenshots, typed text, clipboard text, or raw failure content.

Promote only a sanitized artifact that passes its repository validator. The [benchmark data policy](https://github.com/ashtonchew/modal-computer-use/blob/4425402dbc681133252dbc54d971ea4c95bc0ffc/benchmark-data/README.md) defines that promotion boundary.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.