Skip to main content
Measure startup and warm operations as separate paths. A change that reduces startup time might not change action latency.

Keep one session for the trajectory

Create the desktop once. Keep one connection open across the repeated observe, model, and action loop. Use AsyncComputerSandbox when your application already uses asyncio. Native async keeps the event loop responsive during Modal and daemon I/O. It does not make Sandbox allocation faster. For a deployed Modal Function, enter one borrow_async() context around the complete trajectory. Do not borrow once per action.

Place the caller after measurement

Run the Modal region benchmark from the real caller. Apply the measured region selector to the desktop and to any deployed Function. A shared region request is a scheduling request. It does not prove a shared host or availability zone. Keep public gateways, brokers, and durable stores outside the screenshot and action path. They can still own admission, recovery, and audit records.

Reduce work per turn

Batching saves network round trips. The daemon validates the complete batch before execution. It stops on the first error by default. An immediate post-action screenshot does not prove that an application is ready. Use an application predicate when the next step needs a semantic state.

Prepare the browser and image

Use ResourceConfig(profile="browser") with BrowserConfig for browser work. Prewarm the browser only after browser startup appears in the measured path. Add CPU, memory, or GPU only when measurements show sustained demand. A GPU allocation does not prove that an X11 browser uses hardware rendering. Use named image revisions for stable system and Python dependencies. Keep changing application files in later image layers.

Use warm capacity deliberately

Positive Function capacity reduces Function cold starts. A desktop warm pool reduces request-to-ready time.
Both forms of warm capacity add idle cost. Record the pool hit rate, cold fallback rate, remaining lifetime, and reconciled billing data.

Record a complete measurement

Record these facts with each result:
  • Caller location and topology.
  • Requested and observed placement.
  • Ingress, image revision, resources, and browser setup.
  • Cold or warm state.
  • Exact timer start and end.
  • Raw samples, failures, cleanup, and cost state.
Use at least 30 measured samples when you report p95. Keep the raw samples. Record the clean evidence revision. The July 30, 2026 report contains 30 successful samples per cell. Its optimized path used a synchronous Modal Function caller and tuned daemon settings. Treat those numbers as dated evidence for that recorded topology. Do not treat them as a promise for another workload. Use the benchmarking guide for commands, evidence status, statistics, cost accounting, and publication rules.