Warm-operation benchmarks, 30 July 2026
On 30 July 2026, the benchmark measured six warm operations after the desktop and client connection were ready. Each table reports 30 successful samples for each path. The timer included transport, authentication, request handling, execution, and response collection. It excluded desktop creation and cleanup. The provider paths differed in caller placement, region, resources, screenshot format, and request shape. The tables report those recorded configurations. Each p50 ratio compares the complete recorded path with Modal optimized for the same operation. The benchmark measured screenshots and clicks as separate operations. The complete action-to-frame section reports a later fused Modal Step measurement.Full screenshot
Each path returned its provider-native full-screen image.One click
Each path sent one click to the ready desktop.Four ordered clicks
Each path sent four clicks in order.Type 100 characters
Each path typed a 100-character payload.Type 1,000 characters
Each path typed a 1,000-character payload.Non-login shell command
Each path ran the same logical non-login shell command.
Evidence: warm-operation report, Modal optimized artifact, and provider-default artifact
Warm-operation path configuration and measurement details
Warm-operation path configuration and measurement details
- Modal optimized: One Modal Function and its targets requested
us-west-2. The Function and target each used one CPU and 2 GiB. The path used attested-tunnel ingress, HTTP/1.1, XTest input, zero input pacing, and an isolated asyncio subprocess backend. - Daytona default: An external caller used Daytona 0.175.0 and its provider-default computer-use path.
- E2B default: An external caller used E2B Desktop 2.3.1 and its provider-default computer-use path.
- Modal simple: An external caller used the public Modal Computer Use SDK, standard resources, default input pacing, and the attested-tunnel daemon path.
- Tzafon default: An external caller used Tzafon 2.44.1 and its provider-default computer-use path.
sh -c "printf '42\n'" or its provider equivalent. Every successful sample returned exit code 0 and exact stdout "42\n".The optimized harness used synchronous ComputerSandbox and DaemonClient objects inside the placed Modal Function. It reused one warm client and connection. This benchmark predates the native async owner and daemon client APIs.Modal product benchmarks, 8 August 2026
The 8 August results cover implementation choices and capacity within Modal Computer Use. Each came from a separate benchmark run.Computer Step
On 8 August 2026, the Computer Step benchmark compared two ways to execute one ordered action batch and capture its immediate screenshot within the same topology.
The run retained 100 interleaved pairs per arm. It recorded zero failures and zero cleanup survivors. The harness made zero retries and used zero replacement samples.
Evidence: Computer Step report
The 11 August complete action-to-frame benchmark measured
computer.step() at 43.13 ms p50 and 46.35 ms p95. The same-topology result above came from 8 August.
An immediate frame ended each timer. First visual change and application readiness use separate signals.
Raw binary screenshot default
The optimized-default gate measured a full semantic screenshot followed by one pointer move.
Both arms retained 30 interleaved samples.
Evidence: Optimized-default report
The timer covered two SDK operations. The Computer Step section covers the fused route.
Input-work capacity
The minimum tested runtime completed three independent mixed XTest capacity gates.
Each run completed 80 batches of 48 ordered actions. The product default is a 100-token refill with a 400-token burst.
Evidence: Input-capacity report
Managed Image lifecycle
All 60 measured lifecycles for the managed standard Image completed, and cleanup succeeded each time.
The paired mean confidence interval crossed zero because both arms had large tail outliers. The result supports offering the managed standard Image as an opt-in choice. A universal default requires separate availability, rollback, and variant evidence.
Evidence: Image lifecycle report
Complete action-to-frame benchmarks, 11 August 2026
On 11 August 2026, the benchmark measured four public SDK paths. Each path sent one left click at(512, 384) and then captured the next validated full screenshot. The benchmark used two warmups and collected 100 measured samples from each path. The timer started immediately before ordered action dispatch. It ended after the next full screenshot was decoded and validated.
Configuration and measurement details
Configuration and measurement details
The recorded CPU and memory values describe each desktop target. Modal used the same resource shape for its placed caller and target. The benchmark did not record external caller capacity.
- Modal Computer Use /
computer.step():modal-computer-use2.0.0. An application-owned Modal Function called the target. Both resources requested and observedus-west-2. The target and caller each used one physical CPU and 2048 MiB. The screenshot was a PNG with unrecorded dimensions and a hidden cursor. One SDK call produced one transport request. The SDK did not retry mutations. - Daytona: Daytona 0.175.0. An external caller used the provider defaults for placement and region. The target used one physical CPU and 1024 MiB. The screenshot was a 1024 x 768 PNG with an unreported cursor setting. Two SDK calls produced two transport requests. The provider retry policy applied.
- E2B: E2B Desktop 2.4.2. An external caller used the provider defaults for placement and region. The target used two physical CPUs and 1024 MiB. The screenshot was a 1024 x 768 PNG with an unreported cursor setting. Two SDK calls produced three transport requests. The provider retry policy applied.
- Tzafon: Tzafon 2.44.1. An external caller used the provider defaults for placement and region. The provider did not disclose target resources. The screenshot was a 1280 x 720 JPEG with an unreported cursor setting. Two SDK calls produced two transport requests. The provider retry policy applied.

