Benchmark methodology

Official snapshots are evidence, not a single ranking. Cost, speed, and visual preference stay separate.

Prompts

Each challenge uses one canonical English prompt. Localized UI copy never changes the generation prompt.

Inputs

Editing and image-to-video challenges publish the exact reference, first-frame, or last-frame assets used in the run.

Output count

Each model produces one output per published run. Additional samples are not mixed into the snapshot.

Dimensions

Shared aspect ratio or resolution is compiled independently per model. Unreviewed nearest-value fallbacks are not used.

Provider settings

Transport fields such as sync mode and response format are server-owned. Quality, seed, and guidance stay model-specific.

Retries

Official retries are operator-initiated. A failed model does not rerun siblings.

Timing

Generation time is measured from upstream task creation to a terminal result.

Cost

Displayed prices are run snapshots. Current catalog prices are labeled separately when they differ.

Failures

Unsupported, incompatible, refused, empty, and failed states remain in the grid instead of being filtered out.

Run versions

Publishing points a challenge at one complete run. Previous runs stay private.

Limitations

Benchmarks do not claim equal seeds, quality labels, or prompt-enhancement flags across providers.