Benchmark methodology
Official snapshots are evidence, not a single ranking. Cost, speed, and visual preference stay separate.
Prompts
Each challenge uses one canonical English prompt. Localized UI copy never changes the generation prompt.
Inputs
Editing and image-to-video challenges publish the exact reference, first-frame, or last-frame assets used in the run.
Output count
Each model produces one output per published run. Additional samples are not mixed into the snapshot.
Dimensions
Shared aspect ratio or resolution is compiled independently per model. Unreviewed nearest-value fallbacks are not used.
Provider settings
Transport fields such as sync mode and response format are server-owned. Quality, seed, and guidance stay model-specific.
Retries
Official retries are operator-initiated. A failed model does not rerun siblings.
Timing
Generation time is measured from upstream task creation to a terminal result.
Cost
Displayed prices are run snapshots. Current catalog prices are labeled separately when they differ.
Failures
Unsupported, incompatible, refused, empty, and failed states remain in the grid instead of being filtered out.
Run versions
Publishing points a challenge at one complete run. Previous runs stay private.
Limitations
Benchmarks do not claim equal seeds, quality labels, or prompt-enhancement flags across providers.