Model map

Every model at every reasoning effort, ranked on what one finished task costs rather than on price per token.

effort
paid via
vendor
limits intel ≥ $/task ≤ out tok ≤
columns

Score against cost

Up and left is better. The dashed line is the frontier: nothing in view beats those points on both axes.

What each effort step buys

Where a line flattens, the next effort level is paying for very little. Follows the filters; click a name to hide it.

Pick by the job

Hand-written picks. The numbers under each are read live, so a stale claim shows itself.

    Rows to avoid

    Operating points that lose to something else on every axis that matters.

      What a subscription buys

      Plan quota converted into Index tasks. Read the ratio as a lie detector, not a discount: anything past about 10x means the quota estimate is too generous.

      Which model people prefer the look of

      Design Arena's website board, top twelve. Human votes on a single generated design, which the Intelligence Index does not measure at all.