Every model at every reasoning effort, ranked on what one finished task costs rather than on price per token.
Up and left is better. The dashed line is the frontier: nothing in view beats those points on both axes.
Where a line flattens, the next effort level is paying for very little. Follows the filters; click a name to hide it.
Hand-written picks. The numbers under each are read live, so a stale claim shows itself.
Operating points that lose to something else on every axis that matters.
Plan quota converted into Index tasks. Read the ratio as a lie detector, not a discount: anything past about 10x means the quota estimate is too generous.
Design Arena's website board, top twelve. Human votes on a single generated design, which the Intelligence Index does not measure at all.