llm-architectures-explained

/compare

Compare and calculate

Pick up to four models. The numbers come from the cost model, run in your browser on each model’s sourced configuration: memory for the weights and the KV cache, and the arithmetic and memory traffic of prefill and of each decode step. Closed models whose sizes are not disclosed stay blank unless you include reported estimates, which are labelled wherever they appear.

Loading the comparison…