Practical evaluation guides
Added three connected guides: reading benchmark results carefully, measuring latency, throughput, and context, and building repeatable model comparisons. Expanded the site scope and shared methodology.
Updates
New pages and substantive revisions are recorded here. Minor copy and styling corrections are not listed.
Follow updates via Atom for new and materially revised notes.
Added three connected guides: reading benchmark results carefully, measuring latency, throughput, and context, and building repeatable model comparisons. Expanded the site scope and shared methodology.
Published the project overview, comparison methodology, and introductory notes on benchmark interpretation and inference trade-offs.