I think for the same model wall time is probably a more intuitive metric; at the...

nomel · 2026-04-07T18:56:09 1775588169

As a customer, it's nice that I can quantize and count the units of cost in an understandable way.

For Anthropic, as a business bleeding money, it's probably nice to have value-based pricing, for the tokens, so innovation (like computation efficiency improvements) can result in some extra margin. If they exposed the more direct computation cost, they could never financially benefit from any improved efficiency, including faster hardware!

yencabulator · 2026-04-08T18:03:28 1775671408

> I think for the same model wall time is probably a more intuitive metric; at the end of the day what you’re doing is renting GPU time slices

This is a bit too much of a simplification.

The LLM provider batches multiple customer requests into one GPU/TPU pass over the weights, with minimal latency increase.

The LLM provider may in fact be renting GPUs by the second, but the end user isn't. We the end users are essentially timesharing a pool of GPUs without any dedicated "1 vGPU" style resource allocation. In such a setting, charging by "GPU tick" sounds valid, and the various categories of token costs are an approximation of cost+margin.

nsomaru · 2026-04-07T03:39:48 1775533188

They already bucket when context goes above 200k

refulgentis · 2026-04-07T04:55:21 1775537721

No longer