Self-Hosted Models
Listed inSelf-Hosted ModelsModels & Providerson
Sizing a GPU for a self-hosted model means budgeting for more than the weights. The KV cache grows with concurrency, and exhausting it preempts in-flight requests that are then recomputed from the start.