Open-model API prices, live

0 priced models on 0 independent servers right now. USD per 1M tokens; "from" is the cheapest listing, "median" is the middle of the market.

No priced models are online at the moment. Check back shortly, or be the first server.

How pricing works

Each provider lists their own input and output price per 1M tokens for each model they serve. A request is routed to the best-scoring online server for the model (throughput, latency, uptime, error rate) unless you pin one with X-Server-Key. You are charged that server's price for the tokens actually used, and the platform keeps a fee from the provider's side, not yours.

The "In / Out" column shows the cheapest input and output listing separately; a chat workload is mostly output tokens, a summarisation workload mostly input, so pick by the number that matches yours.

"Served as" is the quantisation each provider reports (Q4_K_M and F16 of the same model are not the same model); "Where" is the provider's country from the connection. Open a model to see every server individually and pin one with X-Server-Key. Send X-Region: DE (any ISO country) to prefer servers near you.

Use these modelsGet an API key, point any OpenAI-compatible client at /v1, pay per token. No minimums. Earn with your GPURun one command on your Ollama / LM Studio / vLLM box, set your price per 1M tokens, keep the rest.

Updated 2026-10-02 02:36 UTC. Machine-readable: /api/models (no key needed).