Every Offer, Live
Node owners set their own discount off each model's reference price. Your requests go to the deepest discount that's online and healthy, and fail over to the next one.
Loading markets…
Discounts are measured against each model's reference price. TTFT is time to first token, tok/s is generation speed and latency is the whole request (all medians over 7 days, including verification checks). Uptime is served ÷ attempted; "Checks" is the share of hidden verification prompts answered correctly. A dash means not enough data yet.
Route by discount
Require a minimum discount with a header, or right in the URL:
X-Min-Discount: 80https://<host>/min80/v1/chat/completionsIf no live offer meets it you get 404 minimum_discount_not_met instead of a pricier fill.
Set your own price
Host a model, then pick your discount and an optional daily cap per model in the dashboard. Deeper discounts get routed first.