You pass verification
We check that your endpoint is OpenAI-compatible, that your model metadata is accurate, and that your prices are readable by machine. We then benchmark latency and throughput from the regions you serve.
AIVory Smart Inference · Providers
AIVory Smart Inference routes every customer request to the cheapest provider that can serve it. Send a test key and we benchmark your endpoint — latency, throughput and error rate, per model, from the regions you serve. You keep those numbers whether or not we ever route to you.
Smart Inference is in early access and the pool is still small. That is the honest argument for joining now rather than later: less competition for the traffic there is, and a longer health record by the time the volume arrives.
We reply within two business days.
How volume is allocated
There is no placement to negotiate and no tier to buy. You pass verification, the router scores you on every request, and the traffic follows the score.
We check that your endpoint is OpenAI-compatible, that your model metadata is accurate, and that your prices are readable by machine. We then benchmark latency and throughput from the regions you serve.
For each request the router removes every provider that lacks a needed capability, has too small a context window, or is unhealthy. It then picks the cheapest candidate that meets the reliability and speed requirements.
Win on price in a region and you get that region's requests. Health probes run continuously, with a five-minute grace period, so a short network blip does not remove you from the pool.
Integration requirements
These are not a wish list. Each one is something the router or the billing system reads directly, so a gap here blocks the review.
/v1/chat/completions with streaming. We do not write
per-provider adapters, so the request and response shapes have to
match.
Exact model IDs, context window and max output per model. Also whether each one supports tool calling, JSON mode and vision, and the quantization you serve it at. The router filters on these fields, so an optimistic answer costs you traffic when requests fail.
Either prices inline in /v1/models, or a price endpoint we
can poll. Tell us which, and name the fields. A price list in a PDF
cannot be routed on.
Streams must honour
stream_options {"include_usage": true}, and non-streaming
responses must always return usage. We bill our customers
on those numbers, so we cannot estimate them.
Requests per minute, tokens per minute and concurrency, per key. Tell
us what a 429 looks like and whether you send
Retry-After. The router needs to know when to fail over
rather than retry.
One URL we can watch, and one human we can reach during an incident. Health probes tell us that you are down; they do not tell us when you will be back.
Set your expectations
Here is the deal in full, both halves of it. It is fairer to say this on the page than in the third email.
What happens next
The key needs a little eval credit on it. Without traffic we cannot measure anything.
Time to first token, throughput and error rate, per model, from the regions you serve. We check that your reported usage matches what we count.
You get the benchmark results, including where you sit against the pool. This happens whether or not we go ahead.
See below. This is the step that decides whether routed volume works for both sides.
We start with a small share of traffic and widen it as the health record builds.
The commercial question
We sell to the customer and we pay you. Our margin sits between your price and theirs, so your price is the input that decides whether the route is ever competitive.
That is a workable starting point. You compete on the same terms as everyone else in the pool, and you win the requests where you are cheapest for the capability the customer asked for.
Say so in the application. Routed volume arrives without you paying to acquire it, and a rate that reflects that usually moves you up the score in far more requests than a marketing budget would.
Questions
Both, depending on what you supply. If you run an OpenAI-compatible inference API, we route token traffic to you and settle on the usage your endpoint reports. If you rent out GPUs, that is the spot marketplace side, and it settles on instance time. Tell us which one you are in the application, or both.
We cannot promise a number, and we will not pretend otherwise. What we can tell you is the mechanism: the router re-scores every request, so a price change on your side shows up in your traffic within the hour rather than at the next contract review. A provider who becomes cheapest on one popular model can take a large share of it overnight, and lose it the same way.
Not by default. Customers see a pool label rather than a provider name. We may change that later, but plan around how it works today. This is why we say volume follows your numbers, not your logo.
No. Sell the same capacity anywhere else you like. We only ask that your published prices stay current, because the router polls them and a stale price means we route on numbers that are wrong.
No. The router does not weight providers by brand or by how long they have been trading. It filters on capability and health, then picks on price. A new provider with an honest model list and a working endpoint competes on the same terms as anyone else.
Yes. Some requests are served by reclaimable capacity, and the response carries an X-SI-Spot header so the customer knows. Say in the application which of your capacity is interruptible and what your reclaim notice looks like.
It depends almost entirely on how complete your first message is. With a working test key and an accurate model list we can usually return benchmark numbers within a week. Without them the review does not start at all.
Health probes take you out of the candidate list, after a five-minute grace period so a short blip does not remove you unnecessarily. In-flight requests fail over to the next candidate, and you come back automatically once the probes recover. That is why we ask for a status page and a human to contact.
There are two ways in. Take the short one if you want to move now, or the full one if you want the review to start immediately.
Your company, your API base URL, the regions you serve, and a test key with a little eval credit on it. That is enough for us to start measuring, and we will ask for the rest when we need it.
Send the short versionAll six at once. Nothing to chase, so the technical review starts the day it arrives.
/v1/models, or a price endpoint we can poll. Name the fields.stream_options {"include_usage": true}, and that non-streaming responses always return usage. We bill on those numbers.One more thing, and it is the one that decides everything: are the prices you publish your public rate, or is there a partner rate for routed volume?
Both buttons open your mail client with the questions ready to fill in. You can also write to [email protected] directly. We reply within two business days.
We will come back with benchmark numbers. That is the fastest way to find out whether this works for both of us.
Apply to the pool