The business case is on the other page. This one answers the questions you actually have: what is on my request path, what happens when you are down, and what do I have to change.
Status: in development. Early adopters get access in Q4 2026.
Every call to your API carries a token, and your code checks it before the call reaches your handler: with our SDK, or with the JWT library or gateway you already use. The token is a standard signed JWT, checked the standard way: signature, issuer, audience, expiry and scope. Any language can do it.
Your API traffic never passes through us. No service of ours sits between your customer and your API. Your code sends us usage, in batches, after the response: no request body, no response body, no headers you did not choose to send. It reads our public keys to check tokens. You can watch every call it makes to us with tcpdump.
Use our SDK, or none of our code at all. The SDK does the check and the usage report for you. Without it, any JWT library does the check, and the usage API is plain HTTP.
Your API keeps its address. Nothing re-points.
Keep Kong, AWS, Azure, Apigee, or none at all. If you have one, it can check the token too.
The check runs inside your application or gateway, where every call already goes.
Your customer presents a token. Your code verifies it against our public keys, cached on your side, and checks that the token's scope covers the route or tool being called. Then it serves the call or refuses it. It does not ask us. The decision was made when the token was minted, and the token carries it. A call without a valid token never reaches your handler.
A call waits on us in one case only: when your code does not yet hold the key its token was signed with. That happens once at start and once after we rotate keys. Our SDK fetches at most once a minute, however many calls arrive.
Every copy of your application reads the same token, so ten replicas cannot double-spend one balance between them. There is no balance or counter to keep, and nothing to make durable.
request ──▶ your API our SDK, your JWT library, or your gateway │ ├ 1 verify the token # our keys, cached ├ 2 check the scope # this route or tool ├ 3 run your handler # or refuse the call └ 4 queue a usage event # after the response └──▶ Keelie, in batches
Once your code holds our keys, steps 1 to 3 make no call to us. Step 4 happens after your customer has their response.
import { createRemoteJWKSet, jwtVerify } from "jose"; // Our public keys, cached, with a minute between fetches. const keys = createRemoteJWKSet(new URL(KEELIE_JWKS_URL), { cooldownDuration: 60_000, }); // A bad token throws: answer 401. // A missing scope returns null: answer 403. export async function check(token: string, scope: string) { const { payload } = await jwtVerify(token, keys, { issuer: KEELIE_ISSUER, audience: YOUR_API, algorithms: ["RS256"], clockTolerance: 30, }); const scopes = String(payload.scope ?? "").split(" "); // azp is the customer's credential. Keep it for the usage event. return scopes.includes(scope) ? payload.azp : null; }
A standard library, not ours. KEELIE_ISSUER and KEELIE_JWKS_URL come from your issuer's discovery document. Any JWT library, in any language, does the same checks.
We read the balance at mint time, not at request time. A healthy balance gets the full token lifetime of your plan. On Growth and Scale, a balance running dry gets a shorter one. An empty balance gets no new token at all.
So when we are unreachable, nothing stops immediately and nothing runs forever. Live tokens keep working until they expire. The next mint is what fails.
| Your customer | What happens |
|---|---|
| Paying, healthy balance | Calls succeed until their token expires, at most your plan's normal lifetime. Then they are refused until we are back. |
| Runs dry during the outage | Calls succeed, then stop. Your exposure is their burn rate multiplied by the time left on the token. |
| Already low on credit | Cut off when their current token expires. On Growth and Scale, their low balance has already made it short. |
| Sitting on a spend cap | The cap holds. It rides the token the same way the balance does. |
| Never seen before | Refused. They have no token and cannot get one. |
The balance can shorten a token but never below your plan's floor. That floor is your worst case during an outage — a longer floor means more of a customer's burn is already authorised when we go dark, and a shorter floor means less.
Upgrading buys a tighter bound. It is the one pricing lever we have that is also a reliability number, and we would rather you knew that than discovered it.
The floors are 30 minutes on PAYG, 5 minutes on Growth and 1 minute on Scale. On PAYG the floor is also the normal lifetime, so a token there never shortens.
A new copy of your application has no keys yet. If it starts while we are unreachable, it refuses calls until it can fetch them, so scaling up during an outage adds copies that refuse. Copies that were already running keep their keys. Our SDK holds unsent usage in memory, so a copy that stops before sending it loses those events. Those calls go uncharged: the error runs in your customer's favour, never against them.
You get your issuer and your data as a standard export, and your money is already yours. The export holds your settings, your signing keys, your customers' credentials and, on Keelie's sign-in, their logins. It also holds your plans, balances and usage. Live tokens keep working until they expire, and after that you run the issuer yourself.
What does not move: our code, so your issuer signs tokens without checking balances until you build that part. People approve their agents again. On PAYG and Growth the issuer's address changes, so your checks need the new one. On Scale it is on your own domain and moves with you.
After your API answers, your code records the call: which route or tool, and what it cost at your prices. It sends events in batches, within seconds, never inside the call. Every event carries its own idempotency key, so a retry after a network failure cannot double-charge your customer.
If you call the usage API yourself, follow the same rules. Send within seconds rather than waiting for a full batch, because the balance only moves when usage arrives, and give each event its own key.
There is only one count, and it comes from you. Your code reports what was used, and Keelie takes exactly that off the balance. Keelie keeps no count of its own.
Their credits live with us, and show on your screens. Keelie holds each customer's balance. Your pages read it from our API, and top-ups go through your Stripe account.
⚠️ The two paths fail in opposite directions, on purpose.
May this call proceed? The token answers it. When there is no valid token, the answer is no. Money questions do not get the benefit of the doubt.
Reporting what was used must never take your API down. Our SDK holds unsent events in memory and sends them again when we come back, with the same keys.
If we disappear mid-request, your customer gets their response and we get the event later. That ordering is deliberate.
Token endpoint, JWKS, key rotation, discovery documents, and scopes bound to the plan that was actually paid for. Client credentials for a service. Authorization code with PKCE when a person approves an agent: they sign in, see what their plan covers, and the agent connects itself to exactly that.
They become your customer, not ours. They sign in through Keelie's sign-in or, from Growth, through your own identity provider, and Keelie keeps their API account linked to that login.
On your own domain, our SDK serves the documents an agent reads first: the metadata that points it to our issuer, and your AI Catalog. Without the SDK, you serve the same files, and we keep their content current.
This is the part that is genuinely months of work, and it is the reason 8.5% of more than 5,000 MCP servers have built it while 53% still hand out static keys.
The integration is a token check and a usage report. The rest is configuration.
Ask them at info@apiable.io. We would rather lose an evaluation on a straight answer than win one on a vague page.