AI Connect
A person links their own Anthropic or OpenAI key once, on the hosted page. You point the official SDK at OpenMarkets and run Claude or GPT on their key — never seeing it — or on managed credits. Requires ai:invoke, which is on every plan.
What you skip
OpenMarkets handles
- Holding provider keys — tested before they are stored, encrypted, never returned (last four characters only)
- Choosing the key per call: the account’s own, or OpenMarkets’ on managed billing
- Metering usage, including cached tokens, on both providers
- Per-app invoke consent, a monthly token cap the user sets, and the credit ledger for managed calls
- Heroku’s 30-second limit: long calls answer 202 and can be polled, or stream
You do
- Ask for
ai:invoke(apps) or hold it on your plan (everyone does) - Point the provider SDK at the pass-through base URL, with your om_ key where the provider key goes
- Pick a billing mode per call
- Decide which model, and stream anything long
How a key gets linked
- Yourself — the console's Connect page → Link a provider. The hosted page tests the key against the provider before it is kept; a rejected key never reaches the database.
- Your app's users — open the hosted page for them with the
account:linkscope:POST /flow/v1/auth/account/link-sessions(see Sign in with OpenMarkets). A key linked there, or an existing key the user chooses to Share with your app there, is consented to your app only. Holdingai:invokeis not enough on its own: until the user shares a key with your app,byokanswers403 scope_not_grantedandautoruns on managed billing — paid from the user's credits, since an app key acts in their workspace.GET /flow/v1/ai/providersshows your app's consent ininvoke_granted. The user can stop sharing on the same page, and uninstalling your app ends it. - Your Connect users — add providers to your Connect flow's AI providers list (it defaults to none). They link on the hosted page; invoking for them also needs
connect:trade.
Billing, decided per call
| Mode | Runs on | Charged | When there is no usable linked key |
|---|---|---|---|
| byok | the account's linked key | never, by OpenMarkets — the provider bills the key owner | 409 ai_provider_not_linked; a key the provider rejected is 409 ai_key_invalid. Never falls through to a charged call. |
| managed | OpenMarkets' key | reasoning credits, at the rate per model (1,000 credits = $1) | — |
| auto | the linked key when usable, else managed | only the managed leg | managed. A rejected key is still an error, never silently replaced. |
Managed calls precheck a worst-case cost (input plus the full output ceiling) and answer 402 out_of_credits rather than starting something they cannot pay for. A person can cap their own key's monthly usage; the cap answers 429 token_cap_reached. AI calls never count against your request quota.
Provider API pass-through
Send the provider's own request to OpenMarkets and get the provider's own response back, streamed or not. Tools and tool results, prompt caching, thinking, images and structured output work exactly as the provider documents them, because nothing is translated. OpenMarkets only picks the key and records usage.
POST /flow/v1/ai/anthropic/v1/messages— Anthropic MessagesPOST /flow/v1/ai/openai/v1/responses— OpenAI ResponsesPOST /flow/v1/ai/openai/v1/chat/completions— OpenAI Chat Completions
The paths mirror the providers', so the official SDKs work with a baseURL change. Your OpenMarkets key goes where the provider key would; set X-OpenMarkets-Account to act for a Connect user, and X-OpenMarkets-AI-Billing to auto (the default), byok or managed.
import Anthropic from '@anthropic-ai/sdk';
const claude = new Anthropic({
baseURL: 'https://api.openmarkets.ai/flow/v1/ai/anthropic',
apiKey: process.env.OPENMARKETS_API_KEY, // om_…, not theirs
defaultHeaders: {
'X-OpenMarkets-Account': userId, // a Connect user; omit for your own account
'X-OpenMarkets-AI-Billing': 'byok', // never fall back to a charged call
},
});
const response = await claude.messages.create({
model: 'claude-sonnet-5',
max_tokens: 16000,
tools: [{ name: 'get_props', description: '…', input_schema: { type: 'object', properties: {} } }],
messages, // tool_use / tool_result blocks and all
});
import OpenAI from 'openai';
const openai = new OpenAI({
baseURL: 'https://api.openmarkets.ai/flow/v1/ai/openai/v1',
apiKey: process.env.OPENMARKETS_API_KEY, // sent as Authorization: Bearer om_…
defaultHeaders: { 'X-OpenMarkets-Account': userId },
});
const response = await openai.responses.create({ model: 'gpt-5', input, tools, max_output_tokens: 8000 });
- Errors come back in the provider's shape, so the SDK raises its own typed errors. A provider error is relayed with its status and body; OpenMarkets' own (
ai_provider_not_linked,out_of_credits,not_available_on_managed,request_too_large) use the same shape. - Every response carries
x-openmarkets-inference-idandx-openmarkets-billing. Usage — cached tokens included — is on the provider's ownusageobject. - Stream long calls (
stream: true). A non-streamed call that runs past 20 seconds gets its 200 early and leading whitespace while it waits (still valid JSON); a provider error after that point arrives in the body under the 200. - Streamed Chat Completions always include usage (
stream_options.include_usageis set for you): expect one final chunk with emptychoices. - Request bodies up to 32 MB. Model output is never stored by OpenMarkets on this surface.
max_tokens / max_output_tokens / max_completion_tokens), and anything that cannot be priced in credits or would keep state on OpenMarkets' provider account is refused with not_available_on_managed: beta headers, code execution, file search, file references, stored or chained responses (previous_response_id, store: true), containers and premium speed tiers. Your own tools, prompt caching, thinking and web search all work. With a linked key everything the provider offers works, and nothing is charged by OpenMarkets.One request shape for both providers
POST /flow/v1/ai/inferences normalizes both providers into one small request for single questions and side-by-side model comparisons. It will never keep up with the providers' feature lists — for tools, multi-turn agents and caching, use the pass-through above.
reasoning: true in the catalog) reject a custom temperature and spend hidden reasoning tokens out of max_tokens. Leave temperature unset for them, and leave room in max_tokens — a budget sized for the answer alone can come back empty./flow/v1/ai/inferencesRun a model. Body: provider ('anthropic' | 'openai'), model (the provider's id, passed through verbatim), messages [{ role: 'user' | 'assistant', content }] (the first must be 'user'; up to 200), and optionally system, max_tokens (1–128000, default 8192; a model with a lower ceiling rejects more), temperature (0–2, sent only when set), output_schema { name, schema, strict } to get a parsed JSON object back as `output`, web_search (true = 5 searches, or { max_uses: 1–10 }), cache (true marks the prompt for Anthropic prompt caching; OpenAI caches on its own), billing ('byok' | 'managed' | 'auto', default auto), wait_seconds (0–25, default 20) and metadata (up to 20 string labels of your own). Any other field is 400 invalid_request. Answers 200 when the call finishes inside wait_seconds, otherwise 202 with an inference_id to poll. Send an Idempotency-Key to retry safely. Errors: 400 model_not_available, 402 out_of_credits, 403 model_not_allowed / scope_not_granted, 409 ai_provider_not_linked / ai_key_invalid, 413 request_too_large, 429 token_cap_reached, 503 managed_unavailable.
// Request body
{
"provider": "openai",
"model": "gpt-5-mini",
"billing": "byok",
"web_search": { "max_uses": 3 },
"messages": [{ "role": "user", "content": "Bills or Dolphins at these prices? Answer as JSON." }],
"output_schema": {
"name": "pick",
"strict": true,
"schema": {
"type": "object",
"properties": { "team": { "type": "string" }, "probability": { "type": "number" } },
"required": ["team", "probability"],
"additionalProperties": false
}
}
}
// Response.data
{
"inference_id": "inf_8f2c…",
"status": "succeeded", // running | succeeded | failed
"provider": "openai",
"model": "gpt-5-mini",
"served_model": "gpt-5-mini-2025-08-07",
"billing": "byok",
"output_text": "{\"team\":\"Buffalo Bills\",\"probability\":0.61}",
"output": { "team": "Buffalo Bills", "probability": 0.61 },
"stop_reason": "end", // end | max_tokens | stop_sequence | refusal | tool_use | paused | context_exceeded
"usage": {
"input_tokens": 14212,
"output_tokens": 96,
"cache_read_input_tokens": 0,
"cache_write_input_tokens": 0,
"credits_charged": 0
},
"web_search": {
"requests": 2,
"sources": [
{ "url": "https://www.espn.com/nfl/injuries", "title": "NFL Injuries - ESPN" },
{ "url": "https://www.nfl.com/news/…", "title": "Week 3 injury report" }
]
},
"error": null,
"metadata": null,
"latency_ms": 2140,
"created_at": "2026-09-21T21:02:11Z",
"completed_at": "2026-09-21T21:02:13Z"
}
/flow/v1/ai/inferences/:inference_idPoll an inference that answered 202. Same shape as the POST response; status moves from running to succeeded or failed. A failed inference carries error { code, message } — a failure is a result, not an HTTP error. An inference left running by a restart is failed on read with code interrupted.
| Name | In | Type | Required | Description |
|---|---|---|---|---|
| inference_id | path | string | yes |
/flow/v1/ai/inferences/inf_8f2c/flow/v1/ai/providersWhich AI providers the acting account has linked, with each key's status, last four characters, whether the caller may invoke it, and the user's monthly token cap. Never returns a key.
/flow/v1/ai/providers{
"providers": [
{
"provider": "anthropic",
"label": "Anthropic",
"key_hint": "9f2a",
"credential_status": "connected", // pending | connected | failed
"invoke_granted": true,
"monthly_token_cap": null,
"last_used_at": "2026-09-21T21:02:13Z",
"connected_at": "2026-09-21T18:40:02Z"
}
]
}
/flow/v1/ai/modelsThe models OpenMarkets runs on its own keys (billing: managed), with the credit rate per 1,000 tokens (1,000 credits = $1) and whether your plan allows each one. A linked key is not limited to this list.
/flow/v1/ai/models{
"models": [
{
"model_key": "gpt-5-mini",
"provider": "openai",
"label": "GPT-5 mini",
"input_credits_per_1k": 0.875,
"output_credits_per_1k": 7,
"reasoning": true,
"allowed": true,
"is_default": false
},
{
"model_key": "claude-sonnet-5",
"provider": "anthropic",
"label": "Claude Sonnet 5",
"input_credits_per_1k": 7,
"output_credits_per_1k": 35,
"reasoning": false,
"allowed": true,
"is_default": false
}
]
}
Web search
Per request: "web_search": true (up to 5 searches) or { "max_uses": 3 } (1–10) lets the model look things up before it answers; omit it for none. It works on every Anthropic and OpenAI model, including with output_schema. The response carries web_search.requests and the sources returned. Searches are billed per search on managed (about $0.01, 35 credits) plus the result text as input tokens — typically 5k–25k per call — so budget max_tokens generously for reasoning models.
For a Connect user
Add X-OpenMarkets-Account on either surface. Your organization needs connect:trade, the user must have linked a provider on the hosted page, and under byok the user must have granted your organization invoke consent — otherwise 403 scope_not_granted. Under auto, withheld consent falls through to managed billing on your credits.