Responses, Messages and embeddings on the inference API
Chat models accept stateless Responses, text Completions and Anthropic Messages requests, including streaming
and client function tools. Named endpoints apply their existing access rules and spending caps to these routes.
BGE-M3 embeddings accept batches of up to 2,048 inputs when the model is available.
Retrying a completed non-streamed request with its idempotency key replays the same answer without a second
charge. Stored Responses, hosted tools and Messages token counting remain unsupported.
Completed chat streams include the OpenAI terminal marker after the final usage report. Interrupted streams retain their incomplete state.