Hosted inference and Composer

Point any OpenAI SDK at https://inference.nodus-compute.ai/v1 with a Nodus API key and call the catalog models through chat completions, streamed or not, plus audio transcription, translation and speech. Every response carries a Nodus-Request-Id whose receipt stays at GET /v1/requests/{id} for 30 days. A request whose outcome is never known is never charged.

nodus/auto (Composer) chooses one catalog model for each request and names it in the x-nodus-routed-model response header. Named inference endpoints give a project its own URL with requests-per-minute, tokens-per-minute, concurrency and spend limits, and can be limited to chosen API keys. The console playground calls the same API. See Inference.