Skip to content

Inference

View Markdown

The inference data plane speaks the OpenAI and Anthropic wire formats. nodus.llm returns the official clients, configured with your key and the Nodus base URL, so every feature of those SDKs works unchanged.

Terminal window
pip install "nodus-compute[openai]" # or [anthropic]
examples/python/inference/chat.py
"""Chat with a catalog model through the OpenAI-compatible inference data plane.
Needs `pip install "nodus-compute[openai]"`. Run it with `python examples/python/inference/chat.py`.
"""
import nodus
def main() -> None:
client = nodus.llm.openai() # the official OpenAI client with your Nodus key and base URL
reply = client.chat.completions.create(
model="nodus/gpt-oss-20b",
messages=[{"role": "user", "content": "Say hello in five words."}],
max_tokens=32,
)
print(reply.choices[0].message.content)
if __name__ == "__main__":
main()
client = nodus.llm.openai(project="nlp") # usage is attributed to the project
claude = nodus.llm.anthropic() # messages API
aclient = nodus.llm.async_openai() # AsyncOpenAI

Inside a Job, Function or Sandbox, nodus.llm uses the container’s built-in proxy: no key is needed, and the calls are billed to, and capped by, the run that makes them.

Use GET /v1/models for the available on-demand catalog. Public on-demand models require verified released weights for the served checkpoint. New or unverified versions stay unavailable until reviewed. Open-weight licenses can still impose conditions; availability does not mean unrestricted licensing. Indra selects among eligible catalog models rather than representing a separate weight release.

An InferenceEndpoint gives a model its own base URL with rate limits, allowed keys and a spending cap.

ep = nodus.InferenceEndpoint.create("support-bot", "nodus/gpt-oss-120b", rpm=120, max_concurrent=8, max_cost=50)
client = ep.openai()
print(ep.usage()) # requests, tokens and cost over the last 24 hours