OAra Labs | Docs

Reference

OpenAI-compatible API

The daemon serves /v1/models and /v1/chat/completions on the OpenAI chat-completions wire. Anything that already speaks it — Open WebUI, LobeChat, Continue, Zed, the openai client libraries, curl — can talk to Prometheus without a Prometheus-specific client. Beacon stays the cockpit; this is the second door.

Not a proxy This is not Prometheus forwarding your request to OpenAI. It is Prometheus being the server: the reply comes from the agent loop, with its memory, its tools and its security gate, running on whatever model the daemon is configured for. The wire format is OpenAI's so that existing clients work unchanged.

Endpoints

RouteWhat it does
GET /v1/modelsThe model catalog as OpenAI-shaped rows. The id is the key you pass as model: local for the configured primary, or a cloud preset. Only models the daemon can currently reach are listed.
POST /v1/chat/completionsOne agent turn. Send the whole conversation; get the reply as a completion object, or as Server-Sent Events in the OpenAI chunk shape when "stream": true.

Both sit behind the same bearer-token middleware as /api/. There is no separate key for this surface.

Authentication

The web API token, as a bearer. oara token show prints it; it also lives in the daemon's env file as PROMETHEUS_API_TOKEN. See Tokens and the open web API before you expose the port beyond localhost or your tailnet — this surface is exactly as public as the rest of the API.

one turnany machine that can reach the daemon
curl http://localhost:8005/v1/chat/completions \
  -H "Authorization: Bearer $PROMETHEUS_API_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"model": "local", "messages": [{"role": "user", "content": "Hello"}]}'

Pointing a client at it is the same three values everywhere: base URL http://<daemon-host>:8005/v1, API key = the token, model = a key from /v1/models. Open WebUI calls these an “OpenAI API connection”; Continue and Zed call it an OpenAI-compatible provider.

What is different from OpenAI, on purpose

WhereWhat happens, and why
stateStateless per request. Each call runs one agent turn in a fresh openai:<id> session that persists nothing to the daemon's conversation memory. Like OpenAI, the client owns the history and sends all of it every call. The daemon's own sessions, memory extraction and retention are untouched; it does not learn from a compat client.
toolsThe tools are the daemon's. A request carrying tools, functions, tool_choice or function_call is refused with 400 tools_unsupported. Refusing is honest; ignoring would let a client believe its tools were in play. The model uses Prometheus's tools, gated by the security gate, and the client only ever sees text deltas. A tool that needs an approval blocks the turn exactly as it would for Beacon — the operator approves it there.
samplingGeneration settings are ignored. temperature, top_p, max_tokens, n, stop, logprobs. The agent loop owns them. They are accepted and dropped rather than refused, so clients that always send them keep working.
systemsystem messages are appended, not substituted. They land under an “Instructions from the connecting client” heading after the daemon's own system prompt, and never replace the identity, tool and safety text the loop is built on.
modelmodel is a catalog key, not a model name. It applies a per-request override and clears it afterwards. An unknown key is 404 model_not_found, with the fix in the message: pick one from GET /v1/models.
contentText only. Message content must be a string or text parts. Image and audio parts are refused with 400 unsupported_content. The gateways handle inbound media; this surface does not.

One extension

"mode": "chat" in the request body runs the turn with no tools at all — a plain model reply. The default, "agent", is the full loop. Anything else is 400 invalid_mode.

Replies also carry a prometheus object beside the standard fields: the session_id the turn ran under and the mode it used. Standard clients ignore it; it is there for people reading logs.

Errors

Every refusal uses OpenAI's error envelope — {"error": {"message", "type", "code", "param"}} — so client libraries raise it as their own error type and the message says what to change. The codes you will meet:

Status and codeMeaning
400 tools_unsupportedRemove the client-side tool fields.
400 messages_requiredmessages must be a non-empty list.
400 last_message_not_userThe final message must be from the user.
400 invalid_roleRoles are system, user, assistant.
400 unsupported_contentText only.
404 model_not_foundNot a key from /v1/models.
503 loop_unavailableThe web server answered but the agent loop is not wired to it yet — usually the daemon is still starting. oara doctor will say.
502 turn_failed (or a more specific kind)The turn started and the loop died — a provider error, most often. Non-streaming replies get the status; a stream that has already sent headers reports it as its last event instead.
A 200 is not proof either A streaming reply that fails mid-turn cannot change its status code after the headers are sent. The stream sends an {"error": …} event and then [DONE] instead. Read the whole stream, not just the first chunk, and treat that event as the failure it is.