Agent Loop
Pydantic-validated tool calls, PreToolUse/PostToolUse hook pipeline, permission governance. Works with any model.
OAra Labs · open-source agent daemon
A fire you own doesn't go out.
The agent daemon that stays on in your house and makes the model you already have do real work.
A daemon is a program that never stops running. The word comes from the Greek daimōn, a guiding spirit that stays with you. Prometheus is both.
After Santucci, 1593. The stand is Beacon; the spinning sky is the daemon.
Active development. Expect rough edges. Fixes land weekly. Feedback welcome.
What doesn't work yet ↓The Stack
Prometheus is the daemon. The pillars plug into it through an open hook contract. Remove them and Prometheus runs exactly as before.
Today
The agent daemon. The base of the stack.
The desktop cockpit. Sits on Prometheus.
The cockpit in your pocket. Same daemon, same control plane.
Next
Fast decisions. Which model, which tool, when to stop.
Planning, and checking the work against the goal before delivering it.
Coordination across several machines.
Instinct lives with the body, on every node, always local. Cognition thinks with whatever brain Instinct wakes.
What leaves your machine
No telemetry upload, no analytics, no update check: what Prometheus records stays on your disk. What does leave, and when:
Origins
Built here; the adapted pieces say where they came from.
The Guards
9,700+ tests passed in CI on the v0.9.5 release commit. A passing test proves the code runs. It does not prove anything calls it. Every guard below exists because a feature was built, tested, green — and never wired.
| Enforced by | The invariant |
|---|---|
| test_no_site_resolves_the_wiki_root_independently | No source file may resolve the wiki root on its own — one resolver, or the build fails. |
| test_every_live_config_key_exists_in_the_default_template | Every key in the live config must exist in the shipped template — live ⊆ template, with no allowlist. Compares against a real install, so it skips on a fresh clone where there is nothing to compare. |
| test_no_new_config_key_without_a_reader | A register of config keys nothing reads, which can only shrink. An unregistered key with no reader fails — and a registered key that gains one fails as stale, so the list cannot rot. |
| test_example_call_uses_real_param_names | A tool's advertised example must validate against its own schema. The example ships inside the tool advertisement, so a wrong one teaches the model a parameter that does not exist. |
| test_every_guard_declares_its_enforcement | Every media and rate check declares fail-closed or fail-open at construction. The registry is built at module level, so an undeclared guard is an import error, not a test failure. |
| test_tripwire_end_to_end | An acceptance test that terminates in a registered test double fails — registration is what makes a double detectable. Wildcard exemptions are refused; only individually-named doubles pass. |
| test_every_advertised_document_extension_is_admitted | Every allowlisted file type must be provably admitted, not merely “not refused”. Breach tests prove the door closes; only admission tests prove it opens. |
The last one is there because its absence shipped. A control suite whose every case asked “does disabling this let something bad through?” — and none asked “does this let the permitted things through?” — stayed green while the document surface silently degraded to PDF-only: 19 of 20 advertised types refused, including two the allowlist explicitly permitted. Over-refusal looks exactly like the control working.
Behavior has a guard of its own. A change to the daemon’s seams, or to anything that shapes a request to the model, replays recorded turns through a real daemon. Any difference fails the Parity check: a tool call, a gate decision, a checkpoint, a memory write, a telemetry row, the final reply, or a single request sent to the model.
Verification
At the top of the page that would be a rule to take on faith. Down here it is what the daemon does when it does not know:
✗ carries a fix: line naming what to change. A red mark with no next step is a mark, not a diagnosis.provider_degraded frame naming the requested model and the served one. Never a silent swap.Core systems
Pydantic-validated tool calls, PreToolUse/PostToolUse hook pipeline, permission governance. Works with any model.
DAG-based conversation compression. Every message persisted to SQLite. Old messages summarised into expandable nodes. FTS5 full-text search. Works within 32K context windows.
Proactive background intelligence. Watches telemetry, sends nudges during idle. AutoDream engine: wiki lint, memory consolidation, telemetry digest, knowledge synthesis. Budget-capped. Never exceeds autonomous trust level.
Compounding knowledge base. Memory Extractor runs every 30 min. WikiCompiler builds entity pages with cross-references. Query results file back as new pages. Obsidian-compatible.
Four trust levels: BLOCKED, APPROVE, AUTO, AUTONOMOUS. Configurable deny lists over a floor no mode waives: rm -r at a protected root is refused, and what substring patterns used to catch is left to what actually holds — resource limits and, on Linux, the kernel floor. Workspace boundary enforcement, bash intent analysis, memory security scanning. Token-authenticated REST and WebSocket control plane.
The chat gateways are the one surface exposed to the internet by design, so they are checked cheapest-and-earliest-first. Per-chat rate budgets under a global ceiling. Declared MIME and the size cap before download — then the download under a hard byte ceiling, because the peer supplied that size and the pre-check believed it. Magic bytes are sniffed after; a declared type that disagrees is refused.
Known token shapes become <redacted> before conversation history, memory, telemetry or training data is written. The live conversation is left alone, so a token you paste still works for that session. oara scrub finds what earlier releases kept, and changes nothing until --apply.
Tool calls that actually work on local models. Validator catches malformed tool calls. Formatter translates between model formats. Enforcer constrains output to valid schemas. Each model runs at a tier: the registry’s, or for a local model it doesn’t know, the one its served chat template implies — and adapter.model_tiers overrides it per model. Retry engine with structured error feedback.
Point the agent at a coding task and it iterates in a sandbox until the build and tests pass. With LSP enabled, diagnostics feed type errors back into the same loop, so it self-corrects against compiler ground truth. The result is a reviewable diff — nothing touches your branches until you merge.
Successful tool-call traces and adapter repair-pairs are captured, stored and mined into an exportable dataset. The data-collection half of a LoRA loop for the local model — training is on the roadmap. Fed by telemetry that never leaves your disk, and one config line turns it off entirely.
Demonstrate a task once — screen and DOM captured — and the daemon distils it into a reusable skill. Two-tier trust: a recording of the page itself becomes a skill once it passes a quality gate, while one read from the screen alone waits as a draft until you accept it. A screen recording of you doing your job is among the most revealing data you own, and every frame stays on your disk.
The skill tool is offered by default, even in a config setup wrote before it was. Every load is counted: /skills shows when each skill was last used, and the Curator ages skills by real use, never by file date. GEPA — off by default, experimental — proposes improvements from real loads, and nothing changes until a person runs oara gepa promote.
A local LLM judge scores task completion, tool accuracy and hallucination using constrained decoding on your own hardware. Every score records which judge produced it — base URL, model, and whether that model was pinned or auto-detected. Records written before this carry no judge at all: that means unknown, permanently, and they are not compared across paths.
Sessions survive daemon restarts. A running turn can be stopped mid-flight. Artifacts produced in chat are downloadable from the session.
One table of every local inference box you own, probed for the served model, its reported context window, detected vision and latency. A switch is probed first and refused if the box is down; the chat is budgeted at that box's window and the choice survives restarts.
Sessions
You can branch it, delete it for real, roll its files back a turn, and point it at a directory. The security boundary moves with it.
Copy a conversation up to any message into a new session and continue from there. The original is untouched: its rows, its cursors, its summaries. The branch keeps the real timestamps — rewriting them would make it look like it happened now.
Deleting a session hides it from the list and leaves the rows. Purging is the other answer — the one for “I pasted a customer's details in there”. It clears the full-text index before the row goes, takes the summaries with it, and overwrites the freed pages so the bytes are not in the file. It reports row counts per table, because a purge that only says ok is indistinguishable from one that matched nothing.
For a session with a workspace, every turn starts by capturing the workspace's files before any tool runs. Restore puts back what a turn overwrote, removes what it created, and reports every path. A workspace too big to capture is refused at the log, and the list shows the gap rather than a checkpoint that lies.
Point a session at a repo. The run's cwd, the security gate's allowed roots, the project instruction files and the checkpoints all follow it, per run — the boundary is the session's, not the process's.
A restored session's history is always readable after a daemon restart. With sessions.rehydrate on, the next message also restores its recent tail into the live working set, so the model resumes with context. Off by default.
All of it is REST under /api/sessions/<id>/ — fork, purge, workspace, checkpoints — and Beacon drives the same routes. Purge asks you to name the session twice; there is no undo behind it.
The kernel floor
A permission gate sees a command string. It cannot see cd ~/.ssh && cat id_*, a $HOME indirection, a glob or a heredoc for what they are. So the floor is enforced by the kernel, on the path as the kernel resolves it.
Bash's writes stop at the workspace, in a mount namespace (bubblewrap). Needs no root. It is what makes the workspace the complete write domain — and what lets the per-turn checkpoint be a real undo.
An AppArmor profile keeps bash out of the secret path families — keys, tokens, credentials — whatever the command string says. Needs root to load, so it is opt-in.
Set either floor to required and, on a host that cannot provide it, bash refuses to run and says why. It never falls through to an unconfined shell: a floor that quietly isn't there is worse than none, because everything downstream is written as though it is.
An exit code of 0 from the wrapper proves the wrapper ran, not that a transition happened. The preflight launches a process through the stack and reads the label the process reports for itself.
Linux only; there is no macOS equivalent. Defaults are honest about that: the write floor is auto — attempted everywhere, and where bubblewrap is missing it degrades once at ERROR and marks every call's result — and the read floor is off. oara doctor tells you which floors you actually have. Keys and modes are in the configuration docs.
Beacon
Beacon is the native desktop client for the daemon — chat with live tool timelines, coding runs and a Loop Manager, documents with AI redlines, Kanban, and per-provider key management across eighteen views. The model picker lists your own boxes, each with its live health, above the cloud presets. It pairs with a one-time six-digit code and works over localhost or your tailnet. The daemon stays the source of truth; Beacon is the window into it. Its Mission view runs the same Armilla you see above, fed real telemetry.
OpenAI-compatible API
Anything that already speaks the OpenAI chat-completions wire — Open WebUI, LobeChat, Continue, Zed — can point at the daemon and get Prometheus: memory, tools, the security gate, your local model. No Prometheus-specific client. Beacon stays the cockpit; /v1 is the door for everything else.
The client sends the whole conversation every call, and each call is one agent turn in a fresh session that persists nothing to memory. A compat client owns its history; the daemon does not learn from it.
A request that supplies its own tools or tool_choice is refused with a 400, not silently ignored. The tools are Prometheus's, gated by the security gate, and the client only ever sees text.
temperature, top_p, max_tokens, stop and logprobs are ignored on purpose. model is a key from GET /v1/models — your local box, or a cloud preset — and applies to that request alone.
curl http://localhost:8005/v1/chat/completions \
-H "Authorization: Bearer $PROMETHEUS_API_TOKEN" \
-H "Content-Type: application/json" \
-d '{"model": "local", "messages": [{"role": "user", "content": "Hello"}]}'Same bearer token as the rest of the daemon's API; oara token show prints it. Endpoints, refusals and streaming are in the docs.
Architecture
click a model — nothing else changes
The dashed boundary is the only part that knows which vendor it is talking to. Gateways, memory, guards and telemetry sit on the daemon's side of that line and do not change when the model does.
Your machine (or split across two)
│
├── Prometheus daemon (systemd)
│ ├── Agent Loop
│ ├── Telegram / Slack / Discord gateway
│ ├── Heartbeat + Cron
│ ├── Memory Extractor
│ ├── Wiki Compiler
│ ├── SENTINEL
│ ├── LCM Engine
│ ├── SYMBIOTE GitHub research → safe grafting → blue-green swap
│ └── AnatomyScanner infra self-awareness
│
├── SQLite databases
│ ├── memory.db
│ ├── telemetry.db
│ └── lcm.db
│
└── ~/.prometheus/
├── wiki/ knowledge base
├── sentinel/ dream logs
├── skills/auto/ learned skills
└── workspace/ sandboxed execution
Runs single-machine, or split brain and GPU across two boxes. All data in SQLite on your filesystem. Nothing phones home.
Install
The packaged release, v0.9.5 on PyPI. The plain package runs the daemon and pairs with Beacon; [full] adds Slack, Discord, MCP, the browser tool, the Anthropic provider and voice output.
pip install 'oara-prometheus[full]'
Apple Silicon (M1 and later) only — Intel Macs aren’t verified yet. It installs in a few minutes, and like the pip install it runs the daemon and pairs with Beacon.
brew install oaralabs/tap/oara
Use the full name: a bare brew install oara fails because the tap isn’t trusted.
For reading or changing the code, or running the test suite: an isolated install straight from the main branch with uv or pipx, or an editable clone.
uv tool install 'oara-prometheus[full] @ git+https://github.com/OAraLabs/Prometheus' # or: pipx install 'oara-prometheus[full] @ git+https://github.com/OAraLabs/Prometheus'
git clone https://github.com/OAraLabs/Prometheus.git && cd Prometheus pip install -e '.[full]'
Prerequisites: Python 3.11+, a running llama.cpp / Ollama / LM Studio / vLLM instance (or a cloud API key), and optional bot tokens for the messaging gateways.
pip install 'oara-prometheus[full]' # or Homebrew, or the developer path above
oara setup # detects llama.cpp / Ollama / LM Studio / vLLM, writes config, smoke-tests # no local GPU yet? start on a cloud model and switch later: oara setup --provider anthropic --api-key-env ANTHROPIC_API_KEY --model claude-sonnet-5
oara daemon # always-on: web API + gateways + cronIf anything misbehaves, oara doctor checks every subsystem and prints a fix hint per failure.
On first daemon start a web API token is minted and printed once — re-print it with oara token show.
On Linux, oara install-service writes a systemd user unit so the daemon survives reboots.
Prefer setup from a couch? Skip oara setup and run oara daemon bare — it boots in setup mode and prints a one-time six-digit pairing code. Beacon's wizard does the rest.
pre-1.0 · stated first
Provenance
Built by OAra Labs, starting with early scaffolding from OpenHarness. Design informed by Andrej Karpathy’s LLM Wiki concept, Lossless-Claw, and Sigrid Jin’s analysis of Claude Code’s agent-loop patterns. Full lineage and upstream licenses: NOTICE.