OAra Labs · open-source agent daemon

Prometheus

A fire you own doesn't go out.

The agent daemon that stays on in your house and makes the model you already have do real work.

A daemon is a program that never stops running. The word comes from the Greek daimōn, a guiding spirit that stays with you. Prometheus is both.

Runs offline with a local model · Nothing phones home · MIT licensed · One-command install

Armilla — Daemon Telemetry demo feed · canned
spin rate · live work heat0.42 ecliptic nodes · open sessions3 horizon ticks · 24h signals11 / 24 link · sky dims when it dropsup

After Santucci, 1593. The stand is Beacon; the spinning sky is the daemon.

Active development. Expect rough edges. Fixes land weekly. Feedback welcome.

What doesn't work yet ↓

The Stack

Everything stands on Prometheus

Prometheus is the daemon. The pillars plug into it through an open hook contract. Remove them and Prometheus runs exactly as before.

Today

Prometheus Live — Open Source

The agent daemon. The base of the stack.

Beacon Live — Hardening

The desktop cockpit. Sits on Prometheus.

Beacon iOS Private Beta

The cockpit in your pocket. Same daemon, same control plane.

Next

OAra Instinct Designing

Fast decisions. Which model, which tool, when to stop.

Free download · closed source

OAra Cognition Designing

Planning, and checking the work against the goal before delivering it.

Free download · closed source

Blackboard Designing

Coordination across several machines.

Free download · closed source

OAra Instinct
Instinct makes the bounded calls in milliseconds: route this to the small model or the big one, pick the tool, halt a run that’s going nowhere. Rules answer first; nothing else acts unless the rules abstain.
OAra Cognition
Cognition does the slow thinking: plans, assembles context, checks the answer against the goal, in seconds.

Instinct lives with the body, on every node, always local. Cognition thinks with whatever brain Instinct wakes.

Capability router
Matches skills to each request and surfaces the right one only when it’s confident. Measured before it’s switched on.Planned
Adding a pillar
One command, or a toggle in Beacon.Planned

What leaves your machine

Nothing, until you point it somewhere

No telemetry upload, no analytics, no update check: what Prometheus records stays on your disk. What does leave, and when:

your model
Prompts go to the model you configure. A local one keeps them on your network; a hosted one sends them to its provider. Escalating a failed tool call to a cloud model is off until you turn it on.
the web
Search and fetch are on by default: a search sends its query to DuckDuckGo, and a fetch requests the page. Image and video generation, YouTube transcripts, downloads, the browser tool and MCP servers reach out only once you enable them.
your chat app
Telegram, Slack or Discord carry the messages you send through them, once you connect one.
your devices
Beacon talks to your daemon over your LAN or tailnet. Push to Beacon iOS goes through Apple, and it is off by default.
Beacon
Beacon is closed source, so here is what it contacts on its own: your daemon, any integration you connect yourself, and GitHub, to check for a newer Beacon. Its built-in browser goes only where you point it.
first-run downloads
Speech recognition and the skills encoder fetch their model files from Hugging Face the first time they run, if you install them.

Origins

An original codebase, not a fork

Built here; the adapted pieces say where they came from.

Model Adapter Layer
Validates, auto-repairs, and schema-enforces open-model tool calls, retrying with specific error context.
SENTINEL
A proactive layer that watches for idle time and acts — nudges, dreams, synthesizes — instead of only reacting to prompts.
Wiki Knowledge System
Turns every conversation into a compounding knowledge base that cross-references itself over time.
Coding engine
Sandboxed iterate-to-green runs that end in a reviewable diff.
Fine-tuning gym
Captures repair-pairs and golden traces into an exportable training dataset.
SYMBIOTE
Code assimilation and self-modification: GitHub research, license gating, AST-level security scanning, safe grafting with provenance headers, and a blue-green hot swap with automatic rollback. Experimental, off by default.

The Guards

A test that fails the build, not a convention to remember

9,700+ tests passed in CI on the v0.9.5 release commit. A passing test proves the code runs. It does not prove anything calls it. Every guard below exists because a feature was built, tested, green — and never wired.

Enforced byThe invariant
test_no_site_resolves_the_wiki_root_independentlyNo source file may resolve the wiki root on its own — one resolver, or the build fails.
test_every_live_config_key_exists_in_the_default_templateEvery key in the live config must exist in the shipped template — live ⊆ template, with no allowlist. Compares against a real install, so it skips on a fresh clone where there is nothing to compare.
test_no_new_config_key_without_a_readerA register of config keys nothing reads, which can only shrink. An unregistered key with no reader fails — and a registered key that gains one fails as stale, so the list cannot rot.
test_example_call_uses_real_param_namesA tool's advertised example must validate against its own schema. The example ships inside the tool advertisement, so a wrong one teaches the model a parameter that does not exist.
test_every_guard_declares_its_enforcementEvery media and rate check declares fail-closed or fail-open at construction. The registry is built at module level, so an undeclared guard is an import error, not a test failure.
test_tripwire_end_to_endAn acceptance test that terminates in a registered test double fails — registration is what makes a double detectable. Wildcard exemptions are refused; only individually-named doubles pass.
test_every_advertised_document_extension_is_admittedEvery allowlisted file type must be provably admitted, not merely “not refused”. Breach tests prove the door closes; only admission tests prove it opens.

The last one is there because its absence shipped. A control suite whose every case asked “does disabling this let something bad through?” — and none asked “does this let the permitted things through?” — stayed green while the document surface silently degraded to PDF-only: 19 of 20 advertised types refused, including two the allowlist explicitly permitted. Over-refusal looks exactly like the control working.

Behavior has a guard of its own. A change to the daemon’s seams, or to anything that shapes a request to the model, replays recorded turns through a real daemon. Any difference fails the Parity check: a tool call, a gate decision, a checkpoint, a memory write, a telemetry row, the final reply, or a single request sent to the model.

Verification

Nothing moves unless something is true.

At the top of the page that would be a rule to take on faith. Down here it is what the daemon does when it does not know:

the context cliff
When the window is too small, Prometheus refuses the turn and shows you the numbers, instead of truncating it into an answer that quietly lost its beginning. It never refuses to start.
oara doctor
Every ✗ carries a fix: line naming what to change. A red mark with no next step is a mark, not a diagnosis.
oara setup
Setup probes the server and detects the model first. If nothing usable answers, or a cloud provider has no key, no config is written — Prometheus refuses to write a config it knows cannot work.
provider fallback
When the model you asked for fails and another serves the turn, the daemon says so on the wire — a separate provider_degraded frame naming the requested model and the served one. Never a silent swap.
the test suite
Two runs whose gates differ are not compared. The gate manifest returns void — not passed, not failed — rather than a green that proves nothing.
the judge
A reply with no verdict it can parse is not a score. The metric is listed as unavailable — never a pass, never a fail — and a pass rate counts scored tasks only.

Core systems

What the daemon is made of

Agent Loop

Pydantic-validated tool calls, PreToolUse/PostToolUse hook pipeline, permission governance. Works with any model.

Lossless Context Management

DAG-based conversation compression. Every message persisted to SQLite. Old messages summarised into expandable nodes. FTS5 full-text search. Works within 32K context windows.

SENTINEL

Proactive background intelligence. Watches telemetry, sends nudges during idle. AutoDream engine: wiki lint, memory consolidation, telemetry digest, knowledge synthesis. Budget-capped. Never exceeds autonomous trust level.

Wiki Knowledge System

Compounding knowledge base. Memory Extractor runs every 30 min. WikiCompiler builds entity pages with cross-references. Query results file back as new pages. Obsidian-compatible.

Security Gate

Four trust levels: BLOCKED, APPROVE, AUTO, AUTONOMOUS. Configurable deny lists over a floor no mode waives: rm -r at a protected root is refused, and what substring patterns used to catch is left to what actually holds — resource limits and, on Linux, the kernel floor. Workspace boundary enforcement, bash intent analysis, memory security scanning. Token-authenticated REST and WebSocket control plane.

Public Surface Hardening

The chat gateways are the one surface exposed to the internet by design, so they are checked cheapest-and-earliest-first. Per-chat rate budgets under a global ceiling. Declared MIME and the size cap before download — then the download under a hard byte ceiling, because the peer supplied that size and the pre-check believed it. Magic bytes are sniffed after; a declared type that disagrees is refused.

Tokens Redacted Before They’re Kept

Known token shapes become <redacted> before conversation history, memory, telemetry or training data is written. The live conversation is left alone, so a token you paste still works for that session. oara scrub finds what earlier releases kept, and changes nothing until --apply.

Model Adapter Layer

Tool calls that actually work on local models. Validator catches malformed tool calls. Formatter translates between model formats. Enforcer constrains output to valid schemas. Each model runs at a tier: the registry’s, or for a local model it doesn’t know, the one its served chat template implies — and adapter.model_tiers overrides it per model. Retry engine with structured error feedback.

Coding Mode

Point the agent at a coding task and it iterates in a sandbox until the build and tests pass. With LSP enabled, diagnostics feed type errors back into the same loop, so it self-corrects against compiler ground truth. The result is a reviewable diff — nothing touches your branches until you merge.

Fine-Tuning Flywheel

Successful tool-call traces and adapter repair-pairs are captured, stored and mined into an exportable dataset. The data-collection half of a LoRA loop for the local model — training is on the roadmap. Fed by telemetry that never leaves your disk, and one config line turns it off entirely.

Record a Skill

Demonstrate a task once — screen and DOM captured — and the daemon distils it into a reusable skill. Two-tier trust: a recording of the page itself becomes a skill once it passes a quality gate, while one read from the screen alone waits as a draft until you accept it. A screen recording of you doing your job is among the most revealing data you own, and every frame stays on your disk.

Skills That Get Used

The skill tool is offered by default, even in a config setup wrote before it was. Every load is counted: /skills shows when each skill was last used, and the Curator ages skills by real use, never by file date. GEPA — off by default, experimental — proposes improvements from real loads, and nothing changes until a person runs oara gepa promote.

Evals with Judge Provenance

A local LLM judge scores task completion, tool accuracy and hallucination using constrained decoding on your own hardware. Every score records which judge produced it — base URL, model, and whether that model was pinned or auto-detected. Records written before this carry no judge at all: that means unknown, permanently, and they are not compared across paths.

Durable Runtime

Sessions survive daemon restarts. A running turn can be stopped mid-flight. Artifacts produced in chat are downloadable from the session.

Backend Registry

One table of every local inference box you own, probed for the served model, its reported context window, detected vision and latency. A switch is probed first and refused if the box is down; the chat is budgeted at that box's window and the choice survives restarts.

Sessions

A conversation is an object, not a scrollback

You can branch it, delete it for real, roll its files back a turn, and point it at a directory. The security boundary moves with it.

Fork at a point

Copy a conversation up to any message into a new session and continue from there. The original is untouched: its rows, its cursors, its summaries. The branch keeps the real timestamps — rewriting them would make it look like it happened now.

Purge, not just hide

Deleting a session hides it from the list and leaves the rows. Purging is the other answer — the one for “I pasted a customer's details in there”. It clears the full-text index before the row goes, takes the summaries with it, and overwrites the freed pages so the bytes are not in the file. It reports row counts per table, because a purge that only says ok is indistinguishable from one that matched nothing.

Per-turn file checkpoints

For a session with a workspace, every turn starts by capturing the workspace's files before any tool runs. Restore puts back what a turn overwrote, removes what it created, and reports every path. A workspace too big to capture is refused at the log, and the list shows the gap rather than a checkpoint that lies.

A working directory of its own

Point a session at a repo. The run's cwd, the security gate's allowed roots, the project instruction files and the checkpoints all follow it, per run — the boundary is the session's, not the process's.

Restart without amnesia

A restored session's history is always readable after a daemon restart. With sessions.rehydrate on, the next message also restores its recent tail into the live working set, so the model resumes with context. Off by default.

All of it is REST under /api/sessions/<id>/ — fork, purge, workspace, checkpoints — and Beacon drives the same routes. Purge asks you to name the session twice; there is no undo behind it.

The kernel floor

What bash may touch is decided below the tool layer

A permission gate sees a command string. It cannot see cd ~/.ssh && cat id_*, a $HOME indirection, a glob or a heredoc for what they are. So the floor is enforced by the kernel, on the path as the kernel resolves it.

The write floor

Bash's writes stop at the workspace, in a mount namespace (bubblewrap). Needs no root. It is what makes the workspace the complete write domain — and what lets the per-turn checkpoint be a real undo.

The read floor

An AppArmor profile keeps bash out of the secret path families — keys, tokens, credentials — whatever the command string says. Needs root to load, so it is opt-in.

Fail loud, never silent

Set either floor to required and, on a host that cannot provide it, bash refuses to run and says why. It never falls through to an unconfined shell: a floor that quietly isn't there is worse than none, because everything downstream is written as though it is.

Verified by outcome

An exit code of 0 from the wrapper proves the wrapper ran, not that a transition happened. The preflight launches a process through the stack and reads the label the process reports for itself.

Linux only; there is no macOS equivalent. Defaults are honest about that: the write floor is auto — attempted everywhere, and where bubblewrap is missing it degrades once at ERROR and marks every call's result — and the read floor is off. oara doctor tells you which floors you actually have. Keys and modes are in the configuration docs.

Beacon

The desktop cockpit

Beacon is the native desktop client for the daemon — chat with live tool timelines, coding runs and a Loop Manager, documents with AI redlines, Kanban, and per-provider key management across eighteen views. The model picker lists your own boxes, each with its live health, above the cloud presets. It pairs with a one-time six-digit code and works over localhost or your tailnet. The daemon stays the source of truth; Beacon is the window into it. Its Mission view runs the same Armilla you see above, fed real telemetry.

Beacon Mission Control connected to a freshly installed Prometheus daemon
Mission Control, freshly paired.

Beacon has its own page →

OpenAI-compatible API

The second door

Anything that already speaks the OpenAI chat-completions wire — Open WebUI, LobeChat, Continue, Zed — can point at the daemon and get Prometheus: memory, tools, the security gate, your local model. No Prometheus-specific client. Beacon stays the cockpit; /v1 is the door for everything else.

Stateless, like OpenAI

The client sends the whole conversation every call, and each call is one agent turn in a fresh session that persists nothing to memory. A compat client owns its history; the daemon does not learn from it.

Tools run server-side

A request that supplies its own tools or tool_choice is refused with a 400, not silently ignored. The tools are Prometheus's, gated by the security gate, and the client only ever sees text.

The loop owns generation

temperature, top_p, max_tokens, stop and logprobs are ignored on purpose. model is a key from GET /v1/models — your local box, or a cloud preset — and applies to that request alone.

curl http://localhost:8005/v1/chat/completions \
  -H "Authorization: Bearer $PROMETHEUS_API_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"model": "local", "messages": [{"role": "user", "content": "Hello"}]}'

Same bearer token as the rest of the daemon's API; oara token show prints it. Endpoints, refusals and streaming are in the docs.

Architecture

The model is a swappable box

click a model — nothing else changes

gateways
Telegram
Slack
Discord
CLI
Web
desktop client
Beacon (Electron)
messages →
← replies · approvals
prometheus daemon systemd
loopagent loop, hooks, permission governance
sentinelheartbeat, cron, memory extractor, wiki compiler
telemetrywhat Armilla reads — per round, measured or marked unmeasured
reads / writes
~/.prometheus/wiki · memory · skills · workspace
prompt + tools →
← tokens · tool calls
model host · swappable
llama.cpp
any GGUF model, on your GPU · nothing leaves the machine
local
cloud

The dashed boundary is the only part that knows which vendor it is talking to. Gateways, memory, guards and telemetry sit on the daemon's side of that line and do not change when the model does.

on disk
Your machine (or split across two)
│
├── Prometheus daemon (systemd)
│   ├── Agent Loop
│   ├── Telegram / Slack / Discord gateway
│   ├── Heartbeat + Cron
│   ├── Memory Extractor
│   ├── Wiki Compiler
│   ├── SENTINEL
│   ├── LCM Engine
│   ├── SYMBIOTE          GitHub research → safe grafting → blue-green swap
│   └── AnatomyScanner    infra self-awareness
│
├── SQLite databases
│   ├── memory.db
│   ├── telemetry.db
│   └── lcm.db
│
└── ~/.prometheus/
    ├── wiki/             knowledge base
    ├── sentinel/         dream logs
    ├── skills/auto/      learned skills
    └── workspace/        sandboxed execution

Runs single-machine, or split brain and GPU across two boxes. All data in SQLite on your filesystem. Nothing phones home.

Install

Three ways to get it

From PyPI — verified

The packaged release, v0.9.5 on PyPI. The plain package runs the daemon and pairs with Beacon; [full] adds Slack, Discord, MCP, the browser tool, the Anthropic provider and voice output.

pip install 'oara-prometheus[full]'

Homebrew (Apple Silicon) — verified

Apple Silicon (M1 and later) only — Intel Macs aren’t verified yet. It installs in a few minutes, and like the pip install it runs the daemon and pairs with Beacon.

brew install oaralabs/tap/oara

Use the full name: a bare brew install oara fails because the tap isn’t trusted.

Developer path — Git or a clone

For reading or changing the code, or running the test suite: an isolated install straight from the main branch with uv or pipx, or an editable clone.

uv tool install 'oara-prometheus[full] @ git+https://github.com/OAraLabs/Prometheus'
# or: pipx install 'oara-prometheus[full] @ git+https://github.com/OAraLabs/Prometheus'
git clone https://github.com/OAraLabs/Prometheus.git && cd Prometheus
pip install -e '.[full]'

Prerequisites: Python 3.11+, a running llama.cpp / Ollama / LM Studio / vLLM instance (or a cloud API key), and optional bot tokens for the messaging gateways.

I
Install
pip install 'oara-prometheus[full]'   # or Homebrew, or the developer path above
II
Setup
oara setup      # detects llama.cpp / Ollama / LM Studio / vLLM, writes config, smoke-tests
# no local GPU yet? start on a cloud model and switch later:
oara setup --provider anthropic --api-key-env ANTHROPIC_API_KEY --model claude-sonnet-5
III
Run
oara daemon     # always-on: web API + gateways + cron

If anything misbehaves, oara doctor checks every subsystem and prints a fix hint per failure.

On first daemon start a web API token is minted and printed once — re-print it with oara token show.

On Linux, oara install-service writes a systemd user unit so the daemon survives reboots.

Prefer setup from a couch? Skip oara setup and run oara daemon bare — it boots in setup mode and prints a one-time six-digit pairing code. Beacon's wizard does the rest.

pre-1.0 · stated first

What doesn't work yet

Provenance

Built by OAra Labs, starting with early scaffolding from OpenHarness. Design informed by Andrej Karpathy’s LLM Wiki concept, Lossless-Claw, and Sigrid Jin’s analysis of Claude Code’s agent-loop patterns. Full lineage and upstream licenses: NOTICE.