A secure, event-driven multi-agent orchestration framework with code-driven capability discovery, SAST-enforced agent admission, asynchronous capability-based I/O, and modular LLM and messaging integrations.
Find a file
2026-10-02 18:57:49 -03:00
identity feat: agent infrastructure v2 — restructuring from core/ monolith into src/vox/ framework 2026-06-02 15:18:52 -03:00
logs feat: agent infrastructure v2 — restructuring from core/ monolith into src/vox/ framework 2026-06-02 15:18:52 -03:00
src/vox fix(tools): support mounted command capabilities 2026-10-02 13:55:52 -03:00
tests docs(tests): remove persona-specific reference from Tool migration notes 2026-10-02 18:57:49 -03:00
tools refactor(security): Vault-only secret resolution and two-gate provisioning 2026-09-22 19:36:26 -03:00
.gitignore release(core): bump version to 0.5.5 2026-08-27 08:46:43 -03:00
AGENTS.md docs(agents): remove stale and duplicated guidance 2026-10-01 13:35:36 -03:00
CHANGELOG.md release(core): bump version to 0.6.5 2026-09-26 00:32:26 -03:00
pyproject.toml release(core): bump version to 0.6.0 2026-09-22 20:31:01 -03:00
README.md refactor(core): close orchestrator architecture debt 2026-09-29 15:17:47 -03:00
requirements.txt arch: checkpoint current architecture specifications for v0.4.0 2026-07-13 00:29:41 -03:00

VOX

VOX is an identity-first orchestration platform for autonomous workloads. It manages long-lived execution units with persistent identity, explicit resource grants, and independent audit storage.

VOX separates reasoning from authority. AI models may produce decisions, but the platform controls what those decisions can act upon. Capabilities, filesystem access, communication paths and execution permissions are governed by the orchestrator, not by the workload itself. This is the central architectural invariant: orchestration owns authority; workloads never own authority.

In VOX, reasoning is advisory; authority is architectural.

Why VOX exists

Autonomous workloads can generate an unbounded sequence of decisions. If the workload carries ambient authority, every decision inherits it, and the only defence is the model's behaviour. Relying on model behaviour does not produce predictable security guarantees.

VOX externalises authority from the workload. The orchestrator grants access to specific capabilities, enforces execution boundaries and records every action in an append-only audit trail. The workload — whether driven by an LLM, a ruleset, or any other decision mechanism — operates within the identity and resource envelope assigned to it.

Architecture

                    ┌─────────────────────┐
                    │   VOXOrchestrator   │
                    │ (authority / fleet) │
                    └─────┬─────────┬─────┘
                          │         │
             ┌────────────┴───┐ ┌───┴────────────┐
             │  VOXWorkload 1 │ │  VOXWorkload 2 │
             │ -------------- │ │ -------------- │
             │  identity      │ │  identity      │
             │  roles         │ │  roles         │
             │  capabilities  │ │  capabilities  │
             └────────┬───────┘ └───────┬────────┘
                      │                 │
                ┌─────┴────┐       ┌────┴─────┐
                │ mounted  │       │ mounted  │
                │ resource │       │ resource │
                └──────────┘       └──────────┘

The runtime exposes two control interfaces:

  • Unix domain socket (/tmp/vox.sock, owner-only) — operator lifecycle commands (start, stop, restart, status, list).
  • HTTP API on port 8000 — fleet queries and lifecycle control for external tooling, authenticated via VOX_API_TOKEN Bearer token.

Design principles

Identity first. Every workload has a persistent identity before execution. Identity is validated at bootstrap and attached to every audit record. No anonymous workloads.

Explicit over implicit. Manifests, roles, lifecycle transitions and operator authentication are configured explicitly. Capabilities are discovered by static analysis of role source code at bootstrap, not by convention or filesystem magic.

Least privilege. A workload has access only to the capabilities, filesystem paths and communication channels required by its roles. Nothing is globally available.

Immutable audit. Every action is recorded in an append-only structured log.

Local-first infrastructure. Backends run on infrastructure the operator controls. No data leaves the local network unless explicitly configured.

Core concepts

Identity

Identity is the foundation of the permission model. Every workload has a UUID, a name, an optional parent reference for hierarchical delegation, and an operator identity enrolled via voice for privileged commands.

Identity is validated at bootstrap. A workload without a valid identity cannot start.

Orchestration

The orchestrator is the single authority for the fleet. It owns every workload's canonical lifecycle state (WorkloadLifecycle), resolves the capability inventory, admits capabilities to workloads at mount time, and routes inbound messages to the workload that declared the source. The orchestrator does not execute workload logic — it governs the conditions under which workloads execute.

Lifecycle is orchestrator-owned: workloads execute the operations the orchestrator requests (initialize, start, stop) and expose no lifecycle state of their own. There is no paused state and no file-watcher hot-restart; the canonical runtime drives lifecycle and control strictly through the orchestrator's public operations.

Capabilities

Capabilities are infrastructure resources — compute, communication, I/O — that the orchestrator manages and assigns to workloads. They are not ambient APIs, and none are mounted implicitly.

Capability requirements are derived from role source code by static analysis: a role requests a capability by using self.workload.capabilities["cap_id"] (or .get(...)). The orchestrator resolves the required set, admits the matching capability instances from its inventory, and the binder mounts exactly that approved inventory — it performs no discovery, health, or lifecycle decisions of its own. A capability the orchestrator did not admit is simply not available to the workload.

Role-level messaging is handled by comm.gateway — a domain-agnostic multi-channel gateway with adapter-driven inbound and outbound (Telegram via webhook or getUpdates long-polling, WhatsApp Cloud API, generic HTTP webhook). It is mounted for a workload only when a loaded role references it, and used via send_text(channel, recipient_id, ...) with an explicit channel and receiver per operation.

Built-in capabilities:

Capability ID Backend
LLM text generation ai.llm Ollama (local) / Gemini (cloud) via explicit adapters
Intent parsing ai.parsing natural-language → CommandSpec via the ai.llm port
Headless browser net.browser Playwright (Firefox)
Messaging (inbound + outbound) comm.gateway Telegram (webhook / getUpdates long-poll) / WhatsApp Cloud API / generic webhook
Email dispatch comm.email SMTP
Speech-to-text comm.voicetotext faster-whisper
Image OCR image.ocr via ai.llm port (vision-capable adapter)

The LLM is one capability among several. Capabilities are interchangeable.

Roles

Roles are behavioural modules attached to a workload. A role exposes command handlers via the @command decorator and subscribes to events via role.on("event"). Only roles listed in the manifest are loaded.

Roles depend on capabilities, not the other way around. A role declares its requirements by using self.workload.capabilities[...] and by declaring the Vault secrets it needs; ASTWorkloadAnalyzer discovers both statically, without executing role code.

Lifecycle

Workloads follow a state machine owned by the orchestrator (WorkloadLifecycle):

INITIALIZING →(config ok) STOPPED →(start) STARTING → RUNNING ⇄(stop) STOPPING → STOPPED

A reported failure while RUNNING moves the workload to REPAIRING; a successful repair returns it to RUNNING, a failed repair to FAILED. Restarting a FAILED workload re-runs configuration and returns it to INITIALIZING. There are no implicit recovery paths: every transition is an explicit event driven by the orchestrator.

The orchestrator initiates and controls all lifecycle transitions. Workloads do not self-start or self-stop. Every transition is an authorisation decision made by the orchestrator. There is no paused state: pause/resume is lifecycle-era behavior that does not exist in the canonical runtime.

Observability and forensics

Every recordable event flows through VOXForensicLogger, which attaches a VOXLogSource (source_type, source_name, source_uuid) to each record and supports dedicated levels including ok. The vox logging hierarchy installed at bootstrap renders every record twice:

  • Console output uses coloured ANSI formatters (VOXColorFormatter).
  • File output uses a machine-friendly plain-text formatter (VOXPlainFormatter) writing to logs/vox.log.

Control plane

A Unix domain socket at /tmp/vox.sock (permissions 0o600) accepts lifecycle commands. The socket is owner-only — no network exposure.

Security model

Security is not a collection of independent features. It emerges from how identity, capabilities, lifecycle, filesystem and communication are structured.

  • Identity enforcement. Every workload must present a valid manifest with name and UUID. Bootstrap rejects workloads without valid identity.
  • Capability gating. Capabilities are admitted by the orchestrator and only admitted capabilities are mounted. A role can only use capabilities that were actually mounted; a capability that was not admitted (missing or denied) is simply unavailable, and no ambient access is granted.
  • Filesystem isolation. File access is scoped to a per-workload sandbox directory. Operations outside the sandbox are rejected at the path level.
  • Control plane isolation. The Unix domain socket is restricted to the owning user.
  • Hierarchical communication isolation. Workloads communicate only with their direct parent or direct children, never laterally.
  • Operator authentication. A pre-enrolled voice embedding (VOXSpeakerProfile) is verified via cosine similarity before privileged operations. All processing is local — no data leaves the machine.
  • Explicit role loading. Only roles listed in the workload's manifest are loaded. Code in the roles directory not listed in the manifest is ignored.
  • Input sanitization. Every inbound payload passes through a pattern scanner (InputSanitizer) before reaching any workload. SQLi, XSS, command injection, path traversal and code execution signatures are blocked at the orchestrator and HTTP API entrypoints.
  • Per-workload rate limiting. Each workload enforces a sliding-window rate limit (RateLimiter) on event emissions. Limits are configurable per workload via manifest keys (rate_limit_max_calls, rate_limit_window); on breach the event is dropped and a critical alert is logged.
  • Append-only audit. Forensic logs cannot be modified or deleted after creation. Every record carries the workload identity.
  • Encrypted secret vault. Per-workload AES-256-GCM vault (WorkloadVault) with PBKDF2HMAC key derivation. Each workload gets an isolated secrets.vault SQLite file. Secrets resolve from the Vault store only — the runtime never falls back to workload config or .env. The workload hosts its vault during initialize and injects each mounted capability's declared secrets at boot; when the vault is unavailable secrets stay unconfigured and capabilities that need them surface a boot failure.

Getting started

Prerequisites

  • Python 3.11+
  • Ollama (local backend for the ai.llm Ollama adapter) or a Google Gemini API key (cloud backend). The Ollama adapter needs no credentials; the Gemini adapter requires the GEMINI_API_KEY Vault secret.
  • Playwright browsers: playwright install firefox
  • Telegram Bot Token — optional, only for the comm.gateway Telegram adapter (if you use Telegram)
  • WhatsApp Cloud API credentials — optional, only for the comm.gateway WhatsApp adapter

Installation

git clone <repo> && cd vox
python -m venv .venv && source .venv/bin/activate
pip install -e .
playwright install firefox

Configuration

Minimal .env in the project root:

TELEGRAM_BOT_TOKEN=...
TELEGRAM_USER_ID=...
Variable Default Purpose
VOX_MASTER_KEY — Master passphrase for the per-workload encrypted vault (WorkloadVault). Without it the vault stays unavailable and capabilities that require Vault secrets surface a boot failure.
VOX_WAR_ROOM_ID — Retired lifecycle-era value. Still read into the resolved config, but no longer required and not consumed by the canonical runtime.
TELEGRAM_BOT_TOKEN — Telegram Bot API token — Vault-provisioned secret for the optional comm.gateway Telegram adapter
TELEGRAM_USER_ID — Telegram chat/user ID injected into every workload config (global workload key) and used as the outbound recipient for comm.gateway Telegram messages
GEMINI_API_KEY — Google Gemini API key — Vault secret for the ai.llm Gemini adapter, provisioned via tools/provision_vault.py (not read from .env). Required only for the Gemini adapter; the Ollama adapter needs no key.
TELEGRAM_LONG_TIMEOUT 25 comm.gateway Telegram getUpdates long-poll timeout (seconds); the HTTP client read timeout is derived from it (+5s buffer)
VOX_API_HOST 127.0.0.1 HTTP API bind address
VOX_API_TOKEN — Bearer token for API authentication
VOX_VERBOSE_LOGGING false Enable verbose debug logging
VOX_UDS_PATH /tmp/vox.sock Unix domain socket path

Adapter credentials for the optional comm.gateway channels are declared per adapter in its config.yml and provisioned per workload through tools/provision_vault.py — e.g. the WhatsApp Cloud API adapter takes WHATSAPP_ACCESS_TOKEN, WHATSAPP_APP_SECRET, and WHATSAPP_VERIFY_TOKEN as Vault secrets, with WHATSAPP_PHONE_NUMBER_ID as an adapter parameter.

The HTTP API always binds to port 8000 (fixed in api_server.py, no env override).

Configuration precedence: CLI args > Environment variables (.env + os.environ) > Code defaults.

Enrolling a speaker identity

python tools/enroll_speaker.py

Records five voice samples and saves the averaged embedding to identity/master_voice.npy.

Running

vox

Starts the orchestrator, reads workload manifests from the filesystem, boots workloads with autostart: true, and serves the control interfaces.

Lifecycle commands

vox start <name>          Boot a workload
vox stop <name>           Stop a workload
vox restart <name>        Restart a workload

Fleet snapshot and registry queries (status, list) are served by the UDS control plane and the HTTP API but are not currently exposed through the vox CLI.

Creating a workload

instance/personas/
  my_workload/
    manifest.yml          name, id, master_id, roles, personality, rate limits
    .env                   optional per-workload configuration override (non-secret only)
    roles/
      handler.py           VOXRole subclass with @command handlers

The manifest declares the workload's identity and which roles to load. Capability requirements are inferred from role source code at bootstrap via AST analysis. Required identity fields: name, id. Secrets required by a workload's roles are provisioned into its vault with tools/provision_vault.py — they are never supplied through .env.

Storage architecture

SQLite storage is fully asynchronous via aiosqlite:

Class File Module
SemanticCache llm_cache.db src/vox/capabilities/ai/llm/cache.py
WorkloadVault secrets.vault src/vox/security/vault.py

WorkloadVault retains synchronous sqlite3 inside __init__ only (one-time bootstrap — schema creation + salt derivation). All runtime public methods (get(), set(), etc.) are async via aiosqlite.

Project structure

vox/
├── instance/personas/                    runtime workload directories
├── identity/                             enrolled operator voice embedding
├── tools/enroll_speaker.py               voice enrolment utility
├── tools/inspect_workload.py             per-workload diagnostics
├── tools/inspect_vault.py                vault inspection / purge
├── tools/provision_vault.py              interactive vault secret provisioning
├── src/vox/
│   ├── cli.py                            CLI entry point, UDS client
│   ├── __main__.py                       python -m vox entry point
│   ├── workloads/                        workload runtime, loader, AST analyzer
│   │   ├── base.py                       VOXWorkload core
│   │   ├── loader.py                     manifest parsing & identity validation
│   │   ├── ast_analyzer.py               static capability dependency scanner
│   │   └── capability_binder.py          mounts orchestrator-approved capabilities
│   ├── capabilities/                     capability definitions and backends
│   │   ├── base.py                       VOXCapability, VOXBoundCapability
│   │   ├── ai/llm/                       provider-agnostic LLM port
│   │   │   ├── capability.py             ai.llm capability entry point
│   │   │   ├── adapters/                 concrete LLM backends (ollama/, gemini/)
│   │   │   ├── cache.py                  semantic response cache (SHA-256 + TTL)
│   │   │   ├── rag.py                    FTS5 retrieval-augmented generation
│   │   │   ├── sanitizer.py              input control-char/boilerplate cleaning
│   │   │   └── models.py                 LLM request/response types
│   │   ├── ai/parsing/                   intent resolution via the ai.llm port
│   │   ├── net/browser/                  headless browser via Playwright
│   │   ├── comm/gateway/                 multi-channel communication gateway
│   │   │   ├── capability.py             CommGatewayCapability wrapper
│   │   │   ├── models.py                 VOXInboundMessage / VOXOutboundMessage
│   │   │   ├── server.py                 IngressServer (shared aiohttp listener)
│   │   │   ├── adapters/
│   │   │   │   ├── base.py               BaseAdapter ABC
│   │   │   │   ├── telegram/             Telegram adapter (+ config.yml)
│   │   │   │   ├── whatsapp/             WhatsApp Cloud API adapter (+ config.yml)
│   │   │   │   └── webhook/              generic HTTP webhook adapter (+ config.yml)
│   │   │   └── capability.yml            capability manifest
│   │   ├── comm/email/                   SMTP email dispatch
│   │   ├── comm/voicetotext/             speech-to-text via faster-whisper
│   │   ├── image/ocr/                    image-to-text via the ai.llm port
│   │   └── file/text_extraction/         file text extraction (scaffold)
│   ├── config/                           configuration loading & resolution
│   │   ├── models.py                     VOXConfig dataclass
│   │   ├── from_env.py                   .env / os.environ parsing
│   │   ├── from_cli.py                   argparse CLI argument parsing
│   │   ├── resolver.py                   merge CLI + env → VOXConfig
│   │   └── loader.py                     load_config() convenience entry
│   ├── messaging/                        inter-workload message envelope schema
│   │   └── models.py                     VOXMessage dataclass
│   ├── observability/                    forensic logger, ANSI formatters
│   │   ├── models.py                     VOXForensicLogger, VOXLogSource
│   │   ├── formatters.py                 colourized console / plain file formatters
│   │   └── constants.py                  custom log levels
│   ├── orchestration/                    lifecycle authority, admission, routing
│   │   ├── orchestrator.py               VOXOrchestrator — fleet coordinator
│   │   ├── lifecycle.py                  WorkloadLifecycle state machine
│   │   └── loader.py                     capability contract/discovery loader
│   ├── roles/                            VOXRole base, @command decorator
│   ├── runtime/                          bootstrap, control plane, daemon
│   │   ├── models.py                     VOXRuntime container dataclass
│   │   ├── daemon.py                     async main loop, start API + UDS
│   │   ├── control_plane.py              UDS request handlers
│   │   └── factory.py                    build_vox() wiring assembly
│   ├── security/                         input sanitizer, rate limiter, vault, voice
│   │   ├── vault.py                      WorkloadVault — AES-256-GCM per-workload store
│   │   ├── guardrails.py                 InputSanitizer — WAF pattern scanner
│   │   ├── rate_limiter.py               sliding-window event rate limiter
│   │   └── speaker_profile.py            VOXSpeakerProfile — cosine-similarity verification
│   ├── services/                         external notification transport
│   │   └── fleet_messenger.py            FleetMessenger — Telegram alert transport adapter
│   ├── api_server.py                     HTTP API (aiohttp, auth + guardrail middleware)
│   └── provider.py                       CapabilityProviderProtocol protocol
├── tests/
└── pyproject.toml

Development

pip install -e ".[dev]"
python -m pytest tests/