In the repository
README · CHANGELOG · Portable agent core spec · AGENTS.md (repo guide)
Everything you need for the all-in-one stack or the kernel. For design rationale and acceptance criteria, see the portable-agent-core spec in the repository.
The stack runs without a key (collection, queries, alerting); the agent needs one.
git clone https://github.com/Joshwong1908/relionaut.git && cd relionaut
cat >> .env <<'EOF'
GLM_API_KEY=your-key # any Anthropic-compatible endpoint
GLM_BASE_URL=https://open.bigmodel.cn/api/anthropic
GLM_MODEL=glm-4.6
EOF
docker compose up -d --build
Open http://localhost:8082. Within 30 seconds the dashboard shows live
telemetry; the alert engine evaluates rules every 60s, opens incidents, and — with a key
configured — runs an automatic investigation on each new incident.
Open any incident and chat: the agent investigates with read-only tools
(query_logs, query_metrics, whitelisted host commands, HTTP probes)
and streams its reasoning. After a run, distill it into a lesson with
POST /api/incidents/:id/distill.
The kernel is a single directory — api/src/core/ — with zero npm dependencies and
zero imports outside itself (enforced by a purity gate in CI). Vendor it into any
Node 24 project (native TypeScript, no build step).
// 1. copy api/src/core → your project (e.g. ./agent-core) // 2. implement the ports you need const myTelemetry: TelemetrySource = { describe: () => ({ logsAvailable: false, metricsAvailable: true, metricNames: ['checkout_error_rate'] }), queryMetrics: async (q) => /* call your time-series store */, }; // 3. assemble and run const agent = createAgent({ llm: openaiProvider({ baseUrl, apiKey, model }), tools: [staticTools(...telemetryTools(myTelemetry))], store: inMemoryRunStore(), // or your own RunStore system: { role, environment, dataSources, outputFormat }, maxSteps: 12, }); for await (const ev of agent.run({ runId, input, mode: 'read-only' })) { … }
Ports you don't implement simply don't exist: a telemetry source without
queryLogs produces no query_logs tool — there is no
"configured but unavailable" state, and the system prompt reflects reality automatically.
Every tool carries a risk tag (single source of truth) and a provenance string. A pure function decides each call inside the loop — the red line cannot be overridden by policy.
| Run mode | read-risk tools | write-risk tools |
|---|---|---|
read-only (default) | execute | always blocked — no policy exception |
advise | execute | blocked; the agent is told to suggest instead |
approve-write | execute | needs a matching Approval for the call id (L4 direction, host UI not shipped) |
MCP tools without an explicit risk mapping are treated as write-risk — fail-closed.
Map trusted tools via riskFor to make them available in read-only runs.
Every run ends in exactly one persisted terminal outcome. Exhausted the step budget? The store
records step_budget_exhausted. The LLM endpoint failed mid-investigation? It records
llm_error. Trajectories are stored as structured, block-paired turns — tool calls
and their results replay in the exact shape the model saw.
// agent_messages row kinds (Postgres host) text | tool_use | tool_result // legacy-compatible audit columns run_outcome // terminal: { outcome, summary, findings } blocks JSONB // structured ChatMessage per turn
| Variable | Default | Purpose |
|---|---|---|
GLM_API_KEY | — | LLM key (required for diagnosis only) |
GLM_BASE_URL | bigmodel Anthropic-compatible | Any Anthropic-compatible endpoint |
GLM_MODEL | glm-4.6 | Model id |
MAX_STEPS | 10 | Max ReAct steps per investigation |
RATE_LIMIT | 60/60 | API rate limit (count/window seconds) |
ALERT_INTERVAL | 60 | Rule evaluation period (seconds) |
AUTO_DIAGNOSE | 1 | Auto-investigate on alert |
Read-only by design: host commands pass a whitelist and a dangerous-character gate; the agent never mutates state and only suggests changes. Before exposing the platform to the public internet, read the auth-tunnel spec — token auth is mandatory, unauthenticated endpoints return 404.
README · CHANGELOG · Portable agent core spec · AGENTS.md (repo guide)
TypeScript strict throughout, Node 24 native TS (zero build), Apache-2.0. CI runs typecheck, the kernel purity gate, unit tests, web build and docker build on every push. Implementation plan & shards.