From a one-command self-hosted stack to an embedded investigation kernel inside your own services.
Your platform already has telemetry, incidents, and a chat surface. What it lacks is the agent. Vendor the kernel, wire your stack through the ports, and every incident gets a ReAct investigation with audited evidence — under your governance defaults.
import { createAgent, anthropicProvider, staticTools, telemetryTools, httpProbeTool } from './agent-core/mod.ts'; const sre = createAgent({ llm: anthropicProvider({ baseUrl: process.env.LLM_BASE_URL!, apiKey: process.env.LLM_API_KEY!, model: 'glm-4.6' }), tools: [staticTools(...telemetryTools(myTelemetry), httpProbeTool)], store: myRunStore, // your persistence, your schema system: { role: 'SRE on-call for the payments platform', environment: '3 AZ, Envoy edge, Postgres primaries', dataSources: ['query_metrics: prom-backed series'], outputFormat: ['Root cause', 'Evidence', 'Recommendations', 'Risk'], }, }); for await (const ev of sre.run({ runId: incidentId, input: alertText })) { // stream text / tool_use / tool_result / done(answered) to your UI }
One docker compose command starts the whole loop: log/metric collection from every container, a rule engine that opens incidents, automatic AI investigation, and self-recovery after three clean checks. No cluster required — any Docker host works.
docker compose up -d --build # 30s later: dashboard shows live telemetry # alert fires → incident opens → agent investigates → recovers when healthy
Incident data is the most sensitive telemetry you produce. Run the agent against an internal gateway or a local Ollama model; no packet ever leaves your perimeter.
# OpenAI-compatible internal gateway LLM_PROVIDER=openai LLM_BASE_URL=https://llm.internal.corp/v1 LLM_MODEL=qwen3-235b # fully air-gapped openaiProvider({ baseUrl: 'http://ollama.internal:11434/v1', model: 'qwen3:32b' })
Any MCP server becomes a tool source. Unknown tools are treated as write-risk and blocked in read-only mode — map the ones you trust and they flow into the investigation loop.
import { mcpToolSource, stdioTransport } from './agent-core/mcp.ts'; const k8s = mcpToolSource( stdioTransport(['npx', '-y', '@modelcontextprotocol/server-kubernetes']), { riskFor: { list_pods: 'read', get_events: 'read' } }, // unmapped ⇒ write ⇒ blocked );
Every completed investigation can be distilled into a lesson — symptom, root cause, recommendation. The next similar incident starts with those lessons already in context, so the agent gets faster as your team accrues history.
// after a run: distill with the host LLM, store, and it auto-injects next time const lesson = await distillLesson(llm, { input, report, evidence }); await memory.record(lesson); // next run: memory.search(input, 3) → system prompt