~/anilsingha

AI engineering project

Incident Context

An incident-investigation agent that follows relationships across operational records, preserves the evidence behind every answer, and makes the boundary between confirmed facts and inference visible.

The problem

Similar symptoms do not prove the same root cause.

Incident evidence is spread across services, deployments, configuration changes, earlier incidents, and version-specific runbooks. Keyword matches can find related records, but they cannot safely establish causation.

Engineering decisions

Connected operational model

Models services, deployments, changes, runbooks, and incidents as referenced Sanity documents instead of isolated text fragments.

Evidence trail

Shows the exact relationship path behind each investigation so operators can verify how the answer was reached.

Facts remain distinct

Renders confirmed evidence separately from model inference rather than collapsing both into authoritative-looking prose.

Gemini with structured output

Uses the Gemini SDK and validates answers, evidence, inferences, next steps, and sources through a Zod schema before rendering them.

Failure-aware dependency

Uses MCP Failure Lab to cover delays, hangs, malformed responses, connection loss, and live compatibility with the knowledge endpoint.

Production boundaries

Keeps credentials server-side and adds request limits, timeouts, explicit error states, and a health endpoint.

Investigation contract

Retrieve evidence. Trace relationships. State only what the sources support.

  1. 01

    Retrieve

    The agent queries a Sanity knowledge base of connected operational records through MCP.

  2. 02

    Trace

    It follows explicit references and records the source path behind each material claim.

  3. 03

    Report

    A validated response separates the answer, evidence trail, confirmed facts, inference, next step, and sources.