Case study

Docs RAG Ask

End-to-end demo (local build recording)

Evaluation

In progress

Evaluation in progress - validated numbers coming soon.

Evaluation methodology

Questions include paraphrases, typos, near-duplicate phrasing, and unanswerable prompts. Retrieval hit means expected doc id in top chunks. Corpus is the public Helio Ops sample only.

Fixtures: eval/fixtures/rag-labeled.json.

In-house regression: npm run eval (internal splits in eval/results/latest.json, not published on the site).

Problem

Internal support and on-call engineers need answers tied to runbooks, not confident guesses. Generic chat UIs hide whether a claim came from documentation.

Who it is for

Ops enablement and AI implementation reviewers who want to see retrieval, citations, and explicit I don't know paths before trusting automation.

What it does

Ask questions against the fictional Helio Ops corpus. Hybrid retrieval ranks chunks, the UI builds an extractive answer with numbered citations, and low-confidence queries return I don't know. Owner unlock allows custom pasted docs and optional LLM synthesis that must preserve citation markers.

How I built it

Documents chunk by paragraph with a character cap. Retrieval blends lexical and embedding scores; answers stitch the best sentences from retrieved chunks. Answerability checks (key-term coverage, chunk margin, local NLI) gate whether an answer is shown.

Tradeoff: refusing out-of-scope questions without hurting in-corpus recall is the main open problem.

Error analysis (in-house held-out)

Internal regression only (not published as headline metrics).

Representative retrieval or grounding misses from the in-house eval set. Lexical search can miss paraphrases and sometimes surface off-topic chunks when unrelated terms overlap Helio doc vocabulary.

  • "Can PII go in exported triage reports?" - retrieval miss, grounded no (expected grounded)
  • "P1 vs P2 difference" - retrieval miss, grounded no (expected grounded)
  • "IDK threshold behavior" - retrieval miss, grounded no (expected grounded)
  • "What is P1? typos: custmer visble" - retrieval miss, grounded no (expected grounded)
  • "RAG must cite when confident" - retrieval miss, grounded no (expected grounded)

Next steps: hybrid retrieval, stricter abstain when top score is close to threshold, and entailment check before showing an answer.

Limitations and next steps

  • Sample corpus only unless owner paste mode is enabled.
  • Out-of-scope refusal rate is still below target on held-out eval.
  • Next: stricter answerability calibration on dev only.

Architecture

[Helio Ops sample docs]
        │
        ▼
┌──────────────────┐
│ Chunk + BM25-ish │  lexical scoring in browser
└────────┬─────────┘
         ▼
┌──────────────────┐     ┌─────────────────────┐
│ Top-k passages   │────▶│ Extractive answer   │
└────────┬─────────┘     │ + [n] citations     │
         │               └─────────────────────┘
         ▼
   below threshold ──▶ "I don't know"
         │
         ▼ (owner + OPENAI optional)
┌──────────────────┐
│ /api/docs-rag/ask│  synthesis must keep cites
└──────────────────┘