Case study

Inbox Triage / Follow-up Copilot

End-to-end demo (local build recording)

Evaluation

In progress

Evaluation in progress - validated numbers coming soon.

Evaluation methodology

I labeled synthetic emails (forwarded threads, typos, mixed marketing and recruiter language) into four buckets. Labels were chosen from scenario intent, not from keyword rules alone. A held-out split is reserved for reporting; tuning uses the dev partition only.

Fixtures: eval/fixtures/inbox-labeled.json. Harness calls production triage. Regenerate fixtures: npx tsx scripts/eval/seed-fixtures.mts.

In-house regression: npm run eval (internal splits in eval/results/latest.json, not published on the site).

Problem

During an active job search my inbox mixed recruiter screens, assessment deadlines, calendar nudges, and newsletter noise. I kept missing time-sensitive threads because everything looked equally urgent in a flat list.

Who it is for

Individual contributors and job seekers who need a fast triage pass before they write. Recruiters visiting julianrieder.com get a safe sample queue with no Gmail connection.

What it does

The copilot classifies each message into act now, follow up, waiting, or noise, explains why in plain language, and drafts a short reply for high-priority buckets. Owner unlock adds paste mode, local session persistence, exports, and optional API enhance without changing public demo behavior.

How I built it

Next.js App Router UI with triage logic in TypeScript. Local embeddings plus a logistic and kNN ensemble classify messages without a paid LLM API. Drafts are template maps keyed by message id for the sample corpus, with fallbacks for pasted mail.

Tradeoff: local models are cheap and auditable, but calibration and promotional noise remain active research areas.

Error analysis (in-house held-out)

Internal regression only (not published as headline metrics).

Representative mis-buckets from the in-house eval set. Most mistakes are follow-up vs waiting vs act-now collisions when marketing keywords or deadline phrases appear without real urgency.

  • email-6: expected follow-up, got act-now
  • inb-021: expected follow-up, got act-now
  • inb-025: expected act-now, got waiting
  • inb-036: expected act-now, got waiting
  • inb-043: expected act-now, got waiting

Next steps: add a second pass with lightweight embeddings for ambiguous threads, and downgrade act-now when the body says no action needed or auto-pay is on.

Limitations and next steps

  • No live mailbox sync by design (privacy and demo safety).
  • Pasted emails depend on a simple delimiter format.
  • Newsletter and promotion recall needs more boundary training.

Architecture

[Paste / sample inbox]
        │
        ▼
┌───────────────────┐
│ Rule + score engine │  (client-side, no Gmail API)
└─────────┬─────────┘
          ▼
┌───────────────────┐     ┌─────────────────────┐
│ Bucket + reason   │────▶│ Draft + next action │
└─────────┬─────────┘     └─────────────────────┘
          │
          ▼ (owner + OPENAI only)
┌───────────────────┐
│ /api/.../enhance  │  optional polish, same buckets
└───────────────────┘