RAG agent with citation grounding

Hybrid retrieval (vector + keyword, merged with reciprocal rank fusion) feeding an agent that verifies its own output sentence by sentence.
  • Policy docs are markdown, one clause per ## Section N heading; upload → parse → split → embed (apps/knowledge/ingest.py, chunking.py), embeddings via Ollama (bge-m3) in Postgres/pgvector.
  • Retrieval is hybrid, not pure vector: apps/knowledge/search.py runs pgvector cosine distance (meaning) and Postgres full-text search (keywords) as two ranked lists, merged with Reciprocal Rank Fusion - score 1/(60+rank) per list, summed, so a clause strong in both wins.
  • Role and effective-date filtering (“which version is in force today, for this employee’s role”) are part of the same query, not a filter bolted on after.
  • Grounding: apps/agents/graph.py:check_sentence() splits the model’s draft sentence by sentence; each needs a real [1]/[T1] reference, and every number in it must appear in that source’s text. Ungrounded sentences are dropped before the employee ever sees them - streaming uses the identical check per sentence as it’s typed.
  • Versioning: a new version is a draft until explicitly published; unchanged sections reuse their embedding; the old version is superseded, not deleted, so past citations still resolve.

LLM gateway

Rate limiting, provider fallback, policy-aware routing and full observability in front of every model call - OpenAI, Claude, Gemini, Groq or local.
  • Routing (apps/llm_gateway/router.py): a pure function. Personal data only ever goes to a provider flagged personal_data_allowed (or local, always allowed); over-budget requests only go local.
  • Rate limiting (ratelimit.py): per team/user requests-per-minute and tokens-per-day budgets in Valkey, checked before any model call leaves the gateway.
  • Fallback: a failure walks to the next provider in priority order. A 429 (quota) gets one bounded retry respecting the provider’s own Retry-After; a 5xx on the last available step also retries once - added after finding a real gap where a dead fallback model turned an ordinary quota blip into a total outage.
  • Circuit breaker (breaker.py), hand-rolled: closed/open/half-open. 5 calls in 30s with ≥50% failing trips it open for 60s; one trial call in half-open decides whether it recloses.
  • Observability: Prometheus counters/histograms for requests, tokens, cost and fallbacks (config/observability.py); OpenTelemetry spans across the whole app, exported to Jaeger.

Multi-tool orchestrator agent

A LangGraph state machine that calls real HR systems as MCP tools, never a direct database import, and can combine several systems in one answer.
  • guard → retrieve → think → act → verify, with a bounded reconsider loop if the model gives up before trying an obvious tool (apps/agents/graph.py).
  • Each HR system (Odoo, OrangeHRM) is a standalone MCP server (bridges/*.py), never a direct import from the Django app. A per-call Ed25519 JWT names the employee; the bridge trusts that identity, never anything the model or the employee typed.
  • Tools are picked by the model via function-calling; one turn can call several tools across several systems and combine them in one answer (manager from Odoo, rating from OrangeHRM).
  • An admin can switch a whole system on or off (apps/connectors/admin.py) with no code change - the agent’s tool list is filtered live by whether its connector is enabled.

Human-in-the-loop approval agent

A write action is a frozen, auditable request that only becomes a real system call after every required human says yes.
  • A write-tier tool call never runs immediately: act() queues it as a pending action with its arguments frozen (apps/approvals/workflow.py).
  • Who approves what is a small declarative table (apps/approvals/rules.py:approval_chain()) by request type - manager for leave, manager+HR for an exception, manager+finance above a reimbursement limit - not scattered conditionals.
  • The employee confirms first, then each approver decides in order; only after the last approval does execute() call the tool - idempotent via a reference tied to the request’s own id, so a retry can never double-book.
  • Every transition (requested, routed, decided, executed or failed) is audited.

Evals and guardrails

Deterministic, not LLM-judged - a known-answer golden set, gated in CI, extended to actually exercise the write path against the real system.
  • A YAML golden set (evals/golden.yaml, 30 cases; evals/golden_hris.yaml, the real systems) scored by exact checks - right citation, right tool, no leaked record, numbers match, hand off when it should (apps/evals/scoring.py). No judge-model flakiness.
  • Two CI gates: a safety gate (access/injection/refusal/actions) on a cheap local model, always; a quality gate on the production model when its key is set, so a fork with no budget still gets some signal.
  • Guardrails reuse the same grounding mechanism as the RAG agent, plus a prompt-injection check at the very first step - blocked before any model call - plus the approval gate for anything destructive.
  • The harness can actually execute a write action end to end against the real system and clean up after itself (a maintenance-only cancel tool, never offered to the agent) - proving the approval chain really lands, not just that it waits.
Built as a proof of concept for an AI agent engineer CV - see the code for the rest.