
Project Intelligence Agent
An agentic RAG system that syncs GitHub issues, PRs, and commits into Postgres + pgvector and answers grounded questions with citations via a LangGraph agent.
Overview
PIA turns fragmented engineering history into a conversational layer - deliberately not an LLM wrapper. A BullMQ sync engine (github-sync → embedding-index) normalizes GitHub data into provider-independent documents, chunks (1000/200), embeds with Gemini, and stores in pgvector. Questions flow through LangGraph: input guardrail → memory retrieval → decompose into 1–5 plans (semantic/activity/hybrid) → parallel retrieve → evidence evaluation → refinement loop (max 2) → grounded generation with [Source N] citations → output guardrail. 35 Vitest tests; 6/6 eval cases ≥ 0.75 with LangSmith traces per run.
What Users Can Do
- Sync any GitHub repo (issues, PRs, commits) idempotently - re-syncs never duplicate.
- Ask natural-language questions about the codebase history with cited sources.
- Resolve temporal queries ('this week', 'last month') over occurredAt/mergedAt ranges.
- Build durable project memory (facts/decisions) that is recalled but never cited as evidence.
- Inspect every run in LangSmith: decomposition, retrieval, evaluation, refinement, generation.
- Run the LLM-as-judge eval suite (6 cases) with scores attached as run feedback.
Why I built this
- To practice retrieval systems where wrong answers cost trust: evaluation + refinement over naive top-k.
- Tradeoff: normalize-before-persist (NormalizedDocument) over raw GitHub schema passthrough - slower ingestion, but retrieval, agent, and UI stay provider-independent (Jira-ready).
- Tradeoff: evidence-evaluation + max-2 refinement loop over single-shot RAG - 2-3 extra LLM calls per hard question, but hallucinations become honest 'could not verify'.
- Tradeoff: guardrails that degrade to safe refusal over 500s - a refused answer beats a confident fabrication.
- To learn pgvector operations: cosine KNN with per-document dedup, 50-chunk context cap, and ANN path for larger corpora.
Tech Stack
After launch & Impact
- Sync engine chains BullMQ jobs reliably: 1k-issue test repo ingested with zero duplicates across 3 re-syncs (unique on sourceType+sourceId+documentType).
- Retrieval capped at 50 unique chunks keeps generation prompts bounded; hybrid plans merge semantic KNN and date-filtered SQL with cross-plan dedup.
- Eval suite scores 6/6 ≥ 0.75 (all 1.00 on seeded facebook/react): technical, activity, refusal-without-fabrication, and memory-vs-evidence separation all pass.
- Memory dedup at ≥ 0.9 cosine and recall at ≥ 0.7 keeps long-running project context without polluting citations.
- 35 Vitest tests cover graph, decomposition, refinement, guardrails, and memory; strict typecheck + lint + build are release gates.
- Honest limits: no auth (CORS + validation only, not for public hosting), single dev GitHub token, credentials in plaintext - auth + encrypted per-workspace credentials are the documented next step.
Future Plans
- Add per-user auth with encrypted per-connection credentials.
- Add Jira provider on the normalized-document abstraction + GitHub App webhooks for incremental sync.
- Add HNSW/ivfflat ANN indexes for larger corpora and SSE streaming answers.
- Expose an MCP tool interface for external agent clients.