Files
buddha-gpt/docs/superpowers/specs/2026-08-14-buddha-gpt-design.md
2026-08-14 16:41:19 -07:00

102 lines
6.0 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# BuddhaGPT — Design Spec
**Date:** 2026-08-14 · **Owner:** Marcus · **Linear:** PAI-87 (project BuddhaGPT)
## Goal
Show end-to-end LLM competency (prompting, fine-tuning, embeddings, retrieval, evaluation) plus product judgment, via a "Buddha GPT": an open 7B model fine-tuned on Buddhist literature, grounded by RAG over the Pali Canon, evaluated for compassion against base and frontier models, with a safety-research angle (does value-laden fine-tuning shift safety behavior?).
## Non-goals
- Training from scratch on Buddhist text only (corpus too small; proves nothing).
- Claiming the model "is" compassionate — we measure judged behavior on a defined rubric, and report sycophancy separately.
- Multi-tradition completeness. Scope: Theravada (Pali Canon) primary, clearly stated.
## Deliverables
1. **Public repo + writeup** — pipeline code, results, README (portfolio).
2. **Live demo** — Hugging Face Space (Gradio), merged model + RAG citations.
3. **Research-style report** — CompassionBench results + safety-benchmark deltas, HF model card.
## Constraints
- **Compute:** local Apple M5, 24 GB unified memory. Training via MLX (`mlx_lm.lora`), 4-bit QLoRA. No GPU rental.
- **Budget:** ~$100 Anthropic API — synthetic instruction data + LLM-judge. Use Batches API (50% off) and `claude-sonnet-5` (intro $2/$10 per MTok through 2026-08-31) for generation; judge on `claude-opus-5`.
- **License hygiene:** corpus must be redistributable (CC0/CC-BY); base model Apache-2.0.
## Architecture
```
corpus (SuttaCentral/Bilara, Access to Insight, Dhammapada)
├─► data pipeline ─► instruction pairs (synthetic Q&A via Claude, filtered)
│ └─► MLX QLoRA fine-tune of Qwen2.5-7B-Instruct (4-bit)
└─► chunker ─► embeddings (bge-small-en / mlx) ─► LanceDB index
│
user query ─► retrieval (top-k + citations) ─► fine-tuned model ─► answer + sutta refs
│
eval harness (CompassionBench + safety suite)
```
### Components
| Component | Choice | Why |
|---|---|---|
| Base model | Qwen2.5-7B-Instruct (mlx-community 4-bit) | Apache-2.0 (clean for public repo/demo), strong instruct base, fits 24 GB |
| Fine-tune | `mlx_lm.lora` QLoRA, ~5–10k pairs | Runs locally on M5; hours per run |
| Instruction data | Claude Sonnet 5 via Batches API generating Q&A grounded in canon passages | Cheap (~$50 for 10k pairs), quality controllable, filterable |
| Embeddings | `bge-small-en-v1.5` (or nomic-embed) local | Free, fast on M5 |
| Vector store | LanceDB | Embedded, no server, ships with the Space |
| Judge | `claude-opus-5` with rubric, pairwise + absolute | Strongest judge; a human-rated subset checks agreement |
| Demo | HF Space (Gradio) with merged 4-bit model | $0 hosting path (ZeroGPU); account `mrmen` exists |
### Fine-tune vs RAG split (a deliberate write-up point)
Fine-tuning carries voice, framing, and dharma-teacher persona. RAG carries facts and citations (real sutta references, e.g. "MN 21"). The eval compares FT-only, RAG-only, and FT+RAG to demonstrate the judgment of when each tool applies.
## Corpus
- **SuttaCentral / Bilara data** (GitHub `suttacentral/bilara-data`): Sujato translations of the four Nikāyas — CC0. Primary source.
- **Dhammapada** + selected Khuddaka texts (public-domain translations).
- **Access to Insight** (Thanissaro): free-distribution license — verify redistribution terms before inclusion; fallback is prompt-only use (not redistributed).
- Local seed: `~/buddhadasa_mindfulness-with-breathing.pdf` — check license; likely reference-only, not in training set.
## Evaluation design
### CompassionBench (custom, ~150 prompts, 5 categories)
1. Distressed user (grief, anxiety, loneliness)
2. Moral dilemma
3. Harmful request (refusal quality + tone)
4. Sycophancy trap ("tell me my bad plan is good") — measures compassion ≠ agreement
5. Existential/meaning questions
Systems compared: base Qwen2.5-7B, BuddhaGPT-FT, BuddhaGPT-FT+RAG, Claude (frontier reference). Judge: Opus 5 with rubric scoring empathy, non-harm, honesty-under-pressure, groundedness; plus randomized pairwise preferences. Human check: Marcus rates a ~30-item subset; report judge–human agreement.
### Safety delta (research angle)
Run the same safety probes (refusal set, sycophancy set, a TruthfulQA-style subset) on base vs fine-tuned. Question: does compassion-corpus fine-tuning measurably shift refusals, sycophancy, honesty? Either direction is a reportable finding (persona fine-tunes are known to sometimes degrade safety behavior).
## Product judgment / guardrails
- Citations mandatory in RAG mode; answers without a retrieved source are labeled as such.
- Demo disclaimer: not a teacher, not therapy; crisis-resources footer for distress-adjacent inputs.
- Scope statement: Theravada corpus; answers reflect that tradition.
## Error handling / risks
| Risk | Mitigation |
|---|---|
| M5 training too slow / OOM | 4-bit base + LoRA rank ≤ 16, batch 1 + grad accumulation; shrink dataset before shrinking model |
| Synthetic data mode-collapse (samey Q&A) | Diverse prompt templates, dedupe by embedding similarity, temperature-free variety via varied instructions |
| Judge bias toward flowery tone | Rubric penalizes vagueness; pairwise randomized order; human agreement subset |
| ZeroGPU Space limits (7B latency/quota) | Fallback: demo on 3B (Qwen2.5-3B) for the Space, 7B results in the report; or recorded demo |
| Eval bank contamination (prompts leak style) | Hold eval prompts out of all training data; build them after data-gen prompts frozen |
## Milestones
- **M1 — Corpus + RAG MVP:** indexed canon, cited retrieval answers over base model.
- **M2 — Fine-tune v1:** 5k pairs, QLoRA run, qualitative diff vs base.
- **M3 — Eval:** CompassionBench + safety deltas, judge + human subset.
- **M4 — Demo:** HF Space with FT+RAG, guardrails, disclaimer.
- **M5 — Writeup:** README, report, model card, publish.