# BuddhaGPT Fine-tune + RAG experiment: a small local model trained and grounded on the Pali Canon (via [SuttaCentral](https://suttacentral.net)'s Bilara texts). ## Data generation Synthetic instruction pairs (`data/instructions.jsonl`) are generated from `corpus/suttas.jsonl` via `scripts/gen_data.py`, using OpenRouter model `~deepseek/deepseek-v4-flash-latest` (the tilde prefix is part of OpenRouter's real catalog ID for this "latest" alias — verified against the live `/api/v1/models` catalog, not a typo). Token usage / cost: - Original full run (`--mode full`, 3,200 calls, variants 0–1 only): totals were not persisted and the generating process died before a report was written, so these figures are an **estimate**, not measured: ~5.2M input / 1.8M output tokens, ≈$0.4–0.7 at list pricing (~$0.08/M in, $0.16/M out). - Template top-up run (`--mode topup`, 2,300 calls, variants 2–5, 0 failures): **measured** — 1,994,673 input tokens / 1,618,867 output tokens, **$0.42** at list pricing (~$0.08/M in, $0.16/M out). See `.superpowers/sdd/2026-08-14-buddha-gpt/task-5-report.md` for the full fix-round report, including per-template pair counts.