docs: switch API layer to OpenRouter (deepseek-v4-flash gen, gemini-flash judge, kimi-k3 reference)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
marcuspaico
2026-08-14 17:40:01 -07:00
parent 37caa87b49
commit 6d6fa8f163
2 changed files with 142 additions and 137 deletions

View File

@@ -21,7 +21,7 @@ Show end-to-end LLM competency (prompting, fine-tuning, embeddings, retrieval, e
## Constraints
- **Compute:** local Apple M5, 24 GB unified memory. Training via MLX (`mlx_lm.lora`), 4-bit QLoRA. No GPU rental.
- **Budget:** ~$100 Anthropic API — synthetic instruction data + LLM-judge. Use Batches API (50% off) and `claude-sonnet-5` (intro $2/$10 per MTok through 2026-08-31) for generation; judge on `claude-opus-5`.
- **Budget:** ~$100 ceiling, ~$6 expected — all API via OpenRouter: generation on `deepseek/deepseek-v4-flash-latest` (~$0.08/$0.16 per MTok), judge on `google/gemini-flash-latest`, frontier reference `moonshotai/kimi-k3`, second-judge agreement on `deepseek/deepseek-v4-pro`.
- **License hygiene:** corpus must be redistributable (CC0/CC-BY); base model Apache-2.0.
## Architecture
@@ -43,10 +43,10 @@ user query ─► retrieval (top-k + citations) ─► fine-tuned model ─► a
|---|---|---|
| Base model | Qwen2.5-7B-Instruct (mlx-community 4-bit) | Apache-2.0 (clean for public repo/demo), strong instruct base, fits 24 GB |
| Fine-tune | `mlx_lm.lora` QLoRA, ~5–10k pairs | Runs locally on M5; hours per run |
| Instruction data | Claude Sonnet 5 via Batches API generating Q&A grounded in canon passages | Cheap (~$50 for 10k pairs), quality controllable, filterable |
| Instruction data | DeepSeek V4 Flash via OpenRouter generating Q&A grounded in canon passages | ~$1-2 for ~9k pairs, quality controllable, filterable |
| Embeddings | `bge-small-en-v1.5` (or nomic-embed) local | Free, fast on M5 |
| Vector store | LanceDB | Embedded, no server, ships with the Space |
| Judge | `claude-opus-5` with rubric, pairwise + absolute | Strongest judge; a human-rated subset checks agreement |
| Judge | `google/gemini-flash-latest` with rubric; `deepseek/deepseek-v4-pro` second-judge subset | Cheap, capable, independent of all compared systems; human-rated subset checks agreement |
| Demo | HF Space (Gradio) with merged 4-bit model | $0 hosting path (ZeroGPU); account `mrmen` exists |
### Fine-tune vs RAG split (a deliberate write-up point)
@@ -70,7 +70,7 @@ Fine-tuning carries voice, framing, and dharma-teacher persona. RAG carries fact
4. Sycophancy trap ("tell me my bad plan is good") — measures compassion ≠ agreement
5. Existential/meaning questions
Systems compared: base Qwen2.5-7B, BuddhaGPT-FT, BuddhaGPT-FT+RAG, Claude (frontier reference). Judge: Opus 5 with rubric scoring empathy, non-harm, honesty-under-pressure, groundedness; plus randomized pairwise preferences. Human check: Marcus rates a ~30-item subset; report judge–human agreement.
Systems compared: base Qwen2.5-7B, BuddhaGPT-FT, BuddhaGPT-FT+RAG, Kimi K3 (frontier reference). Judge: Gemini Flash with rubric scoring empathy, non-harm, honesty-under-pressure, groundedness — deliberately independent of every compared system (no self-preference bias). Agreement checks: Marcus rates a ~30-item subset (judge-human) and DeepSeek V4 Pro re-judges a 100-item subset (judge-judge).
### Safety delta (research angle)
@@ -88,7 +88,7 @@ Run the same safety probes (refusal set, sycophancy set, a TruthfulQA-style subs
|---|---|
| M5 training too slow / OOM | 4-bit base + LoRA rank ≤ 16, batch 1 + grad accumulation; shrink dataset before shrinking model |
| Synthetic data mode-collapse (samey Q&A) | Diverse prompt templates, dedupe by embedding similarity, temperature-free variety via varied instructions |
| Judge bias toward flowery tone | Rubric penalizes vagueness; pairwise randomized order; human agreement subset |
| Judge bias toward flowery tone | Rubric penalizes vagueness; judge independent of all compared systems; human + second-judge agreement subsets |
| ZeroGPU Space limits (7B latency/quota) | Fallback: demo on 3B (Qwen2.5-3B) for the Space, 7B results in the report; or recorded demo |
| Eval bank contamination (prompts leak style) | Hold eval prompts out of all training data; build them after data-gen prompts frozen |