docs: record verified tilde-alias model ID in constraints

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
marcuspaico
2026-08-15 14:25:27 -07:00
parent 7b97992b1b
commit cc86a97d10

View File

@@ -14,7 +14,7 @@
- Training runs ONLY on local Apple M5, 24 GB — `mlx_lm.lora` with 4-bit base; no GPU rental. - Training runs ONLY on local Apple M5, 24 GB — `mlx_lm.lora` with 4-bit base; no GPU rental.
- All API calls go through OpenRouter (OpenAI-compatible, base_url `https://openrouter.ai/api/v1`), key from env `OPENROUTER_API_KEY` or file `.openrouter_key` (gitignored). Budget ceiling: $100; expected ~$6. Log token usage from every response. - All API calls go through OpenRouter (OpenAI-compatible, base_url `https://openrouter.ai/api/v1`), key from env `OPENROUTER_API_KEY` or file `.openrouter_key` (gitignored). Budget ceiling: $100; expected ~$6. Log token usage from every response.
- Model roles (exact OpenRouter IDs): data gen `deepseek/deepseek-v4-flash-latest`; judge `google/gemini-flash-latest`; frontier reference system `moonshotai/kimi-k3`; judge-agreement check `deepseek/deepseek-v4-pro`. The judge must never be one of the compared systems. - Model roles (exact OpenRouter IDs): data gen `~deepseek/deepseek-v4-flash-latest` (OpenRouter's catalog ID for the latest-alias — tilde prefix is part of the real ID); judge `google/gemini-flash-latest`; frontier reference system `moonshotai/kimi-k3`; judge-agreement check `deepseek/deepseek-v4-pro`. The judge must never be one of the compared systems.
- Base model: `mlx-community/Qwen2.5-7B-Instruct-4bit` (Apache-2.0). Do not substitute a non-Apache model. - Base model: `mlx-community/Qwen2.5-7B-Instruct-4bit` (Apache-2.0). Do not substitute a non-Apache model.
- Corpus licensing: only CC0/CC-BY/public-domain texts enter training data or the repo. - Corpus licensing: only CC0/CC-BY/public-domain texts enter training data or the repo.
- Eval prompts (Tasks 7–9) are hold-out: never used in data generation or training. - Eval prompts (Tasks 7–9) are hold-out: never used in data generation or training.