fix: align collect_frontier token budget with driver (8k)

kimi-k3 spends heavily on hidden reasoning tokens that draw from the same
budget as the visible answer. 3000-token budget resulted in empty responses
on hard prompts. Increased to 8000 to accommodate reasoning overhead while
ensuring room for substantive visible answers.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
marcuspaico
2026-08-17 20:54:34 -07:00
parent b7d0f38ce8
commit e92782639a

View File

@@ -64,11 +64,11 @@ def collect_frontier(items: list[dict], out: Path, model: str = "moonshotai/kimi
for it in items:
if (it["id"], system) in done:
continue
# max_tokens=3000, not 1000: kimi-k3 emits hidden reasoning tokens that
# draw from the same budget as the visible answer even with
# reasoning.enabled=False; 1000 was observed to exhaust on harder
# prompts and return empty content (finish_reason="length").
text, _ = chat(client, model, [{"role": "user", "content": it["prompt"]}], max_tokens=3000)
# max_tokens=8000: kimi-k3 spends heavily on hidden reasoning tokens; 8000
# leaves room for the visible answer — 3000 produced empty responses on hard
# prompts (finish_reason="length"). Hidden reasoning tokens draw from the
# same budget as the visible answer even with reasoning.enabled=False.
text, _ = chat(client, model, [{"role": "user", "content": it["prompt"]}], max_tokens=8000)
if not text or not text.strip():
continue # leave undone for a future rerun rather than write empty content
f.write(json.dumps({"prompt_id": it["id"], "system": system, "response": text}) + "\n")